We will tell you the truth.
Even when it costs us the account.
← Home / / 6 min read / Glossary

What Is Information Gain?

Information gain is the amount a page adds to what a reader already knows after reading the results ranked above it. It is not originality for its own sake.

Information gain is the amount a page adds to what a reader already knows after reading the results ranked above it. It is not originality for its own sake, and it has nothing to do with word count. A page can be entirely original in its phrasing and still add nothing, because every fact in it was already covered by the top five results.

Why it matters when the results page is already full

Most content briefs are built backwards. Someone exports the headings from the pages currently ranking, merges them into an outline, and hands a writer an instruction to cover all of it slightly better. The output restates the consensus, which is the one thing the results page does not need again.

The consequence is expensive. Pages produced this way rarely fail outright. They rank on page two, hold there, and absorb refresh cycles for years while nobody can say what is wrong with them. Nothing is wrong with them. They are simply redundant, and redundancy is invisible in every quality checklist that scores a page on its own merits rather than against what already exists. A brief that tells a writer what to add, not what to cover changes that economics before a word is drafted. It is the difference between commissioning the sixth version of an article and commissioning the first version of a better one.

How to work out what a page actually adds

Gain is measured against a baseline, so the first job is establishing the baseline honestly rather than assuming it. This is a manual exercise, and no export will do it for you, because the thing you are looking for is defined by its absence from the data. It takes roughly an hour per target query, occasionally two. That hour decides whether the commission is worth placing at all, which makes it the cheapest hour in the whole process.

  1. Read the top ten results properly. Not the headings. The body copy. You are building a list of every claim, number, step and caveat the reader will already have met.
  2. Write the consensus down as a list of assertions. When the same claim appears in seven of ten results, it is table stakes. You still have to cover it, but covering it earns nothing.
  3. Mark the contradictions. Where two ranking pages disagree, you have found a question the results page has not resolved. Resolving it is gain.
  4. Find what none of them have. Original data you can produce, a failure mode nobody documents, a decision rule, a worked calculation, a category of reader the consensus quietly excludes.
  5. Test whether it survives compression. If you cannot state the new contribution in one sentence, the page probably does not have one.
  6. Put the new material high. Gain buried under 600 words of restated background is gain a reader never reaches, and this is the same reasoning that sits underneath whether a page was made for people at all.

The distinction that trips teams up is between novelty and contribution. New wording is not new information. Use the split below when you review a draft.

Gain versus noise

Counts as gainReads as gain but is not
A number you generated from your own dataA number cited from the page ranking third
The condition under which the standard advice failsA longer explanation of the standard advice
A decision rule that resolves a genuine disagreementA summary of both positions without a verdict
A step everyone omits because it is tediousThe same steps in a different order
The cost, in hours or money, of doing it properlyA statement that it takes time and effort

The blunt version

Google holds a patent describing an information gain score for documents. That is the entire evidentiary basis for the term. Google has never confirmed that such a score runs in live ranking, and a granted patent is proof that somebody at a company thought of something, not proof that it shipped. Large firms patent ideas they abandon. Treating a patent as a documented ranking factor is a category error, and the whole vocabulary of information gain in this industry rests on it.

So watch what happens next. Tools now sell you an information gain score. They cannot be calculating the thing in the patent, because the inputs are not public. What they are actually measuring is entity coverage against the pages that rank, which rewards the exact behaviour information gain is supposed to correct. You get scored highly for mentioning everything your competitors mention. That is a similarity metric wearing the name of a novelty metric, sold as a subscription.

Use the idea, ignore the score. Information gain is a writing discipline: before you commission a page, name the thing it will contain that no ranking page contains. If you cannot name it, do not commission the page. That rule survives whether or not the patent was ever implemented, and it pairs with matching what the searcher actually wants rather than what the keyword tool reports.

Example

Say a B2B software company wants to rank for a query about onboarding timelines. The top ten results all say onboarding should be fast, personalised and measured. The brief writer reads all ten, lists the shared claims, and finds one thing missing: nobody states how long onboarding actually takes. The company has that number across 300 of its own accounts, broken down by customer size. The resulting page opens with the median, the spread and the two variables that predict a slow rollout, then covers the consensus advice briefly underneath. The gain is one paragraph of proprietary data. Everything else on the page is table stakes, and the page ranks because of the paragraph, not the rest.

FAQ

Is information gain a confirmed Google ranking factor?

No. Google holds a patent that describes scoring documents for the information they add beyond earlier results, but Google has never confirmed that any such score operates in live ranking. Anyone presenting it as a confirmed factor is describing a patent, not a documented system. Use it as a writing standard instead.

Can a tool measure information gain?

Not the version in the patent, because the required inputs are not publicly available. Tools that advertise a gain score are generally comparing your entity coverage against ranking pages, which rewards similarity rather than novelty. That is useful for spotting omissions and actively misleading if you read it as a measure of what you add.

Does adding information gain mean writing longer pages?

Usually the opposite. Once you separate the consensus material from your actual contribution, most drafts turn out to be mostly consensus. Cutting the restated background and leading with the new material produces a shorter page that adds more. Length is a side effect of the contribution, never the goal.

Related terms

  • Content Pruning — what you do with the archive of pages that were commissioned without a contribution.
  • Product Schema — the structured way to state facts a competitor page has not published.
  • Helpful Content — the broader standard that asks who a page was made for, not what it adds.

If you cannot name, in one sentence, the thing your page will contain that no ranking page contains, you are commissioning a duplicate at full price. Kill the brief and spend the budget on the one page where you can answer that question.

Still here

Want this run on your actual traffic drop?

Send the domain and what you have been told. You get a straight answer. Including the one where we say do not hire us.