How We Measured It: The AI Search Evidence Index Method

TL;DR: This is the method behind The AI Search Evidence Index. It covers the sampling frame, the single retrieval window, what counts as a sourced claim, how reliable the coding is, what the instrument got wrong on the way, and every file published with the study.

The findings are in the study itself. This page is the apparatus: how the pages were chosen, how they were fetched, what the instrument does and does not count, how well it agrees with a careful reading, and where every file lives. It is separated so the findings can be read in one sitting — not so the method is easier to skip.

How We Measured It: The AI Search Evidence Index Method

What has been measured before

Citation practice has been studied for decades, but almost entirely inside academic publishing. A meta-analysis of twenty-eight studies of medical journals put the total quotation error rate, meaning citations that do not support the claim attached to them, at 25.4% (95% CI 19.5–32.4%). That work assumes a world where citations exist and asks whether they hold up.

Commercial web content has not been measured the same way, and the question there is one step earlier: is there a citation at all?

The nearest well-documented failure mode is circular reporting, and its internet variant citogenesis: an unsourced claim is repeated by a source that looks independent, and the repetition then becomes the citation. Wikipedia's account of circular reporting notes why the online form is hard to catch: modern pages are revised quickly, citations rarely carry an "as of" date, and pages rarely carry a reliable "last updated" one.

That observation shaped two decisions here. Every page in this study was retrieved inside a single 84-second window, and the dataset publishes a cryptographic hash of each document so a reader can tell whether the page they are looking at is the page we scored.

The field being measured has its own foundational study, the paper that named generative engine optimization and tested which content tactics raise visibility in AI answers. Notably, one of its top-performing tactics is citing sources. That is part of what makes the gap interesting: the field's own research says citing sources helps, and the field's own guides mostly do not link them.

What we got wrong the first time

We ran this measurement twice, and the first run produced a different answer.

The first frame used two seed queries and 19 pages. It gave a competitor median of 5%, and the story wrote itself: nobody in this field sources anything. We nearly published that.

Then we widened the frame to six queries and 38 pages. The median moved to 41%.

The first sample had not measured the field. It had caught a cluster of pages that happened not to link sources, and 19 pages was not enough for that to average out. Nothing was wrong with the instrument; the sample was too small to carry the claim we were about to make with it.

Widening the frame moved the number sharply up. Three refinements that followed each moved it back down:

Change Median
Two queries, counting every number as a claim 5%
Six queries 41%
Counting blocks rather than individual numbers 40%
After all 90 sourced blocks were read individually and 10 rejected 33%

The unit change mattered most for the leaders. Counting each number separately let one link take credit for every figure in its paragraph, which rewarded pages that write number-dense prose. On a block basis three pages lost between 10 and 19 points, while the highest-scoring page gained four.

The headline finding survived all four versions. The median moved from 5% to 41% to 40% to 33%, and the spread stayed wide throughout. The within-publisher comparison is younger than that: the first round took one page per site, so it could not be measured there at all. That is the part we would defend.

How we measured this

Sampling frame. Six seed queries on google.com (hl=en&gl=us), 30 August 2026. Every organic result on page one that carried an h3 heading. Publishing platforms (Reddit, Quora, LinkedIn, YouTube, Medium, Substack) and sponsored results excluded. Four of the six queries are ones our own pages target; two were carried over from an earlier round. So most measured pages compete with one of ours.

That produced 38 URLs across 30 domains. A domain enters as many times as Google ranked it; we did not hand-pick extra pages. The full frame, with every URL, is published with the data.

Retrieval. All 38 were fetched back to back inside 84 seconds. Thirty-two came back; six failed — four to HTTP 403 bot protection and two to a connection error. All six are listed by name and status in the published retrieval log. They are not a random six, and their exclusion is discussed below.

What was measured. Of the 32 pages retrieved, 25 carry a measurement row: the instrument skips a page carrying fewer than three blocks in total, counting claim blocks and price blocks together, and seven fell there. (segmetrics.io clears that cut on two claim blocks plus one price block, though it stays below the threshold for a percentage.) Of those 25, thirteen carry ten or more claim blocks and are given a percentage.

Unit. A block is one observation: a paragraph, list item, table row or quote of at least sixty characters containing at least one numeric claim. Six kinds of block are excluded before anything is counted — headings, date lines, the boilerplate templates a page hands its readers to copy, author biographies, "N out of M" phrasings, and a trailing date stamp — and every year is stripped from the text before the instrument looks for a number, so a block whose only figure is a year is not a claim. Counting headings was one of the errors the instrument's first revision fixed. An earlier version counted every number separately, which let a single link take credit for five figures and rewarded pages that write number-dense paragraphs. Switching to blocks moved twelve of the thirteen scored pages: it cost three of them between 10 and 19 points, moved seven up by at most six, and left one unchanged. Across all twenty-five measured pages the movement is wider — one page gained thirteen points, another lost ten. Both versions ship with the data.

Coding. The instrument's verdicts were checked against a blind reading of 120 blocks drawn at random from the 470 across the 32 retrieved pages: 93.3% agreement, Cohen's κ = 0.85 (95% CI 0.75–0.95).

That pool is deliberately wider than the one the table is built on. It applies only two of the instrument's filters — sixty characters, and a numeric claim — so it also contains blocks the instrument would have thrown out. Thirty-five of the 120 are such blocks. This makes the test harder than the instrument's own job, not easier, and the number holds either way: on the 85 blocks the instrument would actually score, agreement is 94.1% and κ = 0.86 (95% CI 0.74–0.98). We report the wider figure as the headline because it is the more conservative of the two. On the conventional reading of κ (Landis & Koch, 1977) both sit in the "almost perfect" band — which is a label for a number, not a warrant that the instrument is right. Lombard, Snyder-Duch & Bracken (2002), the standard reference for reporting reliability in content analysis, asks for the coefficient together with enough about the coding instrument for a reader to judge it. We publish the instrument and every coded item as well, which is more than the checklist asks and less than a second coder would be. A separate twelve-block check by a second person, on an earlier round and on the named-source question only, matched the model on eleven; twelve items is too small to publish as a rate and it is not reported as one. That mismatch produced a codebook rule, disclosed rather than backdated.

The instrument went through seven revisions before the version that produced this table, and the history of all seven is published with the data — along with an eighth entry, added after publication, recording a repair to the code's file handling that changed no measured number: what the error was, how it surfaced, and what changed. Three of those errors would have put a wrong number next to a named company.

The instrument counted a Start Free Trial button as a source, because the sentence beside it mentioned a "14-day trial." It counted Read Post cards under an article. It counted a page's own product page as sourcing four of that page's figures. It also failed to match according to when the A was capitalised, a case-sensitivity bug introduced at the second revision that survived two more before it was caught.

Each of those was found by reading pages, not by testing code.

Data, code and materials. Everything the figures rest on is published: the sampling frame, the retrieval log, both measurement files, every coding sheet, the instrument and its revision history. This follows the availability standard journals now apply to research articles (Nature Portfolio). The page snapshots themselves are held back for copyright reasons; their SHA-256 hashes are published so anyone can check a re-fetch against what we measured.

Use of AI tools. The measurement instrument, the retrieval, the coding and the first draft of this article were produced with Claude (Anthropic). The author designed the study, set its scope, reviewed the output at each stage, and takes full responsibility for the content, including the parts produced with AI assistance. Where a figure in this article depends on someone reading a page rather than a script parsing it, the reader was the model unless stated otherwise.

Cost. Between 27 August and 2 September 2026 this measurement — the retrieval, seven revisions of the instrument, the blind coding, the independent re-implementation, the hand-verification of all 90 sourced blocks, and the two verification rounds that followed and rewrote much of this article — took 3.05 billion tokens across 5,742 model calls. Across everything on this project since June, including the earlier self-audit that produced the instrument and every study before it, 9.6 billion tokens across 21,121 model calls since 10 June 2026.

Those are tokens processed, and 98.4% of that is cached context being read again rather than new work. Newly written output was 5.53 million. The daily breakdown is published as data/compute-log.json, counted from this project's session logs; the logs themselves are not published, but the arithmetic is, and the totals move only upward as work continues. We report this the way compute is reported in machine-learning papers: not as a claim to rigour, but so the scale sits on the record next to the result.

Verification. Every block the instrument credited with a source on the thirteen scored pages was opened and read individually: 90 of them. The thirteen it credited on the twelve sub-threshold pages were not read, because those pages carry no published percentage. Ten were overturned on inspection, leaving the 80 that stand in the table, and each rejection is listed in the published materials by page and by URL. Note which way that review can move a number: it can only take credits away, never add one. A block the instrument passed over was never re-examined. So the published percentages are the instrument's, minus what a reading could disprove — and any figure the instrument missed is still missing. The counter was then re-implemented from scratch by the same model, without reusing any of the original code, as a check — a test of whether the counting is robust to how it is written, not a second opinion from a second party. The re-implementation has no price rule, so the comparison has to be made on totals and sourced counts with price blocks left in; on that basis seventeen of twenty-five pages matched exactly, and every difference was traced to a known cause. Compared on totals alone the figure is nineteen of twenty-five, which is why the basis matters and is stated here.

How reliable is the coding?

Start with what this check is not. The blocks were re-read by the same language model that wrote the instrument, blind to the instrument's verdict. Blindness is real — the verdicts were held in a separate file and opened only after coding — but authorship is not divided. This measures whether the code does what its own codebook says, not whether the codebook is right. The one human check below was coded by the author of this article. Neither is an independent reading in the sense that phrase usually carries, and no figure here should be read as one.

The check. Of the 470 claim-carrying blocks across the 32 pages we retrieved, 120 were drawn at random with a fixed seed, and coded without seeing what the instrument had decided about any of them. The coder saw the block's text, the links inside it and the publisher's domain. The instrument's verdict was withheld.

Blocks 120
Agreement 112 of 120, 93.3%
Cohen's κ 0.85
95% confidence interval 0.75 – 0.95

All 120 blocks are published with their text, their links and the code each was given, so a reader who disagrees with a judgement can find it and say so.

Which way it errs. The disagreements are not symmetrical. In seven of the eight, the instrument credited a source where the blind read found none; in one, it missed a source the read found. The instrument therefore errs towards making a page look better sourced than it is. An earlier version of this article, written on a sample of thirty, reported the opposite direction. That was wrong, and the larger sample corrects it.

A smaller human check, on a different question. In an earlier round, thirty blocks were coded by the model and sealed, and the author of this article then coded twelve of them blind. Eleven of twelve matched. Two things limit what that shows. Twelve items is far too small to publish as a rate, and we do not. More importantly, every one of those thirty blocks had already been established to contain no link at all, so the person was judging only whether a source was named, never whether a link reached it. It is a check on the second tier, not on the percentages in this article. The single mismatch produced a codebook rule: a publisher naming itself is not naming a source unless it points to a specific identified work.

Claim detection. The selection itself was checked. Every block in the 120-block sample carries a field recording whether it is a numeric claim at all, so this is recomputable from the published sheet. Two rates come out of it, and they answer different questions:

not a claim applies to
whole 120-block sample 32 — 27% (95% CI 19–35%) the wider pool, including blocks the instrument filters out
the 85 the instrument would score 19 — 22.4% (95% CI 13–31%) the blocks behind the published percentages

The second is the one that matters for the table, and it is the one used in the correction below. The first is higher because the wider pool contains material the instrument was built to discard — which is what the sample was drawn to test. An earlier pass over 50 blocks gave a consistent 22%, but its per-item verdicts were not kept, so it stands as corroboration rather than as evidence a reader can check.

The two errors run in opposite directions. On the link side the instrument is generous, which pushes the published percentages up. On the claim side it counts things that are not claims, and those sit almost entirely on the unsourced side, which pushes the percentages down. We cannot net them off against each other and we do not try to. Both are stated so a reader can carry the uncertainty in the direction that matters to them.

Data and code

The sampling frame, the measurement tables, the instrument, the independent re-implementation, the codebook and the coding sheets are published under CC BY 4.0:

File What it is
results.csv The final table, one row per published page, machine verdict and human verdict side by side
retrieval-log.csv All 38 URLs including the 6 that failed, with retrieval timestamps and status
snapshot-hashes.csv SHA-256 hash of every snapshot we scored
sampling-frame.json The six queries and every organic result they returned
measure.py The instrument
independent-counter.py The independent re-implementation used to check it
coder1-sealed-codes.json Coder 1's sealed codes, with the SHA-256 they were sealed under
blind-sample-120.json The 120 blocks coded blind against the instrument, each with its text, its links and the code it was given
blind-sample.py The script that drew the sample, with its seed
README Method, codebook and limitations in full

The same files are also published at github.com/minelgunesoglu-code/ai-search-evidence-index, where every change made after publication is visible in the commit history. If we correct a figure, you can see what it was before and why it changed. We would rather you did not have to take our word for that either.

The working notes behind each decision (the design as written before measuring, the revision history, the hand-verification log) are in notes/. They are in Turkish; the README carries the method in English.

We do not republish the pages themselves. Copying someone's article in full in order to criticise how they cite is not a thing we want to do, and the guidance on text and data mining draws the same line: analysing content is one question, redistributing it is another.

So the dataset carries something else instead. For every page we measured it records the URL, the retrieval timestamp, the character count, and the SHA-256 hash of the exact document we scored.

A hash is a fingerprint. Identical file, identical fingerprint; change one character and it changes completely. Fetch any of these URLs today, hash what comes back, and compare:

  • Same hash. You are looking at the document our number describes.
  • Different hash. The page has changed since 30 August 2026, and you now know that without having to take our word for anything.

Here are the thirteen pages in the table above, with the fingerprint of the document each number was taken from. The first sixteen characters are enough to tell a match from a mismatch; the full 64-character hash for all thirty-two retrieved pages is in snapshot-hashes.csv.

Page Characters First 16 characters of the SHA-256
industry-lens.com 171,142 a433d4582df8eaf7
tryprofound.com 233,257 dfdba698c69b3a2b
tryprofound.com-4 255,626 339461e4717e818e
promptzero.tech 60,958 72909dbd006fa72f
writer.com 273,523 418fc6b9cd234a02
frase.io 440,500 667829639b0101b0
ahrefs.com 990,735 4a24d9592ea6e5cb
aisearch.similarweb.com 657,785 f7ecc7bbe9e61ddf
semrush.com 252,925 18e99bb65b5555e3
seocrawl.ai 285,397 4016d2cf7ec98579
tryprofound.com-2 382,549 e0b92f1fe77055bb
llmrefs.com 167,049 c1ccb00c95e6b5e1
zapier.com 650,110 949857d72c1f6df0

To check one yourself, fetch the URL the same way the study did — same user agent, or several of these sites return a bot-block page instead — and hash what comes back:

UA="Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 \
(KHTML, like Gecko) Chrome/131.0 Safari/537.36"

curl -sL -A "$UA" "https://ahrefs.com/blog/how-to-rank-in-ai-overviews/" \
  | shasum -a 256

A caveat on that command: code/fetch.py wrote each page through Python's text layer, so the saved bytes are the re-encoded UTF-8 of the response, not the raw stream. On a page that has not changed the hashes match; if yours differs by a byte or two on an otherwise identical page, that is why. Comparing the two documents will tell you more than comparing the two hashes.

Short excerpts appear in the coding sheets where they are needed to show why a block was scored the way it was, capped at 200 characters in the coding sheets and 700 in the reliability sample. Measured against each page's extracted body text, the published excerpts are a median of 5.6% of it — but the share is uneven, and on two pages it is not small: writer.com 14% and industry-lens.com 39%, the latter because it is a short page from which many blocks were sampled. Any publisher who wants an excerpt removed can write and we will remove it.

We also publish the instrument's revision history: seven revisions, each recording the error found, how it surfaced, and what changed. Three of those errors would have put a wrong number next to a named company.

If you find an error, tell us. We will correct it and say what changed.

References