Two numbers in three have no link to a source
A measurement of whether a reader can reach the source of a numeric claim. Open data, open method, the instrument's full revision history, and every credited block on the scored pages opened and read one at a time.
Open the data and codeDo AI Visibility Guides Link Their Sources? We Sampled 38 and Scored 13
TL;DR: This is round one of the AI Search Evidence Index. We sampled 38 pages that rank for AI-visibility queries, retrieved 32 of them inside one 84-second window, and scored the 13 that carried enough claims to rate. The median page linked a source for 33% of them. The range ran from 0% to 62% across publishers, and from 6% to 55% across the rateable pages of a single one.
The pages that teach you how to get cited by AI are full of numbers. Citation rates, engine counts, correlation coefficients, market shares, prices. We wanted to know a narrow thing about those numbers: can a reader reach the source?
Not whether the figures are right. Whether they are checkable.
The AI Search Evidence Index is a recurring measurement of whether a reader can reach the source of a numeric claim on the pages that teach AI visibility. This is its first round. Naming it matters: the figures below will be quoted, and a reader who meets one of them elsewhere should be able to ask which study it came from.
It is a first round, but not a first attempt. The instrument behind it has been in development since
June 2026, through earlier measurements, a self-audit that turned it on our own pages first, and
seven revisions in which it was wrong and had to be corrected. That whole line of work, including
this round and the two verification passes that followed it, comes to 9.6 billion tokens across
21,076 model calls, counted from the project's own logs and published as
data/compute-log.json. We report it for the same
reason we report everything else here: so the scale of what produced a number sits next to the
number. The revisions, and what each one got wrong, are set out below.
The AI Search Evidence Index
- How often do AI visibility guides link a source?
- Does the same publisher cite sources consistently?
- Which pages linked the most sources?
- Is sourcing a price the same as sourcing a statistic?
- What counts as a source?
- Does naming a source without linking it count?
- What does this mean for a reader checking a number?
- What this measure cannot tell you
- How did our own pages score, and why are they excluded?
- What has been measured before
- Why did the first run give a different number?
- How this was measured
- How reliable is the coding?
- Data and code
- References
- How to cite this
How often do AI visibility guides link a source?
Across 13 pages with enough claims to score, the median page linked a source for 33% of its numeric claims. Across all of them pooled, 80 of 252 claim-carrying blocks (31.7%) carried a link to the figure's source.
That is the headline number, and it is the least interesting thing we found.
Does the same publisher cite sources consistently?
Pages ranged from 0% to 62%. Two pages linked nothing. One linked most things.
Then we looked at publishers that appeared more than once, and the spread got stranger.
tryprofound.com ranked for four of our six queries and entered with four pages:
| tryprofound.com page | Sourced | |
|---|---|---|
| Top experts in generative engine optimization | 6 of 11 | |
| What is answer engine optimization | 5 of 10 | |
| Best AI visibility tools for marketing agencies | 1 of 17 | |
| How ChatGPT sources the web | 0 of 7 | below threshold, not a rate |
The last row is under the ten-block floor this study sets for a percentage, and it was not among the blocks read by hand. It is shown for completeness, not as a score. The finding stands on the three rateable pages: one publisher, 55% on one, 50% on another, 6% on a third.
semrush.com entered the same way (four pages across four queries), but only one of them clears
the floor, so it cannot corroborate this. It is the only other publisher that could have, and it
cannot. That is a limit on the finding, not a footnote to it: the within-publisher result rests on
one company.
What it suggests is that sourcing is not a house style, not a policy a company sets and its writers follow. Something varies page to page inside one publisher. We cannot say what, because we measured pages and not the people or the deadlines behind them. But a reader who trusts a domain because one of its pages was well sourced has learned less about the next one than they think.
Which pages linked the most sources?
Every page below was retrieved inside the same 84 seconds, so no page had an advantage of time over another. The exact timestamps and a hash of each document are published with the data.
| Page | Claim blocks | Sourced | Share |
|---|---|---|---|
| industry-lens.com: AI search intelligence | 21 | 13 | 62% |
| tryprofound.com: top experts in GEO | 11 | 6 | 55% |
| tryprofound.com: what is answer engine optimization | 10 | 5 | 50% |
| promptzero.tech: GEO guide | 15 | 6 | 40% |
| writer.com: GEO, AEO and SEO | 35 | 14 | 40% |
| frase.io: what is GEO | 48 | 19 | 40% |
| ahrefs.com: how to rank in AI Overviews | 15 | 5 | 33% |
| aisearch.similarweb.com: what is GEO | 24 | 8 | 33% |
| semrush.com: best AI SEO tools | 18 | 2 | 11% |
| seocrawl.ai: AI Overview ranking factors | 17 | 1 | 6% |
| tryprofound.com: best AI visibility tools for agencies | 17 | 1 | 6% |
| llmrefs.com: generative engine optimization | 11 | 0 | 0% |
| zapier.com: best AI visibility tools | 10 | 0 | 0% |
Twelve further pages carried fewer than ten claim blocks and are not given a percentage; a figure built on four blocks is not a rate.
One page needs its context stated. industry-lens.com leads the table. Its format attaches a
source by default, which is a different thing from a guide that has to go and find one. Its links
split about evenly between its own /reports/ pages and outside sources: in the published sample,
nine of its seventeen links are its own, so this is a difference of format, not a page citing only
itself.
Two of this study's rules also happen to favour it at once. It is one of the three price-heavy pages, and a link to a pricing page is the one case that passes without the anchor test. A reader who wants to discount the top row has grounds to; we would rather hand them those grounds than have them found. The number is real. The comparison is not like-for-like, and we would rather say so than let the ranking imply otherwise.
Is sourcing a price the same as sourcing a statistic?
Not every number on these pages is the same kind of number.
Some pages are tool comparisons, and most of their figures are prices and feature counts, "$99/mo", "ten engines". For those, sourcing means linking the vendor's pricing page. Other pages are guides, and their figures are research statistics: "68% of searches end without a click". For those, sourcing means linking the study.
The two are not the same task. Checking a price takes one click to a page that either states it or does not. Checking a statistic means finding a study that may not be public.
In this round, three of the thirteen scored pages carry five or more price claims. They scored 62%, 11% and 0%. The other ten are guides, with a median of 36%.
Three pages is not enough to say anything about the genre, and we are not going to. We note the
split because a reader comparing zapier.com with frase.io should know they are looking at two
different jobs, and because anyone repeating this measurement should probably separate them from
the start.
What counts as a source?
One rule decides almost everything: the link has to lead to the figure.
What makes a block sourced
A block scores as sourced when a link inside it points to where the number came from. The instrument never opens those links. It reads the saved copy of the page and judges each link from its address and its anchor text alone. So what is measured is whether a reader is pointed at a source, not whether that source is still reachable, and not whether it says what is claimed.
Table rows
One rule reaches outside the block. Where a short linked paragraph, under 400 characters,
sits within two blocks of a run of table rows, its link credits every row in that run, on the
reasoning that a source note placed above a table covers the table beneath it. It is applied
identically to every page and it is visible in
code/measure.py; it is also one of the two causes of
the eight disagreements with the independent re-implementation reported below. We state it here
because the sentence above says a link inside the block, and for table rows that is not the
whole rule.
The anchor test
One test does most of the work: the anchor text has to contain a number that appears in the sentence making the claim, or share at least two content words with it. Both halves are cruder than they sound, and it is worth being exact about how.
The number test runs against the raw sentence, so a year counts as a number, an anchor reading "…(2026 guide)" can in principle pass on a sentence that mentions 2026. We checked: stripping every year from every sentence changes none of the blocks the instrument credited across the twenty-five measured pages. The weakness is real in the code and did not bite in this corpus.
The second half is looser. There is no rule against an anchor that is simply a name: a one-word brand
cannot pass, because the test needs two shared words, but a two-word name can. ahrefs.com carried a
figure attributed to "Ahrefs alumni Joshua Hardwick" with his LinkedIn profile linked, and the
instrument counted it. So did the blind re-read. It was removed only when the block was opened and
read by hand, and it is one of the ten rejections listed in the published materials. That is the case
for reading the credited blocks rather than trusting the count, and it is why the table reports 80
rather than 90.
Where the rule favours us
That rule is not neutral, and it favours us. Writing anchor text that names the finding is good practice, and it is the style our own pages were rewritten into on 30 August, before this round ran. Several of the measured publishers anchor on a brand name instead, and a brand-name anchor passes only by accident, when the same name happens to appear in the sentence too. So the rule that decides most of this measurement rewards the house style of the people who wrote it. We are not in the table, which limits the damage, but a reader comparing two publishers should know that the one who anchors on figures is being scored on a convention as much as on whether the source is reachable.
What does not count
Several other things that look like sources are not. The examples below come from the earlier round in which these rules were written, not from the thirteen pages in the table; the rules were drawn while reading other people's pages and then applied unchanged here:
- A sign-up or demo link. One page's three "sourced" claims were all a Start Free Trial button whose surrounding sentence happened to contain "14-day trial."
- A related-post card. Read Post boxes under an article are navigation.
- A vendor's homepage. It does not take you to the price you just quoted; its pricing page does.
- The publisher's own product page. One page linked its own product four times for four different figures.
- A pricing page, where the claim is itself about price; this one is an exception that passes without the anchor test, because quoting a price and linking the price list is a complete act.
Two further kinds (a help document, and an author's LinkedIn profile) carry none of the findings
attributed to them, but the instrument has no rule for either. They were removed by hand, on one
page (ahrefs.com), while its credited blocks were being read. The twelve sub-threshold pages were
not read, so links of this kind may still be counted there.
Price claims
Price claims are counted separately and are not in the percentages. A block whose numbers are
all prices goes into its own bucket, because sourcing a price means linking a vendor's pricing page,
which is a different act from sourcing a research finding. Across the twenty-five measured pages
that removes forty-six blocks, seventeen of them sourced. On the thirteen scored pages it removes
eighteen blocks, four sourced. One page shows why this matters in both directions: orchly.ai is
recorded at 0 of 6 on research claims while linking a vendor page for thirteen of its twenty-five
price claims. The price counts ship in the data so a reader can put them back.
A source named but not linked ("according to Gartner", "Ahrefs' own research") is recorded separately and is not included in the percentages above. The blind reading measured the sourced/not-sourced call only, so this round produced no agreement figure for the second tier at all. Where it matters, we say a page names sources without linking them, and give a verified example.
Does naming a source without linking it count?
A page can name where a number came from without linking it: "according to Gartner", "Ahrefs' own research", "a Princeton-led study". The reader can find that source, but has to go looking.
We record this separately and do not fold it into the percentages, for a reason worth stating: we cannot show that we judge it reliably. The blind reading of 120 blocks tested the sourced/not-sourced call, not this one, so there is no agreement figure for it from this round. The only measurement we have is the twelve-item second-coder check from an earlier round, which matched on eleven, far too small to publish as a rate. So we report this tier as a flag with verified examples rather than as a number, and a reader should treat it as the weaker of the two.
Examples, each one published in the coding sheet with the block it came from, its links and its code:
llmrefs.com: "Vercel reports that 10% of new signups now come from ChatGPT referrals"tryprofound.com-2: "backed by a 25-million-person consumer panel", of a named competitorwriter.com: "Only 16% of brands track their AI search performance (McKinsey)"
The last one shows the limit of measuring by block. That block carries no link, but four blocks earlier the same page makes the same claim and links the McKinsey study directly. Writer named its source and linked it; our unit of observation is the block, so the block that repeats the figure without the link is counted as one. We flag it here rather than let a reader find it in our own coding sheet and conclude we did not look.
One page makes the case for keeping this tier separate. In an earlier round of this
measurement, ayzeo.com scored 0%: no claim on the page carried a link to its source. But the
page ends with a nine-entry bibliography: BrightEdge, Reuters twice, Associated Press, Frase,
SingleGrain and others, in full academic style. None of them are hyperlinks.
A reader can absolutely find "Reuters (2024), Reddit in AI content licensing deal with Google." By any reasonable standard that is a sourced claim. Our measure, which reads links block by block, cannot see it. Publishing "0%" for that page and stopping there would have been the single worst thing in this study.
So we ran the scan again on this round, before publishing these numbers. All 32 retrieved pages
were checked for a source list in their final third, and every hit was read by hand
(code/bibliography-scan.py). One page carried such a section, and it was not a bibliography:
ahrefs.com ends with Further reading: two Search Engine Journal articles and a schema
validator, none of them the source of a figure in its body. Its 33% stands.
The two pages published at 0% were read to their last block. llmrefs.com ends with a list of
takeaways, zapier.com with related-post cards and a product call to action. Neither collects its
sources anywhere. That matters more than the rest of this paragraph: a 0% next to a company's name
is the hardest thing this study says, and it is not an artefact of where they put their links.
In the previous round it went the other way: three pages carried an end-of-page bibliography and one of them linked nothing from it. Rare, then, but real, which is why the check now runs every round rather than when we remember.
What does this mean for a reader checking a number?
Sourcing is not a house style. One publisher's rateable pages ran from 6% to 55%, a 49-point spread inside a single company, against a 62-point spread across the whole field. Whatever governs whether a number gets a link, it is not operating at the level of the organisation. We flag the weakness of this one ourselves: it rests on a single publisher, because no other one in the sample had more than one page clearing the threshold.
The practical consequence is for readers. Trusting a domain because one of its articles was well sourced tells you nothing about the next article on the same domain. The unit that earns trust is the page, and it has to earn it again every time.
The field's own research recommends the thing the field mostly does not do. The study that gave generative engine optimization its name found that citing sources was among the highest-performing content tactics for visibility in AI answers. The pages teaching that tactic link a source for about a third of their own figures.
We are not going to call that hypocrisy. A likelier explanation is ordinary: linking a source is work, it happens at the end of writing, and nothing on the page breaks when it is skipped. Nobody sees the missing link. That is exactly the kind of quiet defect a measurement is for.
What we did not test. Whether any of this affects citation by AI engines. It would be a reasonable hypothesis, since engines that reward verifiable content might prefer pages that link their sources, but we did not measure it and this study is not evidence for it. Testing it would take a different design: the same pages, tracked against actual citation over time.
That is the round we would run next.
Everything after this point is the apparatus: how the measurement was built, what it got wrong on the way, how reliable the coding is, and (in What this measure cannot tell you) the four things these numbers do not show. If you take one number from this article, take it with that section.
What this measure cannot tell you
It is not an accuracy check. A page can source every figure and still be wrong, or source nothing and be right. We measured reachability.
Our claim detector counts things that are not claims. The blind sample was coded for this too.
Of the 85 sampled blocks the instrument would actually score, 19 (22.4%, 95% CI 13–31%) were
advice, author biographies or examples of good writing rather than claims. Only 3 of those 19 were
credited with a source, so they sit mostly in the unsourced side and the percentages above run
low. Applying that rate to the 252 blocks in the table removes about 56 non-claims, 9 of them
credited: the pooled figure corrects from 31.7% to roughly 36%. Every input to that arithmetic
is in the published coding sheet, where every block carries an instrument_would_score flag: count the
non-claims among the 85 it marks and you get the 19. Counting all 120 gives 26.7% instead, because
the sample deliberately includes blocks the instrument never sees. We report the raw number as
measured and the correction as an estimate.
Two known biases, pulling opposite ways. The correction above moves the figure up. The reliability check moves it down: across the 120 blind-coded blocks the instrument called 36.7% of them sourced where the blind read called 31.7%, so it credits about five points it should not. The non-claim correction is worth about four points in the other direction. The two are measured on different populations and we do not net them into a single number, but they are close in size and opposite in sign. That is why the figure we report is the one we measured, and why neither adjustment is offered on its own as a better estimate.
Six pages could not be retrieved. Most are large and established (Search Engine Land twice,
Adobe, TechnologyAdvice), but not all: otterly.ai is a small vendor. A rough browser-side
reading of five of them came out at 8%, 14%, 11%, 71% and 100%. Two of those sit far above the
median and three below it, so we cannot say which way including them would have moved the number.
They are out because they were never retrieved inside the single window, not because of what they
would have done to the result.
One day, one snapshot. In an earlier round, one page lost 181 footnote links between two retrievals a day apart. That is why we publish the SHA-256 hash of every page we measured: you can re-fetch any URL, hash what you get, and see for yourself whether it still matches the document behind our number.
Six queries are not a field. Everything here is stated as "these pages, on this date."
How did our own pages score, and why are they excluded?
This is not a neutral comparison study, in the sense Boulesteix, Lauer and Eugster set out: we compete with the pages we measured, on the same queries. The instrument was built here, and while it was being built we could fix our own pages; the measured publishers could not. We ran ours first, linked our sources, fixed five claims that pointed at a page which no longer carried the figure, and re-measured. That is a real advantage, and it is why our four pages are reported here rather than in the table.
On the block unit the table uses, they score 91%, 75%, 74% and 33%. Naming them, because a score you cannot attach to a page is the same defect this study is about:
| Our page | Per block | Per number |
|---|---|---|
| How to get cited by ChatGPT | 91% | 90% |
| Generative engine optimization | 75% | 85% |
| Best AI search visibility tools | 74% | 74% |
| How to track brand mentions in ChatGPT | 33% | 24% |
Placed in the table they
would take three of the top five rows, and the fourth would sit around tenth of seventeen. We are not putting them there, because a number produced under
a condition none of the other pages had is not a comparable number. Both units are published in
notes/unit-of-analysis.md, where they appear under
the codenames OURS-cited, OURS-geo, OURS-tools and OURS-track; a reader who disagrees with that call can merge them and see what
happens.
The exclusion rules cost us nothing, since we do not use sign-up buttons or homepages as sources, and we found those rules while reading other people's pages, not our own. The per-page count of what each rule removed is published, so a reader can judge whether the rules were drawn fairly, though that count covers the earlier round in which the rules were written, not this round's thirteen pages.
What has been measured before
Citation practice has been studied for decades, but almost entirely inside academic publishing. A meta-analysis of twenty-eight studies of medical journals put the total quotation error rate, meaning citations that do not support the claim attached to them, at 25.4% (95% CI 19.5–32.4%). That work assumes a world where citations exist and asks whether they hold up.
Commercial web content has not been measured the same way, and the question there is one step earlier: is there a citation at all?
The nearest well-documented failure mode is circular reporting, and its internet variant citogenesis: an unsourced claim is repeated by a source that looks independent, and the repetition then becomes the citation. Wikipedia's account of circular reporting notes why the online form is hard to catch: modern pages are revised quickly, citations rarely carry an "as of" date, and pages rarely carry a reliable "last updated" one.
That observation shaped two decisions here. Every page in this study was retrieved inside a single 84-second window, and the dataset publishes a cryptographic hash of each document so a reader can tell whether the page they are looking at is the page we scored.
The field being measured has its own foundational study, the paper that named generative engine optimization and tested which content tactics raise visibility in AI answers. Notably, one of its top-performing tactics is citing sources. That is part of what makes the gap interesting: the field's own research says citing sources helps, and the field's own guides mostly do not link them.
Why did the first run give a different number?
We ran this measurement twice, and the first run produced a different answer.
The first frame used two seed queries and 19 pages. It gave a competitor median of 5%, and the story wrote itself: nobody in this field sources anything. We nearly published that.
Then we widened the frame to six queries and 38 pages. The median moved to 41%.
The first sample had not measured the field. It had caught a cluster of pages that happened not to link sources, and 19 pages was not enough for that to average out. Nothing was wrong with the instrument; the sample was too small to carry the claim we were about to make with it.
Widening the frame moved the number sharply up. Three refinements that followed each moved it back down:
| Change | Median |
|---|---|
| Two queries, counting every number as a claim | 5% |
| Six queries | 41% |
| Counting blocks rather than individual numbers | 40% |
| After all 90 sourced blocks were read individually and 10 rejected | 33% |
The unit change mattered most for the leaders. Counting each number separately let one link take credit for every figure in its paragraph, which rewarded pages that write number-dense prose. On a block basis three pages lost between 10 and 19 points, while the highest-scoring page gained four.
The headline finding survived all four versions. The median moved from 5% to 41% to 40% to 33%, and the spread stayed wide throughout. The within-publisher comparison is younger than that: the first round took one page per site, so it could not be measured there at all. That is the part we would defend.
How this was measured
Every number above rests on a method, and it is set out in full below: the sampling frame, the single 84-second retrieval window, what the instrument counts as a sourced claim and what it throws out, how well it agrees with a careful reading, the seven revisions it went through and what each one got wrong, and every file in the dataset.
Sampling frame
Six seed queries on google.com (hl=en&gl=us), 30 August 2026. Every organic
result on page one that carried an h3 heading. Publishing platforms (Reddit, Quora, LinkedIn,
YouTube, Medium, Substack) and sponsored results excluded. Four of the six queries are ones our own
pages target; two were carried over from an earlier round. So most measured pages compete with one
of ours.
That produced 38 URLs across 30 domains. A domain enters as many times as Google ranked it; we did not hand-pick extra pages. The full frame, with every URL, is published with the data.
Retrieval
All 38 were fetched back to back inside 84 seconds. Thirty-two came back; six failed: four to HTTP 403 bot protection and two to a connection error. All six are listed by name and status in the published retrieval log. They are not a random six, and their exclusion is discussed below.
What was measured
Of the 32 pages retrieved, 25 carry a measurement row: the instrument skips
a page carrying fewer than three blocks in total, counting claim blocks and price blocks together,
and seven fell there. (segmetrics.io clears that cut on two claim blocks plus one price block, though it stays below the threshold for a percentage.) Of those 25, thirteen carry ten
or more claim blocks and are given a percentage.
Unit
A block is one observation: a paragraph, list item, table row or quote of at least sixty characters containing at least one numeric claim. Six kinds of block are excluded before anything is counted: headings, date lines, the boilerplate templates a page hands its readers to copy, author biographies, "N out of M" phrasings, and a trailing date stamp; and every year is stripped from the text before the instrument looks for a number, so a block whose only figure is a year is not a claim. Counting headings was one of the errors the instrument's first revision fixed. An earlier version counted every number separately, which let a single link take credit for five figures and rewarded pages that write number-dense paragraphs. Switching to blocks moved twelve of the thirteen scored pages: it cost five of them, three by between 10 and 19 points and two by fewer than four, moved seven up by at most six, and left one unchanged. Across all twenty-five measured pages the movement is wider: one page gained thirteen points, another lost ten. Both versions ship with the data.
Coding
The instrument's verdicts were checked against a blind reading of 120 blocks drawn at random from the 470 across the 32 retrieved pages: 93.3% agreement, Cohen's κ = 0.85 (95% CI 0.75–0.95).
That pool is deliberately wider than the one the table is built on. It applies only two of the instrument's filters (sixty characters, and a numeric claim), so it also contains blocks the instrument would have thrown out. Thirty-five of the 120 are such blocks. This makes the test harder than the instrument's own job, not easier, and the number holds either way: on the 85 blocks the instrument would actually score, agreement is 94.1% and κ = 0.86 (95% CI 0.74–0.98). We report the wider figure as the headline because it is the more conservative of the two. On the conventional reading of κ (Landis & Koch, 1977) both sit in the "almost perfect" band, which is a label for a number, not a warrant that the instrument is right. Lombard, Snyder-Duch & Bracken (2002), the standard reference for reporting reliability in content analysis, asks for the coefficient together with enough about the coding instrument for a reader to judge it. We publish the instrument and every coded item as well, which is more than the checklist asks and less than a second coder would be. A separate twelve-block check by a second person, on an earlier round and on the named-source question only, matched the model on eleven; twelve items is too small to publish as a rate and it is not reported as one. That mismatch produced a codebook rule, disclosed rather than backdated.
Revisions
The instrument went through seven revisions before the version that produced this table, and the history of all seven is published with the data, along with an eighth entry, added after publication, recording a repair to the code's file handling that changed no measured number: what the error was, how it surfaced, and what changed. Three of those errors would have put a wrong number next to a named company.
The instrument counted a Start Free Trial button as a source, because the sentence beside it
mentioned a "14-day trial." It counted Read Post cards under an article. It counted a page's own
product page as sourcing four of that page's figures. It also failed to match
according to when the A was capitalised, a case-sensitivity bug introduced at the second revision
that survived two more before it was caught.
Each of those was found by reading pages, not by testing code.
Data, code and materials
Everything the figures rest on is published: the sampling frame, the retrieval log, both measurement files, every coding sheet, the instrument and its revision history. This follows the availability standard journals now apply to research articles (Nature Portfolio). The page snapshots themselves are held back for copyright reasons; their SHA-256 hashes are published so anyone can check a re-fetch against what we measured.
Use of AI tools
The measurement instrument, the retrieval, the coding and the first draft of this article were produced with Claude (Anthropic). The author designed the study, set its scope, reviewed the output at each stage, and takes full responsibility for the content, including the parts produced with AI assistance. Where a figure in this article depends on someone reading a page rather than a script parsing it, the reader was the model unless stated otherwise.
Cost
Between 27 August and 2 September 2026 this measurement (the retrieval, seven revisions of the instrument, the blind coding, the independent re-implementation, the hand-verification of all 90 sourced blocks, and the two verification rounds that followed and rewrote much of this article) took 3.03 billion tokens across 5,706 model calls. Across everything on this project since June, including the earlier self-audit that produced the instrument and every study before it, 9.6 billion tokens across 21,076 model calls since 10 June 2026.
Those are tokens processed, and 98.4% of that is cached context being read again rather than
new work. Newly written output was 5.53 million. The daily breakdown is published as
data/compute-log.json, counted from this project's session logs; the logs themselves are not
published, but the arithmetic is, and the totals move only upward as work continues. We report this
the way compute is reported in machine-learning papers: not as a claim to rigour, but so the scale
sits on the record next to the result.
Verification
Every block the instrument credited with a source on the thirteen scored pages was opened and read individually: 90 of them. The thirteen it credited on the twelve sub-threshold pages were not read, because those pages carry no published percentage. Ten were overturned on inspection, leaving the 80 that stand in the table, and each rejection is listed in the published materials by page and by URL. Note which way that review can move a number: it can only take credits away, never add one. A block the instrument passed over was never re-examined. So the published percentages are the instrument's, minus what a reading could disprove, and any figure the instrument missed is still missing. The counter was then re-implemented from scratch by the same model, without reusing any of the original code, as a check, a test of whether the counting is robust to how it is written, not a second opinion from a second party. The re-implementation has no price rule, so the comparison has to be made on totals and sourced counts with price blocks left in; on that basis seventeen of twenty-five pages matched exactly, and every difference was traced to a known cause. Compared on totals alone the figure is nineteen of twenty-five, which is why the basis matters and is stated here.
How reliable is the coding?
What this check is not
The blocks were re-read by the same language model that wrote the instrument, blind to the instrument's verdict. Blindness is real: the verdicts were held in a separate file and opened only after coding, but authorship is not divided. This measures whether the code does what its own codebook says, not whether the codebook is right. The one human check below was coded by the author of this article. Neither is an independent reading in the sense that phrase usually carries, and no figure here should be read as one.
The check
Of the 470 claim-carrying blocks across the 32 pages we retrieved, 120 were drawn at random with a fixed seed, and coded without seeing what the instrument had decided about any of them. The coder saw the block's text, the links inside it and the publisher's domain. The instrument's verdict was withheld.
| Blocks | 120 |
| Agreement | 112 of 120, 93.3% |
| Cohen's κ | 0.85 |
| 95% confidence interval | 0.75 – 0.95 |
All 120 blocks are published with their text, their links and the code each was given, so a reader who disagrees with a judgement can find it and say so.
Which way it errs
The disagreements are not symmetrical. In seven of the eight, the instrument credited a source where the blind read found none; in one, it missed a source the read found. The instrument therefore errs towards making a page look better sourced than it is. An earlier version of this article, written on a sample of thirty, reported the opposite direction. That was wrong, and the larger sample corrects it.
The human check
This one is smaller, and it asks a different question. In an earlier round, thirty blocks were coded by the model and sealed, and the author of this article then coded twelve of them blind. Eleven of twelve matched. Two things limit what that shows. Twelve items is far too small to publish as a rate, and we do not. More importantly, every one of those thirty blocks had already been established to contain no link at all, so the person was judging only whether a source was named, never whether a link reached it. It is a check on the second tier, not on the percentages in this article. The single mismatch produced a codebook rule: a publisher naming itself is not naming a source unless it points to a specific identified work.
Claim detection
The selection itself was checked. Every block in the 120-block sample carries a field recording whether it is a numeric claim at all, so this is recomputable from the published sheet. Two rates come out of it, and they answer different questions:
| not a claim | applies to | |
|---|---|---|
| whole 120-block sample | 32, 27% (95% CI 19–35%) | the wider pool, including blocks the instrument filters out |
| the 85 the instrument would score | 19, 22.4% (95% CI 13–31%) | the blocks behind the published percentages |
The second is the one that matters for the table, and it is the one used in the correction below. The first is higher because the wider pool contains material the instrument was built to discard, which is what the sample was drawn to test. An earlier pass over 50 blocks gave a consistent 22%, but its per-item verdicts were not kept, so it stands as corroboration rather than as evidence a reader can check.
The two errors together
The two errors run in opposite directions. On the link side the instrument is generous, which pushes the published percentages up. On the claim side it counts things that are not claims, and those sit almost entirely on the unsourced side, which pushes the percentages down. We cannot net them off against each other and we do not try to. Both are stated so a reader can carry the uncertainty in the direction that matters to them.
Data and code
The whole package is archived at CERN's Zenodo under 10.5281/zenodo.22257563: the concept DOI, which resolves to the most recent version deposited, currently v1.0.1 (10.5281/zenodo.22300386). The version under which these figures were first published is 10.5281/zenodo.22257564. The archive is a third party's copy: if this site changed a figure tomorrow, the version behind the DOI would not move with it.
The files
The sampling frame, the measurement tables, the instrument, the independent re-implementation, the codebook and the coding sheets are published under CC BY 4.0:
| File | What it is |
|---|---|
| results.csv | The final table, one row per published page, machine verdict and human verdict side by side |
| retrieval-log.csv | All 38 URLs including the 6 that failed, with retrieval timestamps and status |
| snapshot-hashes.csv | SHA-256 hash of every snapshot we scored |
| sampling-frame.json | The six queries and every organic result they returned |
| measure.py | The instrument |
| independent-counter.py | The independent re-implementation used to check it |
| coder1-sealed-codes.json | Coder 1's sealed codes, with the SHA-256 they were sealed under |
| blind-sample-120.json | The 120 blocks coded blind against the instrument, each with its text, its links and the code it was given |
| blind-sample.py | The script that drew the sample, with its seed |
| README | Method, codebook and limitations in full |
The same files are also published at github.com/minelgunesoglu-code/ai-search-evidence-index, where every change made after publication is visible in the commit history. If we correct a figure, you can see what it was before and why it changed. We would rather you did not have to take our word for that either.
Working notes
The working notes behind each decision (the design as written before measuring, the
revision history, the hand-verification log) are in
notes/. They are in Turkish; the README carries
the method in English.
Copyright
We do not republish the pages themselves. Copying someone's article in full in order to criticise how they cite is not a thing we want to do, and the guidance on text and data mining draws the same line: analysing content is one question, redistributing it is another.
The hashes
The dataset carries something else instead. For every page we measured it records the URL, the retrieval timestamp, the character count, and the SHA-256 hash of the exact document we scored.
A hash is a fingerprint. Identical file, identical fingerprint; change one character and it changes completely. Fetch any of these URLs today, hash what comes back, and compare:
- Same hash. You are looking at the document our number describes.
- Different hash. The page has changed since 30 August 2026, and you now know that without having to take our word for anything.
Here are the thirteen pages in the table above, with the fingerprint of the document each number was taken from. The first sixteen characters are enough to tell a match from a mismatch; the full 64-character hash for all thirty-two retrieved pages is in snapshot-hashes.csv.
| Page | Characters | First 16 characters of the SHA-256 |
|---|---|---|
| industry-lens.com | 171,142 | a433d4582df8eaf7 |
| tryprofound.com | 233,257 | dfdba698c69b3a2b |
| tryprofound.com-4 | 255,626 | 339461e4717e818e |
| promptzero.tech | 60,958 | 72909dbd006fa72f |
| writer.com | 273,523 | 418fc6b9cd234a02 |
| frase.io | 440,500 | 667829639b0101b0 |
| ahrefs.com | 990,735 | 4a24d9592ea6e5cb |
| aisearch.similarweb.com | 657,785 | f7ecc7bbe9e61ddf |
| semrush.com | 252,925 | 18e99bb65b5555e3 |
| seocrawl.ai | 285,397 | 4016d2cf7ec98579 |
| tryprofound.com-2 | 382,549 | e0b92f1fe77055bb |
| llmrefs.com | 167,049 | c1ccb00c95e6b5e1 |
| zapier.com | 650,110 | 949857d72c1f6df0 |
Checking a hash yourself
To check one yourself, fetch the URL the same way the study did (same user agent, or several of these sites return a bot-block page instead) and hash what comes back:
UA="Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 \
(KHTML, like Gecko) Chrome/131.0 Safari/537.36"
curl -sL -A "$UA" "https://ahrefs.com/blog/how-to-rank-in-ai-overviews/" \
| shasum -a 256
A caveat on that command: code/fetch.py wrote each page through Python's text layer, so the saved
bytes are the re-encoded UTF-8 of the response, not the raw stream. On a page that has not changed
the hashes match; if yours differs by a byte or two on an otherwise identical page, that is why.
Comparing the two documents will tell you more than comparing the two hashes.
Excerpts
Short excerpts appear in the coding sheets where they are needed to show why a block was scored
the way it was, capped at 200 characters in the coding sheets and 700 in the reliability sample.
Measured against each page's extracted body text, the published excerpts are a median of 5.6% of
it, but the share is uneven, and on two pages it is not small: writer.com 14% and
industry-lens.com 39%, the latter because it is a short page from which many blocks were sampled.
Any publisher who wants an excerpt removed can write and we will remove it.
Revision history
We also publish the instrument's revision history: seven revisions, each recording the error found, how it surfaced, and what changed. Three of those errors would have put a wrong number next to a named company.
If you find an error, tell us. We will correct it and say what changed.
References
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2024). GEO: Generative Engine Optimization. KDD '24.
- Lombard, M., Snyder-Duch, J., & Bracken, C. C. (2002). Content Analysis in Mass Communication: Assessment and Reporting of Intercoder Reliability. Human Communication Research, 28(4), 587–604.
- Landis, J. R., & Koch, G. G. (1977). The Measurement of Observer Agreement for Categorical Data. Biometrics, 33(1), 159–174.
- Boulesteix, A.-L., Lauer, S., & Eugster, M. J. A. (2013). A Plea for Neutral Comparison Studies in Computational Sciences. PLOS ONE, 8(4), e61562. The source of the three criteria for a neutral comparison, the second of which this study does not meet.
- Jergas, H. and Baethge, C. (2015). Quotation accuracy in medical journal articles: a systematic review and meta-analysis, PeerJ 3:e1364. The source of the 25.4% quotation error rate quoted above, pooled across twenty-eight studies.
- Nature Portfolio. Reporting standards and availability of data, materials, code and protocols.
- Wikipedia. List of citogenesis incidents and Circular reporting, on how unsourced claims acquire citations by repetition, and why undated web pages make it hard to detect.
- Authors Alliance (2024). Text and Data Mining under U.S. Copyright Law, on the distinction between analysing a corpus and redistributing it, which is why this study publishes document hashes rather than the documents.
- Our own Cross-Engine Citation Study, August 2026 round, for the engine-behaviour figures referenced above.
How to cite this
The dataset is archived at CERN's Zenodo, which pins a copy no one can quietly edit, including us.
Gunesoglu, M. (2026). The AI Search Evidence Index: source determinability of numeric claims on AI-visibility web pages (Round one) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.22257563
The DOI above resolves to the most recent version deposited, currently v1.0.1. To cite the version these numbers were first published under, use 10.5281/zenodo.22257564.