The AI Search Evidence IndexResearchOpen dataset

Two numbers in three have no link to a source

34 min
Read time
28
Data files published
9.6B
Tokens spent

Round 1 · Dataset frozen 30 August 2026, 20:38:51–20:40:15 UTC

A measurement of whether a reader can reach the source of a numeric claim. Open data, open method, the instrument's full revision history, and every credited block on the scored pages opened and read one at a time.

Open the data and code

Do AI Visibility Guides Link Their Sources? We Sampled 38 and Scored 13

TL;DR: This is round one of the AI Search Evidence Index. We sampled 38 pages that rank for AI-visibility queries, retrieved 32 of them inside one 84-second window, and scored the 13 that carried enough claims to rate. The median page linked a source for 33% of them. The range ran from 0% to 62% across publishers, and from 6% to 55% across the rateable pages of a single one.

The pages that teach you how to get cited by AI are full of numbers. Citation rates, engine counts, correlation coefficients, market shares, prices. We wanted to know a narrow thing about those numbers: can a reader reach the source?

Not whether the figures are right. Whether they are checkable.

The AI Search Evidence Index is a recurring measurement of whether a reader can reach the source of a numeric claim on the pages that teach AI visibility. This is its first round. Naming it matters: the figures below will be quoted, and a reader who meets one of them elsewhere should be able to ask which study it came from.

It is a first round, but not a first attempt. The instrument behind it has been in development since June 2026, through earlier measurements, a self-audit that turned it on our own pages first, and seven revisions in which it was wrong and had to be corrected. That whole line of work, including this round and the two verification passes that followed it, comes to 9.6 billion tokens across 21,076 model calls, counted from the project's own logs and published as data/compute-log.json. We report it for the same reason we report everything else here: so the scale of what produced a number sits next to the number. The revisions, and what each one got wrong, are set out below.

The AI Search Evidence Index

Across 13 pages with enough claims to score, the median page linked a source for 33% of its numeric claims. Across all of them pooled, 80 of 252 claim-carrying blocks (31.7%) carried a link to the figure's source.

That is the headline number, and it is the least interesting thing we found.

Does the same publisher cite sources consistently?

Pages ranged from 0% to 62%. Two pages linked nothing. One linked most things.

Then we looked at publishers that appeared more than once, and the spread got stranger. tryprofound.com ranked for four of our six queries and entered with four pages:

tryprofound.com page Sourced
Top experts in generative engine optimization 6 of 11
What is answer engine optimization 5 of 10
Best AI visibility tools for marketing agencies 1 of 17
How ChatGPT sources the web 0 of 7 below threshold, not a rate

The last row is under the ten-block floor this study sets for a percentage, and it was not among the blocks read by hand. It is shown for completeness, not as a score. The finding stands on the three rateable pages: one publisher, 55% on one, 50% on another, 6% on a third.

semrush.com entered the same way (four pages across four queries), but only one of them clears the floor, so it cannot corroborate this. It is the only other publisher that could have, and it cannot. That is a limit on the finding, not a footnote to it: the within-publisher result rests on one company.

What it suggests is that sourcing is not a house style, not a policy a company sets and its writers follow. Something varies page to page inside one publisher. We cannot say what, because we measured pages and not the people or the deadlines behind them. But a reader who trusts a domain because one of its pages was well sourced has learned less about the next one than they think.

Which pages linked the most sources?

Every page below was retrieved inside the same 84 seconds, so no page had an advantage of time over another. The exact timestamps and a hash of each document are published with the data.

Page Claim blocks Sourced Share
industry-lens.com: AI search intelligence 21 13 62%
tryprofound.com: top experts in GEO 11 6 55%
tryprofound.com: what is answer engine optimization 10 5 50%
promptzero.tech: GEO guide 15 6 40%
writer.com: GEO, AEO and SEO 35 14 40%
frase.io: what is GEO 48 19 40%
ahrefs.com: how to rank in AI Overviews 15 5 33%
aisearch.similarweb.com: what is GEO 24 8 33%
semrush.com: best AI SEO tools 18 2 11%
seocrawl.ai: AI Overview ranking factors 17 1 6%
tryprofound.com: best AI visibility tools for agencies 17 1 6%
llmrefs.com: generative engine optimization 11 0 0%
zapier.com: best AI visibility tools 10 0 0%

Twelve further pages carried fewer than ten claim blocks and are not given a percentage; a figure built on four blocks is not a rate.

One page needs its context stated. industry-lens.com leads the table. Its format attaches a source by default, which is a different thing from a guide that has to go and find one. Its links split about evenly between its own /reports/ pages and outside sources: in the published sample, nine of its seventeen links are its own, so this is a difference of format, not a page citing only itself.

Two of this study's rules also happen to favour it at once. It is one of the three price-heavy pages, and a link to a pricing page is the one case that passes without the anchor test. A reader who wants to discount the top row has grounds to; we would rather hand them those grounds than have them found. The number is real. The comparison is not like-for-like, and we would rather say so than let the ranking imply otherwise.

Is sourcing a price the same as sourcing a statistic?

Not every number on these pages is the same kind of number.

Some pages are tool comparisons, and most of their figures are prices and feature counts, "$99/mo", "ten engines". For those, sourcing means linking the vendor's pricing page. Other pages are guides, and their figures are research statistics: "68% of searches end without a click". For those, sourcing means linking the study.

The two are not the same task. Checking a price takes one click to a page that either states it or does not. Checking a statistic means finding a study that may not be public.

In this round, three of the thirteen scored pages carry five or more price claims. They scored 62%, 11% and 0%. The other ten are guides, with a median of 36%.

Three pages is not enough to say anything about the genre, and we are not going to. We note the split because a reader comparing zapier.com with frase.io should know they are looking at two different jobs, and because anyone repeating this measurement should probably separate them from the start.

What counts as a source?

One rule decides almost everything: the link has to lead to the figure.

What makes a block sourced

A block scores as sourced when a link inside it points to where the number came from. The instrument never opens those links. It reads the saved copy of the page and judges each link from its address and its anchor text alone. So what is measured is whether a reader is pointed at a source, not whether that source is still reachable, and not whether it says what is claimed.

Table rows

One rule reaches outside the block. Where a short linked paragraph, under 400 characters, sits within two blocks of a run of table rows, its link credits every row in that run, on the reasoning that a source note placed above a table covers the table beneath it. It is applied identically to every page and it is visible in code/measure.py; it is also one of the two causes of the eight disagreements with the independent re-implementation reported below. We state it here because the sentence above says a link inside the block, and for table rows that is not the whole rule.

The anchor test

One test does most of the work: the anchor text has to contain a number that appears in the sentence making the claim, or share at least two content words with it. Both halves are cruder than they sound, and it is worth being exact about how.

The number test runs against the raw sentence, so a year counts as a number, an anchor reading "…(2026 guide)" can in principle pass on a sentence that mentions 2026. We checked: stripping every year from every sentence changes none of the blocks the instrument credited across the twenty-five measured pages. The weakness is real in the code and did not bite in this corpus.

The second half is looser. There is no rule against an anchor that is simply a name: a one-word brand cannot pass, because the test needs two shared words, but a two-word name can. ahrefs.com carried a figure attributed to "Ahrefs alumni Joshua Hardwick" with his LinkedIn profile linked, and the instrument counted it. So did the blind re-read. It was removed only when the block was opened and read by hand, and it is one of the ten rejections listed in the published materials. That is the case for reading the credited blocks rather than trusting the count, and it is why the table reports 80 rather than 90.

Where the rule favours us

That rule is not neutral, and it favours us. Writing anchor text that names the finding is good practice, and it is the style our own pages were rewritten into on 30 August, before this round ran. Several of the measured publishers anchor on a brand name instead, and a brand-name anchor passes only by accident, when the same name happens to appear in the sentence too. So the rule that decides most of this measurement rewards the house style of the people who wrote it. We are not in the table, which limits the damage, but a reader comparing two publishers should know that the one who anchors on figures is being scored on a convention as much as on whether the source is reachable.

What does not count

Several other things that look like sources are not. The examples below come from the earlier round in which these rules were written, not from the thirteen pages in the table; the rules were drawn while reading other people's pages and then applied unchanged here:

  • A sign-up or demo link. One page's three "sourced" claims were all a Start Free Trial button whose surrounding sentence happened to contain "14-day trial."
  • A related-post card. Read Post boxes under an article are navigation.
  • A vendor's homepage. It does not take you to the price you just quoted; its pricing page does.
  • The publisher's own product page. One page linked its own product four times for four different figures.
  • A pricing page, where the claim is itself about price; this one is an exception that passes without the anchor test, because quoting a price and linking the price list is a complete act.

Two further kinds (a help document, and an author's LinkedIn profile) carry none of the findings attributed to them, but the instrument has no rule for either. They were removed by hand, on one page (ahrefs.com), while its credited blocks were being read. The twelve sub-threshold pages were not read, so links of this kind may still be counted there.

Price claims

Price claims are counted separately and are not in the percentages. A block whose numbers are all prices goes into its own bucket, because sourcing a price means linking a vendor's pricing page, which is a different act from sourcing a research finding. Across the twenty-five measured pages that removes forty-six blocks, seventeen of them sourced. On the thirteen scored pages it removes eighteen blocks, four sourced. One page shows why this matters in both directions: orchly.ai is recorded at 0 of 6 on research claims while linking a vendor page for thirteen of its twenty-five price claims. The price counts ship in the data so a reader can put them back.

A source named but not linked ("according to Gartner", "Ahrefs' own research") is recorded separately and is not included in the percentages above. The blind reading measured the sourced/not-sourced call only, so this round produced no agreement figure for the second tier at all. Where it matters, we say a page names sources without linking them, and give a verified example.

Does naming a source without linking it count?

A page can name where a number came from without linking it: "according to Gartner", "Ahrefs' own research", "a Princeton-led study". The reader can find that source, but has to go looking.

We record this separately and do not fold it into the percentages, for a reason worth stating: we cannot show that we judge it reliably. The blind reading of 120 blocks tested the sourced/not-sourced call, not this one, so there is no agreement figure for it from this round. The only measurement we have is the twelve-item second-coder check from an earlier round, which matched on eleven, far too small to publish as a rate. So we report this tier as a flag with verified examples rather than as a number, and a reader should treat it as the weaker of the two.

Examples, each one published in the coding sheet with the block it came from, its links and its code:

  • llmrefs.com: "Vercel reports that 10% of new signups now come from ChatGPT referrals"
  • tryprofound.com-2: "backed by a 25-million-person consumer panel", of a named competitor
  • writer.com: "Only 16% of brands track their AI search performance (McKinsey)"

The last one shows the limit of measuring by block. That block carries no link, but four blocks earlier the same page makes the same claim and links the McKinsey study directly. Writer named its source and linked it; our unit of observation is the block, so the block that repeats the figure without the link is counted as one. We flag it here rather than let a reader find it in our own coding sheet and conclude we did not look.

One page makes the case for keeping this tier separate. In an earlier round of this measurement, ayzeo.com scored 0%: no claim on the page carried a link to its source. But the page ends with a nine-entry bibliography: BrightEdge, Reuters twice, Associated Press, Frase, SingleGrain and others, in full academic style. None of them are hyperlinks.

A reader can absolutely find "Reuters (2024), Reddit in AI content licensing deal with Google." By any reasonable standard that is a sourced claim. Our measure, which reads links block by block, cannot see it. Publishing "0%" for that page and stopping there would have been the single worst thing in this study.

So we ran the scan again on this round, before publishing these numbers. All 32 retrieved pages were checked for a source list in their final third, and every hit was read by hand (code/bibliography-scan.py). One page carried such a section, and it was not a bibliography: ahrefs.com ends with Further reading: two Search Engine Journal articles and a schema validator, none of them the source of a figure in its body. Its 33% stands.

The two pages published at 0% were read to their last block. llmrefs.com ends with a list of takeaways, zapier.com with related-post cards and a product call to action. Neither collects its sources anywhere. That matters more than the rest of this paragraph: a 0% next to a company's name is the hardest thing this study says, and it is not an artefact of where they put their links.

In the previous round it went the other way: three pages carried an end-of-page bibliography and one of them linked nothing from it. Rare, then, but real, which is why the check now runs every round rather than when we remember.

What does this mean for a reader checking a number?

Sourcing is not a house style. One publisher's rateable pages ran from 6% to 55%, a 49-point spread inside a single company, against a 62-point spread across the whole field. Whatever governs whether a number gets a link, it is not operating at the level of the organisation. We flag the weakness of this one ourselves: it rests on a single publisher, because no other one in the sample had more than one page clearing the threshold.

The practical consequence is for readers. Trusting a domain because one of its articles was well sourced tells you nothing about the next article on the same domain. The unit that earns trust is the page, and it has to earn it again every time.

The field's own research recommends the thing the field mostly does not do. The study that gave generative engine optimization its name found that citing sources was among the highest-performing content tactics for visibility in AI answers. The pages teaching that tactic link a source for about a third of their own figures.

We are not going to call that hypocrisy. A likelier explanation is ordinary: linking a source is work, it happens at the end of writing, and nothing on the page breaks when it is skipped. Nobody sees the missing link. That is exactly the kind of quiet defect a measurement is for.

What we did not test. Whether any of this affects citation by AI engines. It would be a reasonable hypothesis, since engines that reward verifiable content might prefer pages that link their sources, but we did not measure it and this study is not evidence for it. Testing it would take a different design: the same pages, tracked against actual citation over time.

That is the round we would run next.

Everything after this point is the apparatus: how the measurement was built, what it got wrong on the way, how reliable the coding is, and (in What this measure cannot tell you) the four things these numbers do not show. If you take one number from this article, take it with that section.

What this measure cannot tell you

It is not an accuracy check. A page can source every figure and still be wrong, or source nothing and be right. We measured reachability.

Our claim detector counts things that are not claims. The blind sample was coded for this too. Of the 85 sampled blocks the instrument would actually score, 19 (22.4%, 95% CI 13–31%) were advice, author biographies or examples of good writing rather than claims. Only 3 of those 19 were credited with a source, so they sit mostly in the unsourced side and the percentages above run low. Applying that rate to the 252 blocks in the table removes about 56 non-claims, 9 of them credited: the pooled figure corrects from 31.7% to roughly 36%. Every input to that arithmetic is in the published coding sheet, where every block carries an instrument_would_score flag: count the non-claims among the 85 it marks and you get the 19. Counting all 120 gives 26.7% instead, because the sample deliberately includes blocks the instrument never sees. We report the raw number as measured and the correction as an estimate.

Two known biases, pulling opposite ways. The correction above moves the figure up. The reliability check moves it down: across the 120 blind-coded blocks the instrument called 36.7% of them sourced where the blind read called 31.7%, so it credits about five points it should not. The non-claim correction is worth about four points in the other direction. The two are measured on different populations and we do not net them into a single number, but they are close in size and opposite in sign. That is why the figure we report is the one we measured, and why neither adjustment is offered on its own as a better estimate.

Six pages could not be retrieved. Most are large and established (Search Engine Land twice, Adobe, TechnologyAdvice), but not all: otterly.ai is a small vendor. A rough browser-side reading of five of them came out at 8%, 14%, 11%, 71% and 100%. Two of those sit far above the median and three below it, so we cannot say which way including them would have moved the number. They are out because they were never retrieved inside the single window, not because of what they would have done to the result.

One day, one snapshot. In an earlier round, one page lost 181 footnote links between two retrievals a day apart. That is why we publish the SHA-256 hash of every page we measured: you can re-fetch any URL, hash what you get, and see for yourself whether it still matches the document behind our number.

Six queries are not a field. Everything here is stated as "these pages, on this date."

How did our own pages score, and why are they excluded?

This is not a neutral comparison study, in the sense Boulesteix, Lauer and Eugster set out: we compete with the pages we measured, on the same queries. The instrument was built here, and while it was being built we could fix our own pages; the measured publishers could not. We ran ours first, linked our sources, fixed five claims that pointed at a page which no longer carried the figure, and re-measured. That is a real advantage, and it is why our four pages are reported here rather than in the table.

On the block unit the table uses, they score 91%, 75%, 74% and 33%. Naming them, because a score you cannot attach to a page is the same defect this study is about:

Placed in the table they would take three of the top five rows, and the fourth would sit around tenth of seventeen. We are not putting them there, because a number produced under a condition none of the other pages had is not a comparable number. Both units are published in notes/unit-of-analysis.md, where they appear under the codenames OURS-cited, OURS-geo, OURS-tools and OURS-track; a reader who disagrees with that call can merge them and see what happens.

The exclusion rules cost us nothing, since we do not use sign-up buttons or homepages as sources, and we found those rules while reading other people's pages, not our own. The per-page count of what each rule removed is published, so a reader can judge whether the rules were drawn fairly, though that count covers the earlier round in which the rules were written, not this round's thirteen pages.

What has been measured before

Citation practice has been studied for decades, but almost entirely inside academic publishing. A meta-analysis of twenty-eight studies of medical journals put the total quotation error rate, meaning citations that do not support the claim attached to them, at 25.4% (95% CI 19.5–32.4%). That work assumes a world where citations exist and asks whether they hold up.

Commercial web content has not been measured the same way, and the question there is one step earlier: is there a citation at all?

The nearest well-documented failure mode is circular reporting, and its internet variant citogenesis: an unsourced claim is repeated by a source that looks independent, and the repetition then becomes the citation. Wikipedia's account of circular reporting notes why the online form is hard to catch: modern pages are revised quickly, citations rarely carry an "as of" date, and pages rarely carry a reliable "last updated" one.

That observation shaped two decisions here. Every page in this study was retrieved inside a single 84-second window, and the dataset publishes a cryptographic hash of each document so a reader can tell whether the page they are looking at is the page we scored.

The field being measured has its own foundational study, the paper that named generative engine optimization and tested which content tactics raise visibility in AI answers. Notably, one of its top-performing tactics is citing sources. That is part of what makes the gap interesting: the field's own research says citing sources helps, and the field's own guides mostly do not link them.

Why did the first run give a different number?

We ran this measurement twice, and the first run produced a different answer.

The first frame used two seed queries and 19 pages. It gave a competitor median of 5%, and the story wrote itself: nobody in this field sources anything. We nearly published that.

Then we widened the frame to six queries and 38 pages. The median moved to 41%.

The first sample had not measured the field. It had caught a cluster of pages that happened not to link sources, and 19 pages was not enough for that to average out. Nothing was wrong with the instrument; the sample was too small to carry the claim we were about to make with it.

Widening the frame moved the number sharply up. Three refinements that followed each moved it back down:

Change Median
Two queries, counting every number as a claim 5%
Six queries 41%
Counting blocks rather than individual numbers 40%
After all 90 sourced blocks were read individually and 10 rejected 33%

The unit change mattered most for the leaders. Counting each number separately let one link take credit for every figure in its paragraph, which rewarded pages that write number-dense prose. On a block basis three pages lost between 10 and 19 points, while the highest-scoring page gained four.

The headline finding survived all four versions. The median moved from 5% to 41% to 40% to 33%, and the spread stayed wide throughout. The within-publisher comparison is younger than that: the first round took one page per site, so it could not be measured there at all. That is the part we would defend.

How this was measured

Every number above rests on a method, and it is set out in full below: the sampling frame, the single 84-second retrieval window, what the instrument counts as a sourced claim and what it throws out, how well it agrees with a careful reading, the seven revisions it went through and what each one got wrong, and every file in the dataset.

Sampling frame

Six seed queries on google.com (hl=en&gl=us), 30 August 2026. Every organic result on page one that carried an h3 heading. Publishing platforms (Reddit, Quora, LinkedIn, YouTube, Medium, Substack) and sponsored results excluded. Four of the six queries are ones our own pages target; two were carried over from an earlier round. So most measured pages compete with one of ours.

That produced 38 URLs across 30 domains. A domain enters as many times as Google ranked it; we did not hand-pick extra pages. The full frame, with every URL, is published with the data.

Retrieval

All 38 were fetched back to back inside 84 seconds. Thirty-two came back; six failed: four to HTTP 403 bot protection and two to a connection error. All six are listed by name and status in the published retrieval log. They are not a random six, and their exclusion is discussed below.

What was measured

Of the 32 pages retrieved, 25 carry a measurement row: the instrument skips a page carrying fewer than three blocks in total, counting claim blocks and price blocks together, and seven fell there. (segmetrics.io clears that cut on two claim blocks plus one price block, though it stays below the threshold for a percentage.) Of those 25, thirteen carry ten or more claim blocks and are given a percentage.

Unit

A block is one observation: a paragraph, list item, table row or quote of at least sixty characters containing at least one numeric claim. Six kinds of block are excluded before anything is counted: headings, date lines, the boilerplate templates a page hands its readers to copy, author biographies, "N out of M" phrasings, and a trailing date stamp; and every year is stripped from the text before the instrument looks for a number, so a block whose only figure is a year is not a claim. Counting headings was one of the errors the instrument's first revision fixed. An earlier version counted every number separately, which let a single link take credit for five figures and rewarded pages that write number-dense paragraphs. Switching to blocks moved twelve of the thirteen scored pages: it cost five of them, three by between 10 and 19 points and two by fewer than four, moved seven up by at most six, and left one unchanged. Across all twenty-five measured pages the movement is wider: one page gained thirteen points, another lost ten. Both versions ship with the data.

Coding

The instrument's verdicts were checked against a blind reading of 120 blocks drawn at random from the 470 across the 32 retrieved pages: 93.3% agreement, Cohen's κ = 0.85 (95% CI 0.75–0.95).

That pool is deliberately wider than the one the table is built on. It applies only two of the instrument's filters (sixty characters, and a numeric claim), so it also contains blocks the instrument would have thrown out. Thirty-five of the 120 are such blocks. This makes the test harder than the instrument's own job, not easier, and the number holds either way: on the 85 blocks the instrument would actually score, agreement is 94.1% and κ = 0.86 (95% CI 0.74–0.98). We report the wider figure as the headline because it is the more conservative of the two. On the conventional reading of κ (Landis & Koch, 1977) both sit in the "almost perfect" band, which is a label for a number, not a warrant that the instrument is right. Lombard, Snyder-Duch & Bracken (2002), the standard reference for reporting reliability in content analysis, asks for the coefficient together with enough about the coding instrument for a reader to judge it. We publish the instrument and every coded item as well, which is more than the checklist asks and less than a second coder would be. A separate twelve-block check by a second person, on an earlier round and on the named-source question only, matched the model on eleven; twelve items is too small to publish as a rate and it is not reported as one. That mismatch produced a codebook rule, disclosed rather than backdated.

Revisions

The instrument went through seven revisions before the version that produced this table, and the history of all seven is published with the data, along with an eighth entry, added after publication, recording a repair to the code's file handling that changed no measured number: what the error was, how it surfaced, and what changed. Three of those errors would have put a wrong number next to a named company.

The instrument counted a Start Free Trial button as a source, because the sentence beside it mentioned a "14-day trial." It counted Read Post cards under an article. It counted a page's own product page as sourcing four of that page's figures. It also failed to match according to when the A was capitalised, a case-sensitivity bug introduced at the second revision that survived two more before it was caught.

Each of those was found by reading pages, not by testing code.

Data, code and materials

Everything the figures rest on is published: the sampling frame, the retrieval log, both measurement files, every coding sheet, the instrument and its revision history. This follows the availability standard journals now apply to research articles (Nature Portfolio). The page snapshots themselves are held back for copyright reasons; their SHA-256 hashes are published so anyone can check a re-fetch against what we measured.

Use of AI tools

The measurement instrument, the retrieval, the coding and the first draft of this article were produced with Claude (Anthropic). The author designed the study, set its scope, reviewed the output at each stage, and takes full responsibility for the content, including the parts produced with AI assistance. Where a figure in this article depends on someone reading a page rather than a script parsing it, the reader was the model unless stated otherwise.

Cost

Between 27 August and 2 September 2026 this measurement (the retrieval, seven revisions of the instrument, the blind coding, the independent re-implementation, the hand-verification of all 90 sourced blocks, and the two verification rounds that followed and rewrote much of this article) took 3.03 billion tokens across 5,706 model calls. Across everything on this project since June, including the earlier self-audit that produced the instrument and every study before it, 9.6 billion tokens across 21,076 model calls since 10 June 2026.

Those are tokens processed, and 98.4% of that is cached context being read again rather than new work. Newly written output was 5.53 million. The daily breakdown is published as data/compute-log.json, counted from this project's session logs; the logs themselves are not published, but the arithmetic is, and the totals move only upward as work continues. We report this the way compute is reported in machine-learning papers: not as a claim to rigour, but so the scale sits on the record next to the result.

Verification

Every block the instrument credited with a source on the thirteen scored pages was opened and read individually: 90 of them. The thirteen it credited on the twelve sub-threshold pages were not read, because those pages carry no published percentage. Ten were overturned on inspection, leaving the 80 that stand in the table, and each rejection is listed in the published materials by page and by URL. Note which way that review can move a number: it can only take credits away, never add one. A block the instrument passed over was never re-examined. So the published percentages are the instrument's, minus what a reading could disprove, and any figure the instrument missed is still missing. The counter was then re-implemented from scratch by the same model, without reusing any of the original code, as a check, a test of whether the counting is robust to how it is written, not a second opinion from a second party. The re-implementation has no price rule, so the comparison has to be made on totals and sourced counts with price blocks left in; on that basis seventeen of twenty-five pages matched exactly, and every difference was traced to a known cause. Compared on totals alone the figure is nineteen of twenty-five, which is why the basis matters and is stated here.

How reliable is the coding?

What this check is not

The blocks were re-read by the same language model that wrote the instrument, blind to the instrument's verdict. Blindness is real: the verdicts were held in a separate file and opened only after coding, but authorship is not divided. This measures whether the code does what its own codebook says, not whether the codebook is right. The one human check below was coded by the author of this article. Neither is an independent reading in the sense that phrase usually carries, and no figure here should be read as one.

The check

Of the 470 claim-carrying blocks across the 32 pages we retrieved, 120 were drawn at random with a fixed seed, and coded without seeing what the instrument had decided about any of them. The coder saw the block's text, the links inside it and the publisher's domain. The instrument's verdict was withheld.

Blocks 120
Agreement 112 of 120, 93.3%
Cohen's κ 0.85
95% confidence interval 0.75 – 0.95

All 120 blocks are published with their text, their links and the code each was given, so a reader who disagrees with a judgement can find it and say so.

Which way it errs

The disagreements are not symmetrical. In seven of the eight, the instrument credited a source where the blind read found none; in one, it missed a source the read found. The instrument therefore errs towards making a page look better sourced than it is. An earlier version of this article, written on a sample of thirty, reported the opposite direction. That was wrong, and the larger sample corrects it.

The human check

This one is smaller, and it asks a different question. In an earlier round, thirty blocks were coded by the model and sealed, and the author of this article then coded twelve of them blind. Eleven of twelve matched. Two things limit what that shows. Twelve items is far too small to publish as a rate, and we do not. More importantly, every one of those thirty blocks had already been established to contain no link at all, so the person was judging only whether a source was named, never whether a link reached it. It is a check on the second tier, not on the percentages in this article. The single mismatch produced a codebook rule: a publisher naming itself is not naming a source unless it points to a specific identified work.

Claim detection

The selection itself was checked. Every block in the 120-block sample carries a field recording whether it is a numeric claim at all, so this is recomputable from the published sheet. Two rates come out of it, and they answer different questions:

not a claim applies to
whole 120-block sample 32, 27% (95% CI 19–35%) the wider pool, including blocks the instrument filters out
the 85 the instrument would score 19, 22.4% (95% CI 13–31%) the blocks behind the published percentages

The second is the one that matters for the table, and it is the one used in the correction below. The first is higher because the wider pool contains material the instrument was built to discard, which is what the sample was drawn to test. An earlier pass over 50 blocks gave a consistent 22%, but its per-item verdicts were not kept, so it stands as corroboration rather than as evidence a reader can check.

The two errors together

The two errors run in opposite directions. On the link side the instrument is generous, which pushes the published percentages up. On the claim side it counts things that are not claims, and those sit almost entirely on the unsourced side, which pushes the percentages down. We cannot net them off against each other and we do not try to. Both are stated so a reader can carry the uncertainty in the direction that matters to them.

Data and code

The whole package is archived at CERN's Zenodo under 10.5281/zenodo.22257563: the concept DOI, which resolves to the most recent version deposited, currently v1.0.1 (10.5281/zenodo.22300386). The version under which these figures were first published is 10.5281/zenodo.22257564. The archive is a third party's copy: if this site changed a figure tomorrow, the version behind the DOI would not move with it.

The files

The sampling frame, the measurement tables, the instrument, the independent re-implementation, the codebook and the coding sheets are published under CC BY 4.0:

File What it is
results.csv The final table, one row per published page, machine verdict and human verdict side by side
retrieval-log.csv All 38 URLs including the 6 that failed, with retrieval timestamps and status
snapshot-hashes.csv SHA-256 hash of every snapshot we scored
sampling-frame.json The six queries and every organic result they returned
measure.py The instrument
independent-counter.py The independent re-implementation used to check it
coder1-sealed-codes.json Coder 1's sealed codes, with the SHA-256 they were sealed under
blind-sample-120.json The 120 blocks coded blind against the instrument, each with its text, its links and the code it was given
blind-sample.py The script that drew the sample, with its seed
README Method, codebook and limitations in full

The same files are also published at github.com/minelgunesoglu-code/ai-search-evidence-index, where every change made after publication is visible in the commit history. If we correct a figure, you can see what it was before and why it changed. We would rather you did not have to take our word for that either.

Working notes

The working notes behind each decision (the design as written before measuring, the revision history, the hand-verification log) are in notes/. They are in Turkish; the README carries the method in English.

Copyright

We do not republish the pages themselves. Copying someone's article in full in order to criticise how they cite is not a thing we want to do, and the guidance on text and data mining draws the same line: analysing content is one question, redistributing it is another.

The hashes

The dataset carries something else instead. For every page we measured it records the URL, the retrieval timestamp, the character count, and the SHA-256 hash of the exact document we scored.

A hash is a fingerprint. Identical file, identical fingerprint; change one character and it changes completely. Fetch any of these URLs today, hash what comes back, and compare:

  • Same hash. You are looking at the document our number describes.
  • Different hash. The page has changed since 30 August 2026, and you now know that without having to take our word for anything.

Here are the thirteen pages in the table above, with the fingerprint of the document each number was taken from. The first sixteen characters are enough to tell a match from a mismatch; the full 64-character hash for all thirty-two retrieved pages is in snapshot-hashes.csv.

Page Characters First 16 characters of the SHA-256
industry-lens.com 171,142 a433d4582df8eaf7
tryprofound.com 233,257 dfdba698c69b3a2b
tryprofound.com-4 255,626 339461e4717e818e
promptzero.tech 60,958 72909dbd006fa72f
writer.com 273,523 418fc6b9cd234a02
frase.io 440,500 667829639b0101b0
ahrefs.com 990,735 4a24d9592ea6e5cb
aisearch.similarweb.com 657,785 f7ecc7bbe9e61ddf
semrush.com 252,925 18e99bb65b5555e3
seocrawl.ai 285,397 4016d2cf7ec98579
tryprofound.com-2 382,549 e0b92f1fe77055bb
llmrefs.com 167,049 c1ccb00c95e6b5e1
zapier.com 650,110 949857d72c1f6df0

Checking a hash yourself

To check one yourself, fetch the URL the same way the study did (same user agent, or several of these sites return a bot-block page instead) and hash what comes back:

UA="Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 \
(KHTML, like Gecko) Chrome/131.0 Safari/537.36"

curl -sL -A "$UA" "https://ahrefs.com/blog/how-to-rank-in-ai-overviews/" \
  | shasum -a 256

A caveat on that command: code/fetch.py wrote each page through Python's text layer, so the saved bytes are the re-encoded UTF-8 of the response, not the raw stream. On a page that has not changed the hashes match; if yours differs by a byte or two on an otherwise identical page, that is why. Comparing the two documents will tell you more than comparing the two hashes.

Excerpts

Short excerpts appear in the coding sheets where they are needed to show why a block was scored the way it was, capped at 200 characters in the coding sheets and 700 in the reliability sample. Measured against each page's extracted body text, the published excerpts are a median of 5.6% of it, but the share is uneven, and on two pages it is not small: writer.com 14% and industry-lens.com 39%, the latter because it is a short page from which many blocks were sampled. Any publisher who wants an excerpt removed can write and we will remove it.

Revision history

We also publish the instrument's revision history: seven revisions, each recording the error found, how it surfaced, and what changed. Three of those errors would have put a wrong number next to a named company.

If you find an error, tell us. We will correct it and say what changed.

References

How to cite this

The dataset is archived at CERN's Zenodo, which pins a copy no one can quietly edit, including us.

Gunesoglu, M. (2026). The AI Search Evidence Index: source determinability of numeric claims on AI-visibility web pages (Round one) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.22257563

The DOI above resolves to the most recent version deposited, currently v1.0.1. To cite the version these numbers were first published under, use 10.5281/zenodo.22257564.