The AI Search Evidence IndexOpen dataset

They publish the advice, not the evidence behind it

38
Pages measured
30
Domains
3.7M
Tokens written

Round 1 · Dataset frozen 30 August 2026, 20:38:51–20:40:17 UTC

A measurement of whether a reader can reach the source of a numeric claim. Open data, open method, seven documented revisions of the instrument, and every published figure read by a human.

Do AI Visibility Guides Link Their Sources? We Measured 38 Pages

TL;DR: We measured 38 pages that rank for AI-visibility queries, all retrieved inside one 86-second window, and counted how often a reader can reach the source of a numeric claim. The median page linked a source for 33% of them. The range ran from 0% to 62%, and one publisher spanned nearly that whole range across its own four pages.

The pages that teach you how to get cited by AI are full of numbers. Citation rates, engine counts, correlation coefficients, market shares, prices. We wanted to know a narrow thing about those numbers: can a reader reach the source?

Not whether the figures are right. Whether they are checkable.

We call this measurement the AI Search Evidence Index. This is its first round.

The AI Search Evidence Index

What has been measured before

Citation practice has been studied for decades, but almost entirely inside academic publishing. Reviews of the biomedical literature put citation inaccuracy, meaning quotations that do not support the claim attached to them, at roughly 20–26%. That work assumes a world where citations exist and asks whether they hold up.

Commercial web content has not been measured the same way, and the question there is one step earlier: is there a citation at all?

The nearest well-documented failure mode is circular reporting, and its internet variant citogenesis: an unsourced claim is repeated by a source that looks independent, and the repetition then becomes the citation. Wikipedia's own record of these incidents notes why they are hard to catch: modern pages are revised quickly, citations rarely carry an "as of" date, and pages rarely carry a reliable "last updated" one.

That observation shaped two decisions here. Every page in this study was retrieved inside a single 86-second window, and the dataset publishes a cryptographic hash of each document so a reader can tell whether the page they are looking at is the page we scored.

The field being measured has its own foundational study, the Princeton, Georgia Tech and IIT Delhi paper that named generative engine optimization and tested which content tactics raise visibility in AI answers. Notably, one of its top-performing tactics is citing sources. That is part of what makes the gap interesting: the field's own research says citing sources helps, and the field's own guides mostly do not link them.

Across 13 pages with enough claims to score, the median page linked a source for 33% of its numeric claims. Across all of them pooled, 80 of 252 claim-carrying blocks (31.7%) carried a working link to the figure's source.

That is the headline number, and it is the least interesting thing we found.

Does the same publisher cite sources consistently?

Pages ranged from 0% to 62%. Two pages linked nothing. One linked most things.

Then we looked at publishers that appeared more than once, and the spread got stranger. tryprofound.com ranked for four of our six queries and entered with four pages:

tryprofound.com page Sourced
Top experts in generative engine optimization 6 of 11
Best AI visibility tools for marketing agencies 5 of 10
How ChatGPT sources the web 1 of 17
What is answer engine optimization 0 of 11

One publisher, four pages, a range as wide as the entire field's.

Sourcing is not a house style here. It is not a policy that a company sets and its writers follow. It is decided page by page: by who wrote it, on what day, under what deadline. A reader who trusts a domain because one of its pages was well sourced has learned nothing about the next one.

Which pages linked the most sources?

Every page below was retrieved on 30 August 2026 between 20:38:51 and 20:40:17 UTC, a single 86-second window, so no page had an advantage of time over another.

Page Claim blocks Sourced Share
industry-lens.com: AI search intelligence 21 13 62%
tryprofound.com: top experts in GEO 11 6 55%
tryprofound.com: best AI visibility tools 10 5 50%
promptzero.tech: GEO guide 15 6 40%
writer.com: GEO, AEO and SEO 35 14 40%
frase.io: what is GEO 48 19 40%
ahrefs.com: how to rank in AI Overviews 15 5 33%
aisearch.similarweb.com: what is GEO 24 8 33%
semrush.com: best AI SEO tools 18 2 11%
seocrawl.ai: AI Overview ranking factors 17 1 6%
tryprofound.com: how ChatGPT sources the web 17 1 6%
llmrefs.com: generative engine optimization 11 0 0%
zapier.com: best AI visibility tools 10 0 0%

Thirteen further pages carried fewer than ten claim blocks and are not given a percentage; a figure built on four blocks is not a rate.

One page needs its context stated. industry-lens.com leads the table, and most of its links point to its own /reports/ pages. It is a news aggregator: its format attaches a source by default, which is a different thing from a guide that has to go and find one. The number is real. The comparison is not like-for-like, and we would rather say so than let the ranking imply otherwise.

Two kinds of page, two kinds of claim

Not every number on these pages is the same kind of number.

Some pages are tool comparisons, and most of their figures are prices and feature counts, "$99/mo", "ten engines". For those, sourcing means linking the vendor's pricing page. Other pages are guides, and their figures are research statistics: "68% of searches end without a click". For those, sourcing means linking the study.

The two are not the same task. Checking a price takes one click to a page that either states it or does not. Checking a statistic means finding a study that may not be public.

In this round, three of the thirteen scored pages carry five or more price claims. They scored 62%, 11% and 0%. The other ten are guides, with a median of 36%.

Three pages is not enough to say anything about the genre, and we are not going to. We note the split because a reader comparing zapier.com with frase.io should know they are looking at two different jobs, and because anyone repeating this measurement should probably separate them from the start.

What counts as a source?

One rule decides almost everything: the link has to lead to the figure.

A block scores as sourced when a working link inside it takes the reader to where the number came from. Several things that look like sources are not:

  • A sign-up or demo link. One page's three "sourced" claims were all a Start Free Trial button whose surrounding sentence happened to contain "14-day trial."
  • A related-post card. Read Post boxes under an article are navigation.
  • A vendor's homepage. It does not take you to the price you just quoted; its pricing page does.
  • The publisher's own product page. One page linked its own product four times for four different figures.
  • A help document or an author's LinkedIn profile. Neither carries the finding attributed to it.

A source named but not linked ("according to Gartner", "Ahrefs' own research") is recorded separately and is not included in the percentages above. Our instrument agrees with a human coder on that judgment at κ = 0.66, which is not good enough to publish as a rate. Where it matters, we say a page names sources without linking them, and give a verified example.

Sources that are named but never linked

A page can name where a number came from without linking it: "according to Gartner", "Ahrefs' own research", "a Princeton-led study". The reader can find that source, but has to go looking.

We record this separately and do not fold it into the percentages, for a reason worth stating. Our instrument agrees with a human coder on this judgment at Cohen's κ = 0.66, respectable but below the level at which a rate should be published. So we report it as a flag with verified examples rather than a number.

Verified examples from the set:

  • llmrefs.com: "Vercel reports that 10% of new signups now come from ChatGPT referrals"
  • aisearch.similarweb.com: "…30–40% higher AI visibility, according to Princeton's research"
  • seocrawl.ai: "Google reports AI Overviews are driving over a 10% increase in Search usage"
  • writer.com: "the brands McKinsey found are tracking AI search"

One page makes the case for keeping this tier separate. In an earlier round of this measurement, ayzeo.com scored 0%: no claim on the page carried a link to its source. But the page ends with a nine-entry bibliography: BrightEdge, Reuters twice, Associated Press, Frase, SingleGrain and others, in full academic style. None of them are hyperlinks.

A reader can absolutely find "Reuters (2024), Reddit in AI content licensing deal with Google." By any reasonable standard that is a sourced claim. Our measure, which reads links block by block, cannot see it. Publishing "0%" for that page and stopping there would have been the single worst thing in this study.

We checked whether this was common. Across every page in the current round, not one carried an end-of-page bibliography. ayzeo.com is the exception, not the rule, but the exception is why the tier exists.

What we got wrong the first time

We ran this measurement twice, and the first run produced a different answer.

The first frame used two seed queries and 19 pages. It gave a competitor median of 5%, and the story wrote itself: nobody in this field sources anything. We nearly published that.

Then we widened the frame to six queries and 38 pages. The median moved to 33%.

The first sample had not measured the field. It had caught a cluster of pages that happened not to link sources, and 19 pages was not enough for that to average out. Nothing was wrong with the instrument; the sample was too small to carry the claim we were about to make with it.

Three further changes moved the number again, and it is worth listing them because each one moved it down:

Change Median
Two queries, counting every number as a claim 5%
Six queries 41%
Counting blocks rather than individual numbers 40%
After a human read all 89 sourced blocks and rejected 11 33%

The unit change mattered most for the leaders. Counting each number separately let one link take credit for every figure in its paragraph, which rewarded pages that write number-dense prose. On a block basis the three highest-scoring pages each lost between 10 and 19 points.

The finding survived all four versions. The median moved from 5% to 41% to 40% to 33%; the spread stayed wide, and the within-publisher variance stayed as wide as the field. That is the part we would defend.

How we measured this

Sampling frame. Six seed queries on google.com (hl=en&gl=us), 30 August 2026. Every organic result on page one. Publishing platforms (Reddit, Quora, LinkedIn, YouTube, Medium, Substack) and sponsored results excluded. The queries were the ones our own pages target, so every measured page competes with one of ours.

That produced 38 URLs across 30 domains. A domain enters as many times as Google ranked it; we did not hand-pick extra pages. The full frame, with every URL, is published with the data.

Retrieval. All 38 fetched back to back inside 86 seconds. Thirty-two came back. Six returned bot-protection errors and are listed by name with their status; they are not random, and their exclusion is discussed below.

Unit. A block is one observation: a paragraph, list item, table row or quote containing at least one numeric claim. An earlier version counted every number separately, which let a single link take credit for five figures and rewarded pages that write number-dense paragraphs. Switching to blocks cost the three highest-scoring pages between 10 and 19 points. Both versions ship with the data.

Coding. Two independent coders. The first coder's judgments were sealed with a SHA-256 hash before the second coded, and items whose codes had been discussed were excluded from the comparison. Agreement on the remaining twelve: 92%, Cohen's κ = 0.80. The single disagreement was resolved by discussion and produced a new codebook rule; κ is reported from the codes as first recorded and was not recomputed afterwards.

The instrument was revised seven times, and the revision history is published with the data: what the error was, how it surfaced, and what changed. Three of those errors would have put a wrong number next to a named company.

The instrument counted a Start Free Trial button as a source, because the sentence beside it mentioned a "14-day trial." It counted Read Post cards under an article. It counted a page's own product page as sourcing four of that page's figures. It also, for five versions, failed to see According to when the A was capitalised, a case-sensitivity bug in a regular expression that had been sitting there since the first version.

Each of those was found by reading pages, not by testing code.

Cost. Between 27 and 31 August 2026 this measurement (the retrieval, seven revisions of the instrument, the blind coding, the independent re-implementation and the hand-verification of all 89 sourced blocks) took 3.7 million written tokens across 3,605 model calls. Across the whole of August, including the earlier self-audit that produced the instrument, 11.6 million.

Total tokens processed was 1.9 billion, but 98% of that is cached context being re-read rather than new work, so we quote the written figure. We report this the way compute is reported in machine-learning papers: not as a claim to rigour, but so the scale sits on the record next to the result.

Verification. Every sourced block in the table above was opened and read by a human: 89 of them. Eleven machine verdicts were overturned on inspection, and each rejection is listed in the published materials. The counter was then re-implemented from scratch, without reusing any of the original code, as a check: seventeen of twenty-five pages matched exactly, and every difference was traced to a known cause.

How reliable is the coding?

Two questions matter here, and they are different: does the instrument agree with a person, and do two people agree with each other?

Instrument versus human. Thirty blocks were coded by hand and compared against the instrument: 83% agreement, Cohen's κ = 0.66. The disagreements were not random. Every one of them ran the same direction: the instrument missed sources that a human saw, and never invented one that was not there. An instrument that errs one way makes no page look better than it is.

Human versus human. The first coder's judgments were sealed with a SHA-256 hash before the second coder saw the material, and items whose codes had already been discussed were excluded from the comparison. On the remaining twelve: 92% agreement, Cohen's κ = 0.80.

The single disagreement is instructive. One page wrote "GEO went from a fringe idea to a 22,000-search-a-month category" and named itself, but no study, later in the same block. One coder read the self-naming as a source; the other did not. The codebook had no rule for it.

We wrote one: a publisher naming itself is not naming a source unless it points to a specific identified work. We report κ from the codes as first recorded, not from the corrected ones. Agreeing after a conversation is not reliability.

Claim detection. The judgments above are about blocks the instrument had already selected. We also checked the selection itself: 50 blocks were hand-coded for whether they contain a numeric claim at all, and 11 did not: advice, author biographies, examples of good writing. That is 22% (95% CI 11–33%), and because those blocks sit almost entirely on the unsourced side, the percentages in this study run low rather than high.

What this measure cannot tell you

It is not an accuracy check. A page can source every figure and still be wrong, or source nothing and be right. We measured reachability.

Our claim detector counts things that are not claims. We hand-coded 50 blocks to find out how often: 22% were advice, author biographies or examples of good writing rather than claims (95% CI 11–33%). Because these sit almost entirely in unsourced blocks, the percentages above run low. Corrected, the pooled figure is closer to 39% than 31.7%. We report the raw number as measured and the correction as an estimate.

Six pages could not be retrieved, and they skew large and established. Two of them carried enough claims to score; a rough browser-side reading put them at 8% and 14%, which would move the median down rather than up.

One day, one snapshot. In an earlier round, one page lost 181 footnote links between two retrievals a day apart. That is why we publish the SHA-256 hash of every page we measured: you can re-fetch any URL, hash what you get, and see for yourself whether it still matches the document behind our number.

Six queries are not a field. Everything here is stated as "these pages, on this date."

What this means

Sourcing is not a house style. That is the finding we would defend hardest, because it is the one the data supports most cleanly. The gap between a publisher's own best and worst page was as wide as the gap across the whole field. Whatever governs whether a number gets a link, it is not operating at the level of the organisation.

The practical consequence is for readers. Trusting a domain because one of its articles was well sourced tells you nothing about the next article on the same domain. The unit that earns trust is the page, and it has to earn it again every time.

The field's own research recommends the thing the field mostly does not do. The Princeton study that gave generative engine optimization its name found that citing sources was among the highest-performing content tactics for visibility in AI answers. The pages teaching that tactic link a source for about a third of their own figures.

We are not going to call that hypocrisy. A likelier explanation is ordinary: linking a source is work, it happens at the end of writing, and nothing on the page breaks when it is skipped. Nobody sees the missing link. That is exactly the kind of quiet defect a measurement is for.

What we did not test. Whether any of this affects citation by AI engines. It would be a reasonable hypothesis, since engines that reward verifiable content might prefer pages that link their sources, but we did not measure it and this study is not evidence for it. Testing it would take a different design: the same pages, tracked against actual citation over time.

That is the round we would run next.

Why our own pages are not in the table

We ran our own pages through the same instrument before anyone else's. They scored 52%, 42%, 52% and 17%. We linked our sources, fixed four claims that pointed at a page which no longer carried the figure, and re-measured.

We are not in the table because we had the instrument in hand while fixing our pages, and the measured pages did not. Any number we published for ourselves would not be comparable to theirs.

The exclusion rules cost us nothing, since we do not use sign-up buttons or homepages as sources, and we found those rules while reading other people's pages, not our own. Both facts are in the published materials, with a per-page count of what each rule removed, so a reader can judge whether the rules were drawn fairly.

Data and code

The sampling frame, the measurement tables, the instrument, the independent re-implementation, the codebook and the coding sheets are published under CC BY 4.0:

File What it is
results.csv The final table, one row per published page, machine verdict and human verdict side by side
retrieval-log.csv All 38 URLs including the 6 that failed, with retrieval timestamps and status
snapshot-hashes.csv SHA-256 hash of every snapshot we scored
sampling-frame.json The six queries and every organic result they returned
measure.py The instrument
independent-counter.py The independent re-implementation used to check it
coder1-sealed-codes.json Coder 1's sealed codes, with the SHA-256 they were sealed under
intercoder-kappa.json The inter-coder comparison behind κ = 0.80
README Method, codebook and limitations in full

The same files are also published at github.com/minelgunesoglu-code/ai-search-evidence-index, where every change made after publication is visible in the commit history. If we correct a figure, you can see what it was before and why it changed. We would rather you did not have to take our word for that either.

The working notes behind each decision (the design as written before measuring, the revision history, the hand-verification log) are in notes/. They are in Turkish; the README carries the method in English.

We do not republish the pages themselves. Copying someone's article in full in order to criticise how they cite is not a thing we want to do, and the guidance on text and data mining draws the same line: analysing content is one question, redistributing it is another.

So the dataset carries something else instead. For every page we measured it records the URL, the retrieval timestamp, the byte size, and the SHA-256 hash of the exact document we scored.

A hash is a fingerprint. Identical file, identical fingerprint; change one character and it changes completely. Fetch any of these URLs today, hash what comes back, and compare:

  • Same hash. You are looking at the document our number describes.
  • Different hash. The page has changed since 30 August 2026, and you now know that without having to take our word for anything.

Here are the thirteen pages in the table above, with the fingerprint of the document each number was taken from. The first sixteen characters are enough to tell a match from a mismatch; the full 64-character hash for all thirty-two retrieved pages is in snapshot-hashes.csv.

Page Bytes First 16 characters of the SHA-256
industry-lens.com 171,142 a433d4582df8eaf7
tryprofound.com 233,257 dfdba698c69b3a2b
tryprofound.com-4 255,626 339461e4717e818e
promptzero.tech 60,958 72909dbd006fa72f
writer.com 273,523 418fc6b9cd234a02
frase.io 440,500 667829639b0101b0
ahrefs.com 990,735 4a24d9592ea6e5cb
aisearch.similarweb.com 657,785 f7ecc7bbe9e61ddf
semrush.com 252,925 18e99bb65b5555e3
seocrawl.ai 285,397 4016d2cf7ec98579
tryprofound.com-2 382,549 e0b92f1fe77055bb
llmrefs.com 167,049 c1ccb00c95e6b5e1
zapier.com 650,110 949857d72c1f6df0

To check one yourself, fetch the URL and hash the bytes:

curl -sL "https://ahrefs.com/blog/how-to-rank-in-ai-overviews/" | shasum -a 256

Short excerpts appear in the coding sheets where they are needed to show why a block was scored the way it was, capped at roughly 200 characters. For most pages that is 1–3% of the text.

We also publish the instrument's revision history: seven revisions, each recording the error found, how it surfaced, and what changed. Three of those errors would have put a wrong number next to a named company.

If you find an error, tell us. We will correct it and say what changed.

References