Home › Research

Research

The AI Search Evidence IndexResearchOpen dataset

Two numbers in three have no link to a source

34 min
Read time
28
Data files published
9.6B
Tokens spent

Round 1 · Dataset frozen 30 August 2026

A measurement of whether a reader can reach the source of a numeric claim. Open data, open method, the instrument's full revision history, and every credited block on the scored pages opened and read one at a time.

Read the study

Most statistics about AI search are summaries: a vendor runs some prompts and publishes a percentage, and nobody outside the company can check it. We do it the other way round. Every study below is hand-run and its per-run data is published in full, so you can open the file and tell us we read it wrong.

The studies, and where their data is

Four hand-run studies, together 129 logged engine runs and 38 pages measured. The newest ran 150 scripted requests through the Claude API. A run is one question put to one engine once.

Claude Retrieval and Citation Study

When the query names your brand, its own site is usually not what the search returns

The brand's own site was among the retrieved domains in 36 of 108 run-brand pairs, 33.3%. Twelve of twenty-five named brands were never retrieved in any run.

Minel Gunesoglu · September 2026

Read the study →

Cross-Engine Citation Study

Four engines agree on the brands, not on the sources

The four engines agreed on 32% of the brands they recommended, and 1% of the sources they cited.

Minel Gunesoglu · June to August 2026

Read the study →

AI Overview Cited-Domains Study

Reddit is the source Google's AI Overview reaches for first

Reddit was cited in 13 of the 21 AI Overviews, more than twice the next site.

Minel Gunesoglu · June 2026

Read the study →

AI Overview Industry Study

Every industry is answered from a different set of sources

Review and authority media appeared in 24 of 27 AI Overviews, but which outlets changed completely by industry.

Minel Gunesoglu · June 2026

Read the study →

AI Search Evidence Index

Two numbers in three have no link to a source

Two numbers in three on AI-visibility pages have no link to a source a reader can reach.

Minel Gunesoglu · August 2026

Read the study →

Generative Engine Optimization StatisticsNot our measurement

The most-recycled GEO stats do not say what they are quoted as saying

The most-recycled figures do not say what they are quoted as saying: a 525% revenue jump is one publisher's own number, and 71.5% describes surveyed marketers, not consumers.

Minel Gunesoglu · Kept current

Read the study →
Collection conditions, limits and raw data

Claude Retrieval and Citation Study

Collected 7 September 2026 · Raw data: Repository: data, code and protocol

Fifty B2B-SaaS buyer questions, three runs each, 150 requests against claude-sonnet-5 with the web search tool called directly, so the search results come back alongside the answer and the retrieved list can be recorded separately from the cited one. Twenty-three of the fifty queries name a brand in the query text, which is where the headline figure comes from: which brand is at issue needs no interpretation.

Limit. The surface measured is the Claude API's web search tool, not claude.ai. Whether the two use the same search provider is not documented by Anthropic. In the first pass, twenty runs' searches failed silently and were re-run about twenty minutes later; that defect is published with the data rather than quietly corrected.

Cross-Engine Citation Study

Collected Round 1: 19 Jun 2026 · Round 2: Aug 2026 · Raw data: Round 2 CSV · JSON · Round 1 CSV · JSON

Round two (August 2026). All 48 runs: 189 recommended brands and 322 cited sources, one row per engine-question pair. Brands are listed in the order the engine named them; every source carries a type code (vendor site, review media, independent listicle, Reddit, forum, YouTube, reference). The JSON carries the same rows plus the run conditions, the source-type key and the stated limits, so the file stands on its own if it travels without this page.

Round one (19 June 2026). Smaller and less even than August: Perplexity was run on all twelve questions, the other three engines on fewer. Several Perplexity source panels were truncated behind a “show more” control we could not fully expand, so the recorded list is what was visible and written down at the time. The churn figures quoted across this site (80 sources in June, 42 still cited in August, 48% turnover) are computed from this file and the August one.

AI Overview Cited-Domains Study

Collected 10 Jun 2026 · Raw data: CSV · JSON

24 buyer searches, collected by hand on 10 June 2026. Twenty-one of the twenty-four produced an AI Overview; the other three returned a Shopping grid. Each row carries the domains the source panel listed, which is how the most-cited table on that page is counted. The rows themselves were not downloadable until now.

Limit. Site type was assigned by hand from each source's title and brand, not by reading the page. That classification is our judgement, and the rows are published so you can disagree with any single one of them.

AI Overview Industry Study

Collected 11-12 Jun 2026 · Raw data: CSV · JSON

30 searches across six industries, collected by hand on 11–12 June 2026. Twenty-seven produced an AI Overview. Each row carries the source count, whether Reddit appeared, and the site types recorded for that answer.

Limit. Site type was assigned by hand from each source's title and brand, not by reading the page. That classification is our judgement, and the rows are published so you can disagree with any single one of them.

AI Search Evidence Index

Collected Round 1: 30 Aug 2026 · Raw data: Results CSV · Retrieval log · Hashes · Frame JSON

Round one, dataset frozen 30 August 2026. 38 pages that rank for AI-visibility queries were sampled, 32 of them retrieved inside one 84-second window, and the 13 that carried enough numeric claims were scored. Every credited block on the scored pages was opened and read one at a time.

The dataset is deposited at Zenodo with a DOI, and the instrument's full revision history is published with it, so a later round can be compared against this one.

Generative Engine Optimization Statistics

Collected Kept current; every row carries its source and year

This page is not one of our measurements. It reads other people's published figures back against their sources and records, for each one, what the source actually says, the year, and the sample it rests on. It is listed here because it belongs to the same discipline as the studies, not because the numbers in it are ours.

Everything here is CC BY 4.0. Four studies sit behind this page, so cite the one your number came from rather than the page as a whole; each study page carries its own attribution and its own conditions. The shared method notes are on the methodology page, and journalists or researchers who want the working notes can get in touch.