How We Measured It: The AI Answer Evidence Index Method
TL;DR: This is the method behind The AI Answer Evidence Index. It covers the 50 questions and how they were drawn, how each of the five AI engines was asked, the fetch of every cited page, what counts as a found figure, the two model readings that test the check, the measures, the repeat run, the preregistration, the analysis of sources, and the scope of claims.
The findings are in the study itself. This page is the apparatus: how the questions were chosen, how the answers and their cited pages were saved, what the check does and does not count, how its distance from a judgement of support was measured, and what the study claims.
The method was written before the data. It was published on Zenodo on 5 October 2026, before the first answer was collected: doi:10.5281/zenodo.23145556. That preregistration named Google AI Overviews, ChatGPT and Gemini; Perplexity and Claude (API) were added on 6 October, and their answers went through the same four steps on 7 and 8 October, after the results of the first three had been computed, as an exploratory addition (deviation log, entries 3, 5, 6 and 7).
How We Measured It: The AI Answer Evidence Index Method
- What question does the study ask?
- How we measured this
- How was each engine asked?
- Questions and collection
- Cited pages, figures and the check
- What can happen to a figure?
- How are the figures that were not found read?
- How far is the check from a judgement of support?
- What is reported?
- How stable is one answer?
- What was registered, and what was recorded afterwards?
- The second analysis: which sources the engines cite
- Scope of claims
- Use of AI models
- Data and code
- References
What question does the study ask?
When an AI answer engine states a figure and shows its sources, can a reader find that figure on the pages shown?
The unit is one figure in one answer, together with the pages that answer cites. The study is a descriptive measurement. It tests no hypothesis, and every result is published, whatever it shows.
It is round two of the Evidence Index series. Round one is The AI Search Evidence Index (doi:10.5281/zenodo.22257563).
How we measured this
| Item | Detail |
|---|---|
| Questions | 50 buyer questions about prices, costs, fees and value, from six sectors |
| Engines | Google AI Overviews, ChatGPT, Gemini, Perplexity, Claude (API) |
| Askings | 250: each engine is asked each question once |
| Locale | English questions; New York (US) as the location |
| Window | 5 to 7 October 2026 (UTC+3); every answer carries its timestamp |
| Steps | Four, the same for every engine: extraction of the figures, check against the cited pages, classifier for the figures not found, summary |
How was each engine asked?
| Engine | How it was asked | Dates (UTC+3) | Answers, of 50 questions |
|---|---|---|---|
| Google AI Overviews | Google Search, logged out, in a Chrome browser driven by the collection script, from a New York network exit | 5 to 7 October 2026 | 49 |
| ChatGPT | The public web interface as a guest in a temporary chat, logged out, by the collection script, from the same exit | 5 and 6 October 2026 | 50 |
| Gemini | The public web interface, logged out, by the collection script, from the same exit | 5 and 6 October 2026 | 50 |
| Perplexity | The web interface through a signed-in free account, from the same exit. The study's AI assistant opened each question, one at a time, in the same asking order, and saved the answer text and the source list as the page showed them | 6 October 2026 | 50 |
| Claude (API) | Model claude-sonnet-5-5 through the Anthropic API with the API's web search tool, New York, NY, United States as its location setting, no system prompt and the API's default settings. At most 3 searches were allowed per answer; one was used in 48 answers and none in 2 |
6 October 2026 | 50 |
Facts that belong to one route:
- Perplexity. The operator signed in to the account by hand, and it held five earlier searches. The interface language was Turkish, and the answers are in English. 5 of the 50 records carry the interface's notice that a preview of the advanced search was switched on, and for 10 further records the notice could not be observed. Perplexity was planned as an engine of the scripted collection; in the trial of 4 October its verification page did not admit the script-driven browser window, and the scripts do not work around a verification, so the preregistration left Perplexity out. One answer (SVC-07) was read from the account's stored thread after an interruption, so Perplexity's result is given with and without it.
- Claude (API). The answers are what the model returns to a bare question with a search tool. They are not a reading of the consumer interface, and Claude is named with its access route wherever a result of its is reported. The preregistration left Claude out because the API is a different access route from the public interfaces. Two answers ran no search and cite no page; their figures get the outcome "no source".
- AI models in the collection. The collection of Google AI Overviews, ChatGPT and Gemini is scripted, and no AI model takes part in it. An AI assistant collected Perplexity's answers, and Claude's came through the API.
Questions and collection
Questions. The frame holds 100 questions taken verbatim from Google autocomplete (hl=en,
gl=us) on 22 September 2026, by a written selection rule. A script draws 50 of them. The 10
pilot questions are excluded, because the instrument was adjusted while their answers were
visible. Quotas follow the frame's sector order: 9 each for business software and AI tools, and 8
each for personal finance, legal and local services, consumer electronics and travel.
The seed of the draw is taken from the SHA-256 of the frame, which was fixed on 22 September, so no number was chosen for the draw. The 50 questions are then shuffled once with the same generator, and every engine is asked in that order. Each question is sent exactly as autocomplete wrote it.
Scripted collection (Google AI Overviews, ChatGPT, Gemini). A fixed script per engine drives a Chrome browser. Questions go one at a time, 60 to 120 seconds apart, in blocks of 8 fixed by the asking order. Each block goes to all three engines before the next block starts. The network exit is checked before each engine and at the end of a block.
The script does not log in, accept terms, solve a challenge or work around a limit. At a verification page it clicks nothing and waits up to five minutes for the operator, who may pass the page by hand. An attempt is complete when the interface shows the answer as finished. An incomplete attempt is saved and the question is asked again in a later block, up to three attempts. No complete answer is replaced. Perplexity and Claude (API) were asked each question once, and every answer was complete.
What is saved. For every scripted attempt: the page HTML, the visible text, a screenshot, the time, the address, the model label where the interface shows one, and the interface language. A parser frozen with the scripts reads the answer text, the source list and the sources attached to each block from the saved HTML. Sponsored and shopping cards are not part of the answer or its sources. For Perplexity the record holds the answer text and the source list as the page showed them; its citation chips stand in the saved text as lines of their own (a site name, and a counter such as +4), and these 534 lines are left out before the extraction. For Claude (API) the record holds the text and the citations the API returned.
What a cited page is. For Google AI Overviews, ChatGPT and Gemini: a page the answer's source list or a citation marker names. For Perplexity: any entry of the source list the interface shows, whether or not the answer text carries a citation chip for it (608 entries, a median of 10 per answer). For Claude (API): a page that a citation of the response names (272 entries).
Google's links. Google AI Overviews writes its links as opaque tokens. Where the saved page holds no address for a cited page, the script opens Google's own link once, as a reader's click would, and records the first address outside Google. A cited page whose address stays unknown counts as cited and unreadable.
Perplexity's addresses. Its records hold host and path without the query string. 22 of its 608 source entries had one, and their pages were fetched without it, so the page read for them may differ from the page cited.
Cited pages, figures and the check
Cited pages. Every cited address is fetched from the New York exit: once with a plain HTTP
request and once rendered in headless Chrome. PDF text is extracted. The fetcher identifies itself
as an automated research fetcher in its user agent (AIAnswerEvidenceIndex/1.0). A page that
blocks it or shows a challenge is recorded as an access failure.
A page is readable when either fetch returns its text. A fetched text under 500 characters, or under 2,000 characters with a phrase from a fixed list of check-page phrases, is not the page's text.
When the pages were fetched. For Google AI Overviews, ChatGPT and Gemini the addresses of a block are fetched after that block; their answers cite 483 addresses, a fetch returned text for 429 of them, the check sets aside 19 whose text is under 500 characters, and 410 are readable. The answers of Perplexity and Claude (API) cite 722 addresses, 601 of them readable. 147 of the 722 reuse the fetch already made for the other engines, and the check reads the saved text of that fetch; 108 of the 147 were fetched before an answer of Perplexity or Claude (API) that cites them. The other 575 were fetched on 7 and 8 October 2026, 25 to 34 hours after the answers. That fetch required both IP echo services to report New York, United States; the exit's address changes from request to request, so the two reported addresses may differ.
The study publishes each page's address, fetch time, status and SHA-256. It does not redistribute page text, and quotes are at most 200 characters.
Figures. A local model (gemma4:12b at temperature 0, with a standing briefing and two worked
examples) lists the figures in each answer.
- A figure is a price, percentage, rate, count, duration, size or limit, or a range of these.
- A figure is not a date or year, a list number, a citation marker, a number inside a product name, a phone number or an address.
- A figure is kept when it occurs character for character in its sentence. A figure the extractor returns in a changed form is counted as an instrument gap.
The extractor listed 3,465 figures in the answers of the five engines. No extraction failed.
The check. A script searches the cited pages for each figure. Found means that every part of the figure occurs on one page with the same unit (currency, percent or multiple; cents and K, M, B are normalized), and that the parts of a range lie within 300 characters of each other. A figure with no currency or percent sign is matched on the number alone, so the check leans toward "found".
The same four steps for every engine. The extraction, the check, the classifier and the summary are the frozen functions, with the same two local model passes. For Perplexity and Claude (API) a script wrote the records those functions read, and the functions were pointed at a folder of their own. A step stops when a name still points at the folder of the other three engines, or when the local model's digest is another than the one their outputs name. No frozen file is changed.
What can happen to a figure?
Each figure gets exactly one outcome.
| Outcome | Meaning |
|---|---|
| Found, attached | On a page whose citation marker sits in the same block as the figure: the figure's sentence lies inside the marker's paragraph, list item or table row |
| Found, elsewhere | On another page the answer cites |
| Not found | Every cited page was readable, and none has the figure |
| No source | The answer cites no page |
| Access failure (unknown) | Not on any readable page, and at least one cited page was unreadable or has no address in the record |
| Instrument gap (unknown) | The extractor returned the figure in a form that is not in its sentence, or the check could not read it as a value |
How a source is attached to a figure differs by engine.
| Engine | A source is attached to a figure when |
|---|---|
| Google AI Overviews | A chip or text link in the figure's block names the page |
| ChatGPT | A marker in the figure's block names a site; "attached" then means any page of that site in the answer's source list |
| Gemini | A chip in the figure's block gives the address of one page, also when its label names two sources |
| Perplexity | Not available: the records tie no passage to a source address. No mapping is guessed, so every found figure counts as found on a cited page |
| Claude (API) | A cited text block of the API response lies in the figure's line |
The share found on any cited page, the primary measure, does not depend on the attached source.
When a marker names more pages than the record has addresses for, the further pages cannot be seen. Such blocks are flagged in the record, and the results of Google AI Overviews, ChatGPT and Gemini are reported with and without the figures of flagged blocks. The collections of Perplexity and Claude (API) recorded neither such markers nor sponsored labels.
How are the figures that were not found read?
The check is a screen. Two readings measure its distance from a judgement of support. Both are model readings and are reported as such.
The first is a classifier. It gives every "not found" figure one of five labels. It was built on pilot data and frozen before the window opened. It works in three steps:
- A script cuts the passages of the cited pages that hold the figure's digits, a value within 10% of it, or the words of what it measures.
- The local model (
gemma4:12bat temperature 0, with a standing briefing and six worked examples) reads the passages. It names the passage that fits the figure most, copies a quote and the page's own figure from it, says whether that figure measures the same thing, and writes a calculation where the answer's figure needs one. The model does not name a label. - The script confirms the quote in the saved page and derives the label by fixed rules.
| Label | Rule |
|---|---|
| Present | Every part of the figure equals the page's figure, so the check missed it |
| Rounded | Every part equals the page's figure or is its rounding, and at least one part is a rounding. Rounding is half up at the figure's last non-zero digit and at most 10% away ($149 for $150) |
| Derived | A calculation gives the figure or a value that rounds to it. It may use numbers in the quote, numbers in the answer sentence and 12, 52 or 365, and at least one number must come from the quote |
| Partial | Some parts of the figure fit and others do not (one end of a range), or the page's figure fits but carries another qualifier |
| Absent | Everything else: no passage about the same thing, a number that measures something else, or a quote the script cannot confirm |
Figures that are absent for the last reason alone are counted separately.
How far is the check from a judgement of support?
The second reading is a validation sample. 40 figures are drawn with the study's seed, 20 found and 20 not found, from the figures of the three engines named in the preregistration. Claude Opus 5.5 read each one against the saved pages.
- A found figure is read as supports, supports weakly or coincidental.
- A not-found figure gets one of the five labels above.
- The reader works from the saved text of the cited pages and from the passages the classifier's script cut. The reader does not see the classifier's label.
- Every decision carries a verbatim quote of at most 200 characters, which a script confirms in the saved page. A reading of absent may come without a quote.
What the sample showed in round two:
| Reading | Figures |
|---|---|
| Found figures read as supported | 19 of 20 |
| Found figures read as a coincidence of digits | 1 of 20 |
| Not-found figures where reader and classifier agree | 16 of 20 |
| Readings with a quote that a script confirmed | 35 of 40 |
The readings are model readings of 40 figures, drawn from the three engines named in the preregistration (Google AI Overviews, ChatGPT and Gemini), and describe those three together.
A placebo tests the check from the other side. Each answer's pages are replaced by the same number of readable pages cited for questions of other sectors; the pool, the same for all five engines, is the readable pages that Google AI Overviews, ChatGPT and Gemini cite. The check then finds 28%, 18%, 20%, 40% and 24% of the figures of Google AI Overviews, ChatGPT, Gemini, Perplexity and Claude (API), against 79%, 57%, 69%, 85% and 88% on the pages the answers cite. The placebo is an additional computation beside the primary measure.
What is reported?
Primary measure. For each engine, the share of an answer's figures that are found on any page the answer cites. Every figure the extractor listed is in the denominator, and the two "found" outcomes count toward the share. The statistic is the median across answers, with a 95% percentile bootstrap interval (10,000 resamples of questions, with the study's seed).
The interval. It shows how much the median depends on which questions are in the set. Several questions are about the same product, so it is narrower than it would be for unrelated questions.
The frame. For each engine the 50 questions are accounted for: answered, no AI Overview, no answer, parse failed, and no record. So are the answers: with figures, without figures, empty text, extraction failed. An answer without a figure is outside the per-answer share.
Also reported, per engine:
- the count of every outcome, also per sector;
- the share found on the attached page, where the attached source is available;
- both shares without the unknowns;
- the share found in the answers whose cited pages were all readable;
- the shares without the figures of flagged blocks, and without the answers that hold a "Sponsored" label, for the engines whose collection recorded them;
- the classifier's labels;
- the validation counts, with the classifier's agreement on the sampled figures.
How to read the unknowns. A figure is unknown when its answer cites no source, when a cited page could not be read, or when the check could not read the figure. A figure becomes unknown only when it was not found, so the shares without the unknowns lean upward. The share in the answers whose pages were all readable does not depend on the outcome, and it is the view to read beside the primary measure. An engine that cites many pages has few "not found" figures, because one unreadable page makes a missing figure unknown.
Comparing engines. Engines are compared on the questions all of them answered, and engines whose bootstrap intervals overlap are not ranked. The comparison of the five engines stands on the 41 questions that all five answered with a figure. The three engines named in the preregistration share 44 such questions, and their comparison on those 44 is reported beside it.
Views beside the primary measure. The like-for-like sets, the paired differences, the views by number of pages and the placebo are additional computations. They stand beside the primary measure and replace nothing in it. Every interval follows the registered rule.
Minimum cell size. No percentage is given for a reporting cell with fewer than 5 units, and no median for fewer than 5 answers.
How stable is one answer?
One answer per question and engine does not show how much an answer changes from one asking to the next. For that reason the first 10 questions of the asking order are asked a second time, with the same scripts (30 more answers). The order is the seeded shuffle, so no question is chosen by hand.
The repeat run covers the three engines of the scripted collection: Google AI Overviews, ChatGPT and Gemini. It is reported as a stability measure: for each engine, the share found in the first and in the second answer to the same question, side by side. It is not pooled with the result of the first asking of the 50 questions.
What was registered, and what was recorded afterwards?
Freeze. The method note and every file it names (the question files, the selection and collection scripts, the parsers, the fetcher, the extractor, the check, the classifier, the validation and summary scripts and the run plan) were deposited on Zenodo with a SHA-256 manifest before the window opened. Nothing in them changes after that. The outputs of the two local model passes carry the digest of the model that wrote them.
Review before the freeze. The scripts were run on the ten pilot questions before the freeze. A gate that reads and cannot edit reviewed the instrument, and its findings were resolved in the files before the freeze. Trial answers are not study data.
Deviation log. What happens after the registration is written to a dated deviation log and published with the results. Registered files are never edited, and raw data are never edited by hand. A corrected computation may appear beside the registered result, never in its place: the values of Google AI Overviews, ChatGPT and Gemini are reported as registered. The log holds seven entries:
- one exit check at the end of block 1, for which the results of Gemini are given with and without the eight answers saved after the block's last passing check;
- two entries for blocks that were run a day earlier than the run plan;
- two entries for the analysis of sources and the collection of Perplexity and Claude (API). The run record of 6 October gives the route used for Perplexity, a signed-in account;
- one entry for the four steps run on the answers of Perplexity and Claude (API) on 7 and 8 October 2026;
- one entry of 8 October 2026: the results of the five engines are shown in the same tables.
The second analysis: which sources the engines cite
The collected answers are also used for a second analysis that the registration does not contain: which sources the five engines cite, and how far they share them. It is the third round of the Cross-Engine Citation Study.
It was decided on 6 October 2026, after six of the seven collection blocks, before any figure had been extracted and before any count of sources by engine had been made. It is exploratory.
Counting sources. A page is its host in lower case without "www." plus its path without a trailing slash. The query string and the fragment are dropped, except the video id of a YouTube address. A domain is the registrable domain of the host. Overlap is the Jaccard index, computed per question and reported as the mean of the per-question ratios. Containment is the share of one engine's own sources that another engine also cites for the same question. The Jaccard of several lists cannot exceed the shortest list divided by the longest, so the analysis reports that ceiling beside the observed overlap. The repeat run serves as a yardstick for the three engines it covers: the same engine asked the same question twice.
Scope of claims
The results describe these 50 questions and these five engines through the access routes named above, asked in English with New York as the location, on the collection dates, with one answer per question and engine.
The study does not claim:
- that a figure which is not found is false or invented. It measures whether the cited pages contain the figure, not whether the figure is true, and a page the engine did not cite may contain it;
- that a found figure is supported by its page. The check matches numbers, and a number without a currency or percent sign matches on the number alone. The 20 found figures of the validation sample describe three engines together, not each engine;
- that any reading is a person's or an outside party's: every reading is a model reading;
- anything from the technical pilot of 23 September 2026 (10 questions), whose counts are not pooled with these;
- a ranking of engines whose bootstrap intervals overlap, or a comparison on questions that not all of them answered.
Use of AI models
The collection of Google AI Overviews, ChatGPT and Gemini is scripted, and no AI model takes part
in it. An AI assistant collected Perplexity's answers, and Claude's answers came through the
Anthropic API. A local model (gemma4:12b) lists the figures and reads the passages for the
classifier. Claude Opus 5.5 read the validation sample.
The method was approved by the author on 4 October 2026. The first draft of this page was produced with Claude (Anthropic).
Data and code
- Preregistration. doi:10.5281/zenodo.23145556, version 1.0.0, published 5 October 2026: the method note, the question files, the scripts and a SHA-256 manifest.
- Tables and files. Every table on one page, each with its own link, and the files behind them: the computed views, the method note, the deviation log, the questions and the scripts.
- Results. The answers, the source addresses, the page hashes, the extraction, check and reading outputs, the code and the deviation log are published under CC BY 4.0 with a Zenodo DOI: doi:10.5281/zenodo.23238429.
- Findings. The AI Answer Evidence Index and round three of the Cross-Engine Citation Study.
- Site-wide approach. The methodology page.
If you find an error, tell us. We will correct it and say what changed.
References
- Minel Gunesoglu (2026). The AI Answer Evidence Index: method note (Preregistration, version 1.0.0). Zenodo. https://doi.org/10.5281/zenodo.23145556
- Minel Gunesoglu (2026). The AI Search Evidence Index: source determinability of numeric claims on AI-visibility web pages (Round one) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.22257563
- Our own Cross-Engine Citation Study, for the analysis of sources.