Is the figure on the page the answer cites?
A measurement of whether the figures in the answers of Google AI Overviews, ChatGPT, Gemini, Perplexity and Claude (API) can be found on the pages those answers cite. A figure that is not found is not thereby false, and engines whose intervals overlap are not ranked.
Open the data and codeDo the Numbers in AI Answers Have Sources? 3,465 Figures, 5 Engines
TL;DR: Round two of the Evidence Index series put 50 price and cost questions to five AI engines. On the 41 questions all five answered with a figure, the median share of an answer's figures found on a cited page runs from 57% (ChatGPT) to 92% (Perplexity). The intervals of Perplexity and Claude (API) do not overlap ChatGPT's; the other pairs overlap and are not ranked. A figure that is not found is not thereby false.
AI answers to buying questions are full of numbers: a monthly price, a fee, an average cost, a limit. Most of those answers also show their sources. People ask whether ChatGPT, Gemini and Google's AI Overviews cite their sources correctly. This study measures one checkable part of that question: when an AI answer states a figure and cites pages, is the figure on those pages?
Not whether the figure is true. Whether a reader who opens the cited pages can find it.
This is round two of the Evidence Index series. Round one, The AI Search Evidence Index (September 2026), measured whether AI-visibility guides link the sources of their numbers. Round two, The AI Answer Evidence Index, measures AI answers instead: whether the figures an answer states can be found on the pages that answer cites.
The method was published on Zenodo on 5 October 2026, before the first answer was collected (doi:10.5281/zenodo.23145556), and it is set out in full on the method page. That preregistration named Google AI Overviews, ChatGPT and Gemini; Perplexity and Claude (API) were added on 6 October, and their answers went through the same four steps on 7 and 8 October, after the results of the first three had been computed, as an exploratory addition (deviation log, entries 3, 5, 6 and 7).
The AI Answer Evidence Index
- When an AI answer gives a number, is it on the pages it cites?
- What was measured?
- Is the difference about the engine, or about how many sources it shows?
- What is each engine's median on its own answers?
- Why do ChatGPT's answers fall into two groups?
- Is the figure on the source attached to the sentence, or on another cited page?
- Does it matter whether the figure carries a unit?
- What about the figures that were not found?
- Does the result differ by sector?
- How was the check itself tested?
- Which other views are reported?
- Does the answer change when the same question is asked again?
- What does this mean for a brand?
- What do these results cover?
- Data and code
- How to cite this
When an AI answer gives a number, is it on the pages it cites?
Five engines answered the same 50 questions. On the 41 questions that all five answered with a figure, half of an engine's answers have at least this share of their figures on a page the answer cites:
| Engine | The same 41 questions: median share per answer | 95% bootstrap interval | Figures found on a cited page, all answers pooled |
|---|---|---|---|
| Google AI Overviews | 82% | 72% to 92% | 79% |
| ChatGPT | 57% | 42% to 75% | 57% |
| Gemini | 71% | 67% to 91% | 69% |
| Perplexity | 92% | 86% to 100% | 85% |
| Claude (API) | 91% | 88% to 96% | 88% |
On these 41 questions Perplexity and Claude (API) are above ChatGPT: neither interval overlaps ChatGPT's. The other 8 of the 10 pairs of intervals overlap, and those engines are not ranked against each other.
A figure that is not found is not thereby false, and a found figure is not thereby supported by its page: the check matches numbers.
How far the difference holds, on the same questions:
- Figures with a unit only (39 questions): 80%, 50%, 70%, 90% and 92%. The intervals of Perplexity and Claude (API) stay apart from ChatGPT's.
- Answers that state a figure and cite a page (35 questions): 82%, 57%, 78%, 88% and 91%. Only the intervals of ChatGPT and Claude (API) do not overlap.
- Each answer cut to two of its readable pages (41 questions): 54%, 52%, 67%, 60% and 69%. All 10 pairs of intervals overlap. The answers cite a median of 7, 2, 3, 10 and 6 pages.
- Placebo. On pages cited for questions of other sectors (one pool for all five engines: the readable pages that Google AI Overviews, ChatGPT and Gemini cite) the check finds 28%, 18%, 20%, 40% and 24% of the figures.
- Share minus placebo (41 questions): 51, 38, 59, 50 and 67 points. Only the intervals of ChatGPT and Claude (API) do not overlap.
The views that bring the page counts closer disagree, and the study attributes the difference to neither the engine nor the number of pages. The engines were also reached by different routes (logged-out interfaces, a signed-in account and an API), and the study cannot separate the route from the engine.
How to read the table:
- Share per answer. Of the figures one answer states, the part that is found on any page that answer cites.
- Median. The middle answer of an engine: half of its answers have at least this share.
- Interval. A 95% bootstrap interval, from 10,000 resamples of the questions. It shows how much the median depends on which questions are in the set. Several questions are about the same product, so it is narrower than it would be for unrelated questions.
What was measured?
Fifty questions that ask for a price, a cost or another figure were put to five AI search engines, in English, with New York as the location, on 5 to 7 October 2026. Dates and times are in UTC+3, the operator's time zone. Each answer was saved with the sources it cites.
| Engine | How it was asked | Answers collected | Answers, of 50 questions |
|---|---|---|---|
| Google AI Overviews | Logged-out public interface, by the collection script | 5 to 7 October | 49 |
| ChatGPT | Logged-out public interface, by the collection script | 5 and 6 October | 50 |
| Gemini | Logged-out public interface, by the collection script | 5 and 6 October | 50 |
| Perplexity | Signed-in free account, collected by the study's AI assistant | 6 October | 50 |
| Claude (API) | Model claude-sonnet-5-5 through the Anthropic API, with New York as its location setting | 6 October | 50 |
The full route of each engine is on the method page.
Then three steps, and a fourth that writes the summary:
- A local model listed the figures in each answer.
- A script looked for each figure in the saved text of the pages that answer cites.
- The same model, with a second briefing, labelled the figures the script did not find.
Four words carry the rest of this article:
- Figure. A price, percentage, rate, count, duration, size or limit, or a range of these.
- Cited page. A page the answer shows as a source.
- Readable. A cited page is readable when the study's fetch returned its text.
- Found. Every part of the figure occurs on one cited page, with the same unit.
The frame
| Count | Google AI Overviews | ChatGPT | Gemini | Perplexity | Claude (API) |
|---|---|---|---|---|---|
| Answers collected, of 50 questions | 49 | 50 | 50 | 50 | 50 |
| Answers with at least one figure | 47 | 47 | 50 | 47 | 50 |
| Figures the extractor listed | 640 | 470 | 661 | 755 | 939 |
| Found on any cited page | 507 of 640 (79%) | 267 of 470 (57%) | 453 of 661 (69%) | 644 of 755 (85%) | 822 of 939 (88%) |
- 249 of 250 answers were collected. One question has no Google answer after three attempts.
- The extractor listed 3,465 figures in 241 answers. No extraction failed, and no figure was left unlabelled.
- Six ChatGPT answers cite no source: the saved pages hold no "Sources" control and no outbound link, which indicates that ChatGPT answered them without a search. Two of the six return a question to the user and state no figure. Two answers of Claude (API) ran no search and cite no page.
- The answers of Google AI Overviews, ChatGPT and Gemini cite 483 addresses, 410 of them readable as the check counts them. The answers of Perplexity and Claude (API) cite 722 addresses, 601 of them readable.
Is the difference about the engine, or about how many sources it shows?
These data do not separate engine from number of sources. 45 of 46 answers with figures of Google AI Overviews that cite a page cite three or more, every such answer of Perplexity and Claude (API) does, and 32 of 43 such answers of ChatGPT cite one or two.
Each row below is a median per answer on the same questions for all five engines.
| View, as a median per answer (interval) | Questions | Google AI Overviews | ChatGPT | Gemini | Perplexity | Claude (API) |
|---|---|---|---|---|---|---|
| All figures | 41 | 82% (72% to 92%) | 57% (42% to 75%) | 71% (67% to 91%) | 92% (86% to 100%) | 91% (88% to 96%) |
| Figures with a unit only, on the questions where all five answers hold one | 39 | 80% (67% to 91%) | 50% (27% to 67%) | 70% (47% to 90%) | 90% (78% to 100%) | 92% (85% to 100%) |
| Answers that state a figure and cite a page | 35 | 82% (71% to 92%) | 57% (43% to 82%) | 78% (68% to 92%) | 88% (75% to 95%) | 91% (88% to 97%) |
| Each answer cut to two of its readable pages at random | 41 | 54% (46% to 69%) | 52% (42% to 67%) | 67% (55% to 71%) | 60% (52% to 65%) | 69% (55% to 76%) |
| Share found minus the answer's placebo | 41 | 51 points (44 to 60 points) | 38 points (30 to 50 points) | 59 points (41 to 67 points) | 50 points (38 to 58 points) | 67 points (57 to 70 points) |
Which intervals lie apart:
- All figures. ChatGPT and Perplexity, and ChatGPT and Claude (API). The other 8 pairs overlap.
- Figures with a unit only. The same two pairs. The intervals of Google AI Overviews and ChatGPT meet at one value (67%).
- Answers that state a figure and cite a page. ChatGPT and Claude (API) only.
- Cut to two pages. No pair. An answer with fewer than two readable pages is not cut: 1 of 41 answers of Google AI Overviews, 20 of 41 of ChatGPT, 11 of 41 of Gemini, 0 of 41 of Perplexity and 1 of 41 of Claude (API).
- Share minus placebo. ChatGPT and Claude (API) only.
A paired view asks the same thing question by question, with both engines resampled on the same questions. Against ChatGPT, on the 41 questions, the difference between the medians is 35 points for Perplexity (interval 18 to 50 points), 34 points for Claude (API) (interval 16 to 51 points) and 25 points for Google AI Overviews (interval 8 to 42 points). For Gemini it is 14 points, with an interval that includes zero (-0 to +36 points).
Two views bring the page counts closer, and they disagree:
- Each answer cut to two of its readable pages. All intervals overlap.
- The questions where both answers have two or more readable pages. The mean difference against ChatGPT is 15 points for Google AI Overviews (interval 6 to 25 points, 20 questions), 13 points for Perplexity (interval 5 to 22 points, 22 questions) and 19 points for Claude (API) (interval 9 to 29 points, 22 questions); for Gemini it is 7 points, with an interval that includes zero (-10 to +23 points, 17 questions).
The study reports both and attributes the difference to neither the engine nor the number of pages.
What is each engine's median on its own answers?
The primary measure uses every answer with a figure, each engine on its own answers. The medians of this table are not compared with each other, because the engines do not stand on the same questions.
| Engine | Answers with figures | Median share per answer | 95% bootstrap interval |
|---|---|---|---|
| Google AI Overviews | 47 | 82% | 72% to 92% |
| ChatGPT | 47 | 55% | 40% to 69% |
| Gemini | 50 | 73% | 66% to 83% |
| Perplexity | 47 | 94% | 86% to 100% |
| Claude (API) | 50 | 93% | 89% to 98% |
The three engines of the preregistration share 44 questions that all of them answered with a figure. On those 44 the medians are 83% for Google AI Overviews (73% to 92%), 55% for ChatGPT (40% to 69%) and 71% for Gemini (65% to 85%), and the intervals of Google AI Overviews and ChatGPT do not overlap. On the 41 questions of the five engines they overlap. The paired difference between the two medians is 28 points on the 44 (interval 14 to 46 points) and 25 points on the 41 (interval 8 to 42 points).
The 41 are the 44 without three questions for which Perplexity's answer states no figure; ChatGPT's answers to those three have 1 of 3, 0 of 1 and 0 of 2 figures found.
Why do ChatGPT's answers fall into two groups?
The engines differ in how many pages an answer cites, and in how many of those pages were readable.
| Pages per answer, and figures found by the number of pages | Google AI Overviews | ChatGPT | Gemini | Perplexity | Claude (API) |
|---|---|---|---|---|---|
| Median cited pages per answer | 7 | 2 | 3 | 10 | 6 |
| Median readable pages per answer | 6 | 1 | 2 | 9 | 5 |
| Answers with figures that have at most one readable page | 1 of 47 | 25 of 47 | 12 of 50 | 0 of 47 | 2 of 50 |
| Answers with one readable page: figures found | no such answer | 81 of 171 (47%), 19 answers | 43 of 85 (51%), 8 answers | no such answer | no such answer |
| Answers with two or more readable pages: figures found | 507 of 625 (81%), 46 answers | 186 of 262 (71%), 22 answers | 410 of 541 (76%), 38 answers | 644 of 755 (85%), 47 answers | 822 of 926 (89%), 48 answers |
| Each answer cut to two of its readable pages: pooled expectation of the share found | 57% | 54% | 63% | 59% | 67% |
25 of 47 ChatGPT answers with figures have at most one readable cited page. Those 25 answers have 39% of their figures found, and ChatGPT's other 22 answers have 71%. In 9 of the 25 a cited page could not be read. Readability is a property of the study's fetch, not of the answer.
Each engine stands on its own answers in this table. The last row is a pooled expectation over the figures, exact over every choice of pages, and not a median of answers.
A rank correlation measures whether answers with more pages tend to have a higher share. From two readable pages on, the rank correlations are +0.03, +0.11, +0.15, +0.23 and +0.33. The interval includes zero for the first four engines and not for Claude (API) (+0.07 to +0.56, on 48 answers).
Is the figure on the source attached to the sentence, or on another cited page?
An engine can attach a source to a sentence with a marker, a chip or a link. A source is "attached" to a figure when its marker sits in the same paragraph, list item or table row as the figure.
| Outcome, as a share of all figures | Google AI Overviews | ChatGPT | Gemini | Perplexity | Claude (API) |
|---|---|---|---|---|---|
| Found on the page attached to the sentence | 32% | 27% | 43% | not available | 61% |
| Found on another page the answer cites | 48% | 30% | 26% | not available | 26% |
| Not found, every cited page readable | 5% | 23% | 21% | 2 of 755 | 3% |
| Unknown: no source cited, a page that could not be read, or a figure the check could not read | 15% | 20% | 11% | 14% | 9% |
| Figures whose sentence has no source attached | 48% | 61% | 44% | not available | 25% |
| Where a source is attached: figure found on it | 61% | 69% | 76% | not available | 81% |
Perplexity's records tie no passage to a source, so this view is not available for it. For ChatGPT a marker names a site, and "attached" means any listed page of that site. For Claude (API) a source is attached when a cited text block of the response lies in the figure's line.
Does it matter whether the figure carries a unit?
A figure with a currency, percent or multiple sign must occur on the page with that unit. A figure without one is matched on the number alone: 26% of the found figures (695 of 2,693) carry no unit.
The placebo asks how many figures the check finds on pages that have nothing to do with the answer. It replaces each answer's pages with the same number of readable pages cited for questions of other sectors, as an exact expectation over every choice of such pages. The pool is the same for all five engines: the readable pages that Google AI Overviews, ChatGPT and Gemini cite.
| Share found | Google AI Overviews | ChatGPT | Gemini | Perplexity | Claude (API) |
|---|---|---|---|---|---|
| All figures | 79% | 57% | 69% | 85% | 88% |
| Placebo, all figures | 28% | 18% | 20% | 40% | 24% |
| Figures with a unit | 75% | 51% | 66% | 81% | 87% |
| Placebo, figures with a unit | 19% | 9% | 10% | 33% | 17% |
| Figures without a unit | 92% | 77% | 78% | 100% | 88% |
| Placebo, figures without a unit | 55% | 50% | 47% | 66% | 54% |
An answer draws as many placebo pages as it has readable pages, a median of 9 for Perplexity, whose placebo is the highest of the five. All 177 of Perplexity's figures without a unit were found.
What about the figures that were not found?
314 figures were not found although every page their answer cites was readable. An engine that cites many pages has few such figures, because one unreadable page makes a missing figure unknown: 42 of Perplexity's 47 answers with figures cite at least one page that could not be read.
| Figures not found on a cited page | Google AI Overviews | ChatGPT | Gemini | Perplexity | Claude (API) |
|---|---|---|---|---|---|
| Not found, every cited page readable | 34 | 109 | 137 | 2 | 32 |
| Not found, counted with the figures of answers that have an unreadable page | 18% | 37% | 29% | 14% | 11% |
| Figures in answers that cite no source | 2% | 6% | 2% | 0% | 1% |
| Figures the check could not read | 1 | 0 | 5 | 2 | 1 |
| All figures not found on a cited page | 21% | 43% | 31% | 15% | 12% |
A figure that is not found is not thereby false. The study measures whether the cited pages contain the figure. A page the engine did not cite may contain it.
What the classifier said about the 314
The classifier is the local model's second pass, with fixed rules that turn its reading into one of five labels.
| Label | Google AI Overviews | ChatGPT | Gemini | Perplexity | Claude (API) | All five engines |
|---|---|---|---|---|---|---|
| Absent | 18 | 86 | 80 | 0 | 19 | 203 |
| Partial: one end of a range fits, or the page's figure carries another qualifier | 15 | 21 | 51 | 2 | 8 | 97 |
| Rounded | 1 | 2 | 5 | 0 | 5 | 13 |
| Present | 0 | 0 | 1 | 0 | 0 | 1 |
| Derived | 0 | 0 | 0 | 0 | 0 | 0 |
Eleven of the 203 are absent only because the script did not confirm the classifier's quote in the saved page.
The unknowns
A figure is unknown when its answer cites no source, when a cited page could not be read, or when the check could not read the figure. The view that does not depend on the outcome is the set of answers whose cited pages were all readable.
| Answers whose cited pages were all readable | Google AI Overviews | ChatGPT | Gemini | Perplexity | Claude (API) |
|---|---|---|---|---|---|
| Answers | 18 | 30 | 36 | 5 | 25 |
| Share found, pooled over their figures | 89% | 67% | 71% | 97% | 92% |
| Share found, median per answer (interval) | 95% (81% to 100%) | 61% (46% to 87%) | 75% (66% to 91%) | 100% (88% to 100%) | 100% (91% to 100%) |
Each engine stands on its own questions in this table, so the medians of the row are not compared with each other.
Does the result differ by sector?
The 50 questions come from six sectors.
| Sector: figures found on any cited page | Google AI Overviews | ChatGPT | Gemini | Perplexity | Claude (API) |
|---|---|---|---|---|---|
| AI tools | 112 of 113 (99%) | 78 of 85 (92%) | 87 of 88 (99%) | 115 of 116 (99%) | 118 of 119 (99%) |
| Business software | 71 of 90 (79%) | 40 of 92 (43%) | 118 of 157 (75%) | 200 of 250 (80%) | 229 of 240 (95%) |
| Consumer electronics | 61 of 89 (69%) | 51 of 89 (57%) | 32 of 74 (43%) | 107 of 117 (91%) | 104 of 127 (82%) |
| Legal and local services | 111 of 135 (82%) | 25 of 65 (38%) | 106 of 150 (71%) | 90 of 105 (86%) | 168 of 182 (92%) |
| Personal finance | 89 of 122 (73%) | 55 of 84 (65%) | 65 of 120 (54%) | 80 of 98 (82%) | 126 of 162 (78%) |
| Travel | 63 of 91 (69%) | 18 of 55 (33%) | 45 of 72 (62%) | 52 of 69 (75%) | 77 of 109 (71%) |
The share found is highest for questions about AI tools, pooled over 8 or 9 answers per engine. In legal and local services 26 of ChatGPT's 65 figures are in three answers that cite no source.
The sector cells are pooled counts over 7 to 9 answers and carry no interval. The study does not rank engines within a sector.
How was the check itself tested?
The check is a screen: it matches numbers. Two model readings measure its distance from a judgement of support. One is the classifier above. The other is a validation sample.
A sample of 40 figures was drawn with the study's seed from the figures of Google AI Overviews, ChatGPT and Gemini, 20 found and 20 not found, and read against the saved pages by Claude Opus 5.5. The reader worked from the saved text of the cited pages and from the passages the classifier's script cut, and did not see the classifier's labels. 35 of the 40 readings carry a quote that a script confirmed in the saved page.
| Sample | Google AI Overviews | ChatGPT | Gemini |
|---|---|---|---|
| Found figures (20) | 6 | 3 | 11 |
| Not-found figures (20) | 1 | 9 | 10 |
- Found figures. Of the 20, 19 were read as supported and one, a number without a unit, as a coincidence of digits: the page gives the number for something else. None was read as weakly supported.
- Not-found figures. Of the 20, reader and classifier agree on 16. In the sample, every disagreement moved a figure out of absent: 4 of the 15 figures the classifier called absent were read as partial (2) or as derived by a calculation from the page (2).
The readings are model readings of 40 figures and describe those three engines together.
Which other views are reported?
The first three are views the registration asks for. The fourth concerns Perplexity, which was added after it.
- Markers that name an unlisted page. Without the figures whose citation marker names a page the interface did not list, the pooled shares are 54% for ChatGPT and 67% for Gemini. Google AI Overviews is unchanged at 79%. The collections of Perplexity and Claude (API) did not record such markers.
- Sponsored labels. No answer of Google AI Overviews, ChatGPT or Gemini holds a "Sponsored" label in its text. Two of Google's result pages carried the label outside the AI Overview. The collections of Perplexity and Claude (API) did not record the label.
- Gemini, block 1. The network exit was checked during collection. Gemini without the eight answers of block 1 that were saved after the block's last passing exit check: 42 answers, median 69%, interval 61% to 85%. With them: 50 answers, median 73%, interval 66% to 83%.
- One Perplexity answer. Its answer to one question (SVC-07) was read from the account's stored thread after an interruption. Without it: 46 answers, median 93%, interval 86% to 100%. With it: 47 answers, median 94%, interval 86% to 100%. That question is not among the 41.
Does the answer change when the same question is asked again?
The first ten questions of the asking order were asked a second time on 7 October, of Google AI Overviews, ChatGPT and Gemini (the repeat run covers these three engines). 28 of the 30 second answers are complete: Google gave no complete AI Overview for two of the questions.
Pooled over the questions that have figures in both answers:
| Engine | Questions | First answers | Second answers |
|---|---|---|---|
| Google AI Overviews | 8 | 84 of 118 (71%) | 69 of 100 (69%) |
| ChatGPT | 8 | 30 of 56 (54%) | 31 of 60 (52%) |
| Gemini | 10 | 100 of 132 (76%) | 118 of 149 (79%) |
Single questions move more than the pooled shares. ChatGPT's answer on local moving prices went from 5 of 6 figures found to 0 of 6, and Gemini's answer on the cost of a gaming laptop from 5 of 17 to 14 of 22.
A second answer can state other figures and cite other pages. Google AI Overviews cited again 32 of the 51 pages of its first answers, ChatGPT 5 of 17 and Gemini 17 of 29, counted on the 7, 9 and 10 questions where both answers of the engine cite a page.
The repeat is a stability measure on ten questions and is not pooled with the result above. Which sources the five engines cite, and how far they share them, is the subject of round three of the Cross-Engine Citation Study.
What does this mean for a brand?
This section is a reading of the results above, not a result.
- For all five engines, more of the figures were found on a cited page than not: 79%, 57%, 69%, 85% and 88%, pooled.
- Many figures are on a cited page other than the one attached to their sentence: 48%, 30% and 26% of all figures of Google AI Overviews, ChatGPT and Gemini, and 26% of the figures of Claude (API), were found that way. A check of an answer that opens the attached source and stops there does not see those figures.
- Answers differ in how much they give a reader to check: 25 of ChatGPT's 47 answers with figures had at most one readable cited page, against 1 of 47 for Google AI Overviews and 0 of 47 for Perplexity.
- One answer is one draw. In the repeat run single questions moved more than the pooled shares, so a check of one answer describes that answer.
The manual way to check what an engine says about a brand is in how to track brand mentions in ChatGPT and how to measure AI search visibility.
What do these results cover?
The results describe these 50 questions and these five engines through the access routes named above, asked in English with New York as the location, on the collection dates, with one answer per question and engine. Claude's answers are what the model returns through the API with a search tool; they are not a reading of its consumer interface.
The study does not claim that a figure which is not found is false, or that a found figure is supported by its page. Every reading in it is a model reading. It does not rank engines whose intervals overlap, and it does not attribute the difference between the medians to the engine or to the number of sources it shows.
Data and code
- Method. How we measured it: the questions, the collection, the check, the readings and the scope of claims. The site-wide approach is on the methodology page.
- Preregistration. doi:10.5281/zenodo.23145556, published 5 October 2026 before the first answer was collected.
- Tables and files. Every table on one page, each with its own link, and the files behind them: the computed views, the method note, the deviation log, the questions and the scripts.
- Data and code. The answers, the source addresses, the page hashes, the extraction, check and classification outputs, the validation readings, the code and the deviation log are published for all five engines under CC BY 4.0 with a Zenodo DOI: doi:10.5281/zenodo.23238429.
The study publishes each cited page's address, fetch time, status and SHA-256. It does not redistribute page text, and quotes are at most 200 characters.
If a count here does not match what you get from the files, tell us. We will correct it and say what changed.
How to cite this
Minel Gunesoglu (2026). The AI Answer Evidence Index: do the numbers in AI answers have sources? (Round two) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.23238429
The preregistration is cited separately:
Minel Gunesoglu (2026). The AI Answer Evidence Index: method note (Preregistration, version 1.0.0). Zenodo. https://doi.org/10.5281/zenodo.23145556