# The AI Answer Evidence Index: method note **Do the numbers in AI answers have sources?** Approved by the owner on 4 October 2026, and revised after the collection trial and one review round (section 13). This note is deposited on Zenodo as the study's preregistration before the collection window opens. Author: Minel Gunesoglu, Is My Brand in AI. Second study in the series that began with The AI Search Evidence Index (DOI 10.5281/zenodo.22257563). ## 1. Question When an AI answer engine states a figure and shows its sources, can a reader find that figure on the pages shown? The unit is one figure in one answer, together with the pages that answer cites. The study is a descriptive measurement. It tests no hypothesis, and every result is published, whatever it shows. ## 2. Design | | | |---|---| | Questions | 50 buyer questions about prices, costs, fees and value, from six sectors | | Engines | Google AI Overviews, ChatGPT, Gemini | | Askings | 150: each engine is asked each question once. Google may show no AI Overview, so the answers can be fewer | | Access | The free public web interface. ChatGPT: guest, temporary chat. Google Search: logged out. Gemini: logged out. One Chrome profile is kept for this study only. | | Locale | English interface; network exit in New York (US), verified before every block | | Window | Planned for three consecutive days after the deposit; every answer carries its timestamp | Claude answered through its API in the pilot. That is a different access route from the public interfaces, so Claude is not part of the main run. Perplexity was planned as a fourth engine. In the trial of 4 October its verification page did not admit the script-driven browser window, also after a person passed the check in that window. The study does not work around a verification, so Perplexity is not collected. The trial attempts are published with the code. ## 3. Questions - **Frame.** `questions-v1.tsv` holds 100 questions taken verbatim from Google autocomplete (`hl=en`, `gl=us`) on 22 September 2026, by the written rule in `SELECTION-RULE.md`. - **Main run.** `select_main_questions.py` draws 50 of them into `questions-main-v1.tsv`: - The 10 pilot questions are excluded, because the instrument was adjusted while their answers were visible. Each sector keeps 15 eligible questions. - Quotas follow the frame's sector order: 9 each for business software and AI tools, and 8 each for personal finance, legal and local services, consumer electronics and travel. - The seed is the first 32 bits of the frame's SHA-256 (1109872877). That hash was fixed on 22 September, so no number was chosen for this draw. - The 50 questions are then shuffled once with the same generator. Every engine is asked in that order. - **Text.** Each question is sent exactly as autocomplete wrote it. ## 4. Collection - **Scripted.** A fixed script per engine drives a Chrome browser. Collection is automated, and no AI model takes part in it. - **Pace.** One question at a time, 60 to 120 seconds apart, in blocks of 8 questions fixed by the asking order (`RUN.md`). Each block goes to all three engines (Google AI Overviews, ChatGPT, Gemini, in that order) before the next block starts, so the three answers to a question fall within the same hours, unless a question is asked again in a later block. Before a block's first Google question the script opens Google's home page once. The network exit is checked before each engine and at the end of a block, and a failed check ends the block. - **Limits.** The script does not log in, accept terms, solve a challenge or work around a limit. A check page that moves on by itself within 20 seconds is not a stop. At any other verification page the script clicks nothing and waits up to five minutes for the operator, who may pass the page by hand. Every attempt records how long such a page was shown; a page shown for 20 seconds or longer that then gave way is recorded as passed by the operator. If the page stays, or at a usage limit, the script stops that engine for the block and records the status. A stop is read from the page only while no answer is on it, so the wording of an answer cannot stop an engine. - **Attempts.** An attempt is complete when the interface shows the answer as finished. An incomplete attempt is saved, and the question is asked again in a later block, up to three attempts. An attempt stopped by a verification page or a limit is saved and does not count toward the three, and no question is opened more than six times. The first complete attempt is the record, and no complete answer is replaced. A Google results page that shows no AI Overview once it has stopped changing is a complete observation; any other page without one is an incomplete attempt. A question with no complete attempt stays in the frame as "no answer". - **What is saved.** For every attempt: the page HTML, the visible text, a screenshot, the time, the address, the model label where the interface shows one, and the interface language. A parser frozen with the scripts reads the answer text, the source list and the sources attached to each block from the saved HTML. Sponsored and shopping cards are not part of the answer or its sources. Gemini places shopping cards inside an answer; the parser removes them from the answer text, keeps their text in the record and counts them. In the other two interfaces the trial showed no card inside an answer, so no card is removed there: the record counts the stand-alone "Sponsored" labels on the saved page and inside the answer text, and each engine is also reported without the answers that hold such a label. - **Google's link tokens.** Google AI Overviews writes its links as opaque tokens. Where the saved page holds no address for a cited page, the script opens Google's own link once, as a reader's click would, and records the first address outside Google. These addresses are kept in a separate file and marked as such in the record. A link that a verification page stopped is opened again in a later run. A cited page on one of Google's own sites cannot be told apart from Google's link pages and stays without an address. A cited page whose address stays unknown counts as cited and unreadable. ## 5. Cited pages - Every address the answers of a block cite is fetched from the same New York exit after that block, planned within 24 hours of the answer: once with a plain HTTP request and once rendered in headless Chrome. The time of every answer and of every fetch is published. PDF text is extracted. An address cited by several answers is fetched once, and all of them are checked against that snapshot. - The fetcher identifies itself as an automated research fetcher in its user agent (`AIAnswerEvidenceIndex/1.0`). A page that blocks it or shows a challenge is recorded as an access failure. - A page is readable when either fetch returns its text. A fetched text under 500 characters, or under 2,000 characters with a phrase from a fixed list of check-page phrases (`chain_common.py`), is a stub or a check page and not the page's text. - The study publishes each page's address, fetch time, status and SHA-256. It does not redistribute page text, and quotes are at most 200 characters. ## 6. Figures `evidence-extractor:v1` lists the figures in each answer. It is `gemma4:12b` at temperature 0 with a standing briefing and two worked examples, and it runs locally. - **A figure is** a price, percentage, rate, count, duration, size or limit, or a range of these. - **A figure is not** a date or year, a list number, a citation marker, a number inside a product name, a phone number or an address. - A figure is kept only when it occurs character for character in its sentence. A figure the extractor returns in a changed form cannot be checked and is counted as an instrument gap. Figures the extractor does not list are not measured. - A model call that fails is repeated once. A second failure is final, and the answer is reported as "extraction failed". - The briefing was written when the design had five engines. It is unchanged, so the extractor is the one the pilot ran. ## 7. Check `check_figures.py` searches the cited pages for each figure. **Found** means that every part of the figure occurs on one page with the same unit (currency, percent or multiple; cents and K, M, B are normalized), and that the parts of a range lie within 300 characters of each other. A figure with no currency or percent sign is matched on the number alone, so the check leans toward "found". Each figure gets exactly one outcome: | Outcome | Meaning | |---|---| | Found, attached | On a page whose citation marker sits in the same block as the figure: the figure's sentence lies inside the marker's paragraph, list item or table row | | Found, elsewhere | On another page the answer cites | | Not found | Every cited page was readable, and none has the figure | | No source | The answer cites no page | | Access failure (unknown) | Not on any readable page, and at least one cited page was unreadable or has no address in the record | | Instrument gap (unknown) | The extractor returned the figure in a form that is not in its sentence, or the check could not read it as a value | How a marker names its page differs by engine. - A Google chip or text link names its pages. - A Gemini chip gives the address of one page, also when its label names two sources. - A ChatGPT marker names a site, and "attached" then means any page of that site in the answer's source list. The guest interface does not open a marker's details (it asks for a login). When a marker names more pages than the record has addresses for, the further pages cannot be seen. Such blocks are flagged in the record, and each engine's results are reported with and without the figures of flagged blocks. A marker that carries no address at all is a cited page without an address. ## 8. Reading The check is a screen. Two readings measure its distance from a judgement of support. Both are model readings and are reported as such. - **Classifier.** `classify_figures.py` gives every "not found" figure one of five labels. It is built on pilot data only and frozen before the window opens. It works in three steps: 1. The script cuts the passages of the cited pages that hold the figure's digits, a value within 10% of it, or the words of what it measures. 2. `evidence-classifier:v1` reads the passages. It is `gemma4:12b` at temperature 0 with a standing briefing and six worked examples, and it runs locally. It names the closest passage, copies a quote and the page's own figure from it, says whether that figure measures the same thing, and writes a calculation where the answer's figure needs one. The model does not name a label. 3. The script confirms the quote in the saved page and derives the label by fixed rules: - *present*: every part of the figure equals the page's figure, so the check missed it; - *rounded*: every part equals the page's figure or is its rounding, and at least one part is a rounding. Rounding is half up at the figure's last non-zero digit and at most 10% away ($149 for $150); - *derived*: the calculation gives the figure or a value that rounds to it. It may use numbers in the quote, numbers in the answer sentence and 12, 52 or 365, and at least one number must come from the quote; - *partial*: some parts of the figure fit and others do not (one end of a range), or the page's figure fits but carries another qualifier; - *absent*: everything else: no passage about the same thing, a number that measures something else, or a quote the script cannot confirm. A page figure without a unit sign fits on the number alone. Figures that are absent only because the quote was not confirmed are counted separately. A model call that fails is repeated once; after a second failure the figure is reported as unlabelled. - **Validation sample.** 40 figures are drawn from all engines together with the seed above: 20 found and 20 not found. Claude reads each one against the saved pages, and the model version is recorded. - A found figure is read as *supports*, *supports weakly* or *coincidental*. - A not-found figure gets one of the five labels above. - The reader works from the saved text of the cited pages and from the passages the classifier's script cut. The reader does not see the classifier's label. - Every decision carries a verbatim quote of at most 200 characters, which a script confirms in the saved page. A reading of *absent* may come without a quote. - If fewer than 20 figures are not found, all of them are read and the remainder is drawn from found figures. - `validation_sample.py` draws the sample once. It accepts the readings only when the file names the reader, every figure has one reading, every quote is confirmed and every sampled not-found figure has a classifier row. A figure the classifier could not label is reported as unlabelled and left out of the agreement counts. ## 9. Measures - **Primary.** For each engine, the share of an answer's figures that are found on any page the answer cites. Every figure the extractor listed is in the denominator, and only the two "found" outcomes count toward the share. The statistic is the median across answers, with a 95% percentile bootstrap interval (10,000 resamples of questions, the seed above). - **The frame.** For each engine the 50 questions are accounted for: answered, no AI Overview, no answer, parse failed, and never opened (no record). So are the answers: with figures, without figures, empty text, extraction failed. An answer without a figure is outside the per-answer share. - **Also reported, per engine:** - the count of every outcome, also per sector; - the share found on the attached page; - both shares without the unknowns; - the share found in the answers whose cited pages were all readable; - the shares without the figures of flagged blocks, and without the answers that hold a "Sponsored" label; - the classifier's labels; - the validation counts, with the classifier's agreement on the sampled figures. - **How to read the unknowns.** A figure becomes unknown only when it was not found, so the shares without the unknowns lean upward. The share in the answers whose pages were all readable does not depend on the outcome, and it is the view to read beside the primary measure. An engine that cites many pages has few "not found" figures, because one unreadable page makes a missing figure unknown; the classifier's labels therefore describe mostly the answers that cite few pages. - **Small cells.** No percentage is given for a reporting cell with fewer than 5 units, and no median for fewer than 5 answers. The per-answer share is computed for every answer with at least one figure. ## 10. Repeat run One answer per question and engine cannot show how much an answer changes from one asking to the next. The first 10 questions in the order of `questions-main-v1.tsv` are therefore asked a second time, on a later day of the window, with the same scripts (30 more answers). The order is the seeded shuffle of section 3, so no question is chosen by hand. The repeat is reported as a stability measure: for each engine, the share found in the first and in the second answer to the same question, side by side. It is not pooled with the main result. ## 11. Freeze and publication - **Freeze.** This note and every file it names (the question files, the selection and collection scripts, the parsers, the fetcher, the extractor, the check, the classifier, the validation and summary scripts and the run plan `RUN.md`) are deposited on Zenodo with a SHA-256 manifest before the window opens. Nothing in them changes after that. The outputs of the two local models carry the digest of the model that wrote them. - **Deviations.** A defect found later is recorded as a dated deviation and published with the results. A corrected computation may appear beside the main result, never in its place. - **Raw data** are never edited by hand. - **Publication.** The answers, the source addresses, the page hashes, the extraction, check and reading outputs, the code and the deviation log are published under CC BY 4.0 with a Zenodo DOI. ## 12. Scope of claims The results describe these 50 questions and these three engines through their free public interfaces, in English from a New York exit, on the collection dates, with one answer per question and engine. The study does not claim: - that a figure which is not found is false or invented. It measures whether the cited pages contain the figure, not whether the figure is true, and a page the engine did not cite may contain it; - that a found figure is supported by its page. The check matches numbers, and a number without a currency or percent sign matches on the number alone. The 20 found figures of the validation sample describe the engines together, not each engine; - that any reading is human or independent; - anything from the technical pilot of 23 September 2026 (10 questions), whose counts are not pooled with these; - a ranking of engines whose bootstrap intervals overlap, or a comparison on questions that not all of them answered. The interval shows how much the median depends on which questions are in the set. Several questions are about the same product, so it is narrower than it would be for unrelated questions. ## 13. Record The owner's decisions are in `OWNER-DECISIONS.md`. Decisions 2, 3, 5, 6 and 8 and the decisions of 1 and 4 October 2026 (9 to 13) apply. The earlier multi-round codebook protocol is retired, and its files stay in the archive. **Trial and rehearsal.** The scripts were run on the ten pilot questions before the freeze: three in a trial on 4 October 2026 and seven as a rehearsal block on 5 October. The trial's answer records and attempt notes, with every stopped attempt, are in `raw//trial/`; `code/collection/engines/README.md` describes what was observed. Trial answers are not study data. **Review round.** A gate that reads and cannot edit reviewed the instrument once, on 4 October 2026. Its findings were resolved in the files before the freeze; the README above lists them. A gate read the fixes once more on 5 October 2026, and those findings were resolved as well. **Hashes.** The SHA-256 of every registered file is in `MANIFEST.sha256` of the deposit.