Where Do ChatGPT, Perplexity, Gemini and Google Get Their Information? (We Tested All Four)
TL;DR: ChatGPT gets its information from a few big publishers and Reddit; Perplexity casts the widest net; Gemini cites the fewest and most obscure. Asked the same buyer questions, all four engines shared just 1% of their sources, but 32% of the brands they recommended.
Most advice about "getting cited by AI" treats the AI as one thing. Optimize your page, earn some mentions, and you will show up in the answers. But "the answers" are not one place. ChatGPT, Perplexity, Gemini and Google's AI Overview each run their own retrieval, and we wanted to know a simple thing: where does each one actually get its information, and when you ask them the same question, do they reach for the same sources?
So we tested it. We took a set of ordinary B2B-SaaS buyer questions ("best CRM for startups," "HubSpot vs Salesforce," "how to reduce customer churn," "cheaper alternatives to HubSpot") and ran them through all four engines in June 2026, logging every source each one cited. (Full method at the bottom; we ran real prompts and recorded what actually appeared, nothing modeled.)
The short answer: no, they barely overlap. The same question produces four almost entirely different source lists. If you are planning your visibility around "AI," you are planning around four different games.
Then we re-ran the whole thing in August 2026 and recorded something we had missed the first time: not just the sources each engine read, but the brands each one recommended. That produced the more useful half of this study, and it points the opposite way: the four engines share almost no sources but agree on nearly a third of their recommendations.
Where Do ChatGPT, Perplexity, Gemini and Google Get Their Information? (We Tested All Four)
- What sources did each engine cite for the same question?
- Do ChatGPT and Perplexity cite the same sources?
- Where each engine gets its information
- Does the type of question change which sources get cited?
- The August round in numbers
- Which sources show up in more than one AI engine?
- Do they at least recommend the same brands?
- How long does a citation last? We checked seven weeks later
- How to get cited when every engine reads something different
- Methodology
- Cite this study
What sources did each engine cite for the same question?
Here is the cleanest way to see it. This is the source list each engine cited for the exact same query, "best CRM for startups", on the same day, in round 1 (19 June 2026):
| Engine | Sources it cited |
|---|---|
| Perplexity | crm.org, tech.co, zendesk.com, techradar.com, hubspot.com, digitalocean.com, salesflare.com |
| ChatGPT | techradar.com, tekpon.com, toolcheckr.com, Best CRM Reviews, saasprobe.com, Better Launch, reddit.com |
| Gemini | UpliftGTM, Lightfield, Dalil AI, ZoFlowX |
| Google AI Overview | Podium, plus reddit.com, salesforce.com, techradar.com, zoho.com, monday.com, quickbase.com |
Look at how little they share. Only two sites appear in more than one list: techradar.com (in Perplexity, ChatGPT and the AI Overview) and reddit.com (in ChatGPT and the AI Overview). Neither reaches Gemini. In this June round, not a single source is cited by all four engines. (In the August round two were; see the limitations below.) Gemini's four sources overlap with the other three engines exactly zero times. The same buyer, asking the same question, gets recommended pages from four non-overlapping corners of the web.
This was not a fluke of one question. We saw the same pattern across every query we ran on all the engines.
Do ChatGPT and Perplexity cite the same sources?
For six questions we ran head-to-head through Perplexity, ChatGPT and Gemini (covering recommendation, comparison, how-to and alternative queries). Here is the overlap between the two engines that cite the most sources, Perplexity and ChatGPT, measured as shared domains on the identical query:
| Question | Shared sources (Perplexity ∩ ChatGPT) |
|---|---|
| best CRM for startups | techradar.com (1) |
| best project management software | none (0) |
| HubSpot vs Salesforce | forbes.com, reddit.com (2) |
| Mailchimp vs Klaviyo | technologyadvice.com (1) |
| how to reduce customer churn | none (0) |
| cheaper alternatives to HubSpot | reddit.com (1) |
Two of six questions had zero shared sources. The rest had one or two. The average was under one shared domain per question, out of eight to ten cited by Perplexity and four to seven by ChatGPT.
Bring Gemini in and the overlap does not shrink. It disappears. Across all six questions in the June round, not one domain was cited by all three engines. Not on the recommendation questions, not on the comparisons, not on the how-tos. The closest any question came was "HubSpot vs Salesforce," where Perplexity shared two domains with Gemini and two different ones with ChatGPT, but nothing that all three reached for. On "how to reduce customer churn," Gemini cited no sources at all.
The takeaway is blunt: ranking in one AI engine tells you almost nothing about whether you will be cited in another. They are not reading from a shared shortlist.
Where each engine gets its information
The lists are not just different. They are different in character. After running the set, each engine had a recognizable citation style.
Google's AI Overview is the most mainstream. Its sources are essentially Google's own organic winners, condensed: the big review sites (TechRadar), the established vendors (Salesforce, Zoho, HubSpot, Monday), the well-known listicles (Podium), and Reddit. If a page already ranks on page one of Google, it has a real shot at the AI Overview. Classic SEO still does most of the work here. (We have two deeper studies on this engine alone, the most-cited sources in AI Overviews and who it cites by industry, and Reddit was a top source in both.)
Perplexity casts the widest net. In round 1 it consistently cited the most sources per answer (often eight to ten; in round 2 the range widened to eight to nineteen) and the broadest mix: review sites, vendor pages, Reddit threads, even YouTube videos. For comparison questions like "Notion vs Asana," it pulled in YouTube three times, including a Japanese-language video. Perplexity is the most transparent engine, and the one with the most doors to get in through.
ChatGPT concentrates on a few big publishers, then a long tail of obscure blogs. When its web search is on, a handful of names recur across almost every answer (TechRadar, Forbes, Reddit) alongside a surprising volume of small, unfamiliar sites (think "StackFYI," "Prodify Flow," "RetentionCheck") that look like SEO-built content. Once it even cited an academic arXiv paper for a churn question. One important quirk: with web search off, ChatGPT often cited no third-party sources at all. It answered from memory and linked only to the vendors' own homepages. That is the default experience for a lot of free users, and it means the only "source" many people see is the brand's own site.
Gemini cites the least, and the most obscure. It returned the fewest sources of any engine (often just one or two, sometimes none) and they were almost entirely small, unrecognized sites you have likely never heard of (UpliftGTM, ZoFlowX, RichClicks). On the "how to reduce customer churn" question, it cited nothing at all and answered purely from its own training. Gemini is the hardest engine to influence and the hardest to predict.
You can line these up on a spectrum, from the most predictable engine to the least:
| Engine | Sources per answer | Source character | What moves it |
|---|---|---|---|
| Google AI Overview | medium | Most mainstream (Google's organic winners) | Classic SEO; rank on page one |
| Perplexity | most (8–10) | Broad: reviews, vendors, Reddit, YouTube | Be everywhere; many entry points |
| ChatGPT | fewest (0–9, median 2.5) | A few big publishers + many niche blogs | Earn big-publisher and Reddit mentions |
| Gemini | fewest (0–4) | Most niche/obscure; sometimes none | Hardest; least predictable |
Does the type of question change which sources get cited?
There was a second pattern hiding in the data: the type of question changes which kind of source gets cited, and it changes the same way across engines even when the specific domains do not match.
- "Best [category]" questions pulled from review sites and listicles (TechRadar, PCMag, G2, Cloudwards). Reddit and YouTube were mostly absent.
- "X vs Y" comparisons leaned on YouTube and Reddit, plus each vendor's own comparison page. This is where brand-owned pages actually got cited.
- "How to" / strategy questions pulled from company blogs: for "reduce customer churn," Perplexity cited ten company blogs and vendor pages (Stripe, HubSpot, Qualtrics, Optimove, Contentsquare among them) and no Reddit or YouTube at all.
- "Alternatives" and "cheaper than" questions were the most community-driven, surfacing Reddit and even LinkedIn posts alongside competitor listicles.
If you only have budget to earn one kind of placement, this tells you where to spend it for the queries you care about.
The August round in numbers
Every figure below is recomputed from the published CSV. Nothing here is estimated.
| Citations logged | 322 |
| Distinct sources | 193 |
| Cited by one engine only | 151 (78.2%) |
| Cited by all four engines | 4 (2.1%) |
| Runs that showed no source at all | 7 of 48 |
Per engine, across the same twelve questions:
| Engine | Citations | Distinct sources | Average per answer |
|---|---|---|---|
| Perplexity | 138 | 106 | 11.5 |
| Google AI Overview | 104 | 77 | 8.7 |
| ChatGPT | 43 | 35 | 3.6 |
| Gemini | 37 | 35 | 3.1 |
And by the kind of page cited:
| Source type | Citations | Share |
|---|---|---|
| The vendor's own site | 156 | 48.4% |
| Independent listicle or blog | 85 | 26.4% |
| Review media | 50 | 15.5% |
| YouTube | 16 | 5.0% |
| 14 | 4.3% | |
| Forum | 1 | 0.3% |
Reddit and YouTube together made up 17.3% of Google's AI Overview citations, 18 of 104, far more than any other engine. ChatGPT was next at 7%.
Because the engines overlap so little, no single engine covers much of the set. Averaged across the twelve questions, this is how much of each question's cited set you would see by watching one engine, or two:
| What you watch | Share of the cited set |
|---|---|
| ChatGPT alone | 14% |
| Perplexity + Google AI Overview | 82% |
| Those two plus ChatGPT | 91% |
| Gemini's unique contribution on top | 9 points |
Sourcing cuts the other way too. If engines disagree this much about which sources to cite, it is worth asking whether the pages they cite can be checked at all, and we measured that in a companion study, reading 252 numeric claims one at a time.
Which sources show up in more than one AI engine?
Almost nothing is universal, but three things came closest, and they are the practical core of a cross-engine strategy.
- Reddit. It was the single most common bridge between engines, showing up in comparison, alternative and "best free" answers across Perplexity, ChatGPT and the AI Overview. If there is one off-site place to be discussed, it is the relevant subreddit. (This matches what we found studying Google's AI Overview on its own, where Reddit was the most-cited site by a wide margin.)
- Your own site, but mostly on comparison queries. For "HubSpot vs Salesforce" and "Mailchimp vs Klaviyo," the engines frequently cited the vendors' own comparison pages. A clear, honest "us vs them" page on your domain is one of the few brand-owned assets AI engines reliably reach for.
- A small set of trusted publishers per engine. TechRadar showed up in three of the four engines in the June round (Perplexity, ChatGPT and the AI Overview); in August two sources reached all four engines on a single question (zapier.com and project-management.com) and that question was a recommendation query, not a comparison. On the comparison questions no source at all reached all four. Getting included in the big roundups still pays off in more than one place.
Do they at least recommend the same brands?
Everything above is about sources, the pages an engine reads. That left the more interesting question unanswered: if four engines read four different corners of the web, do they end up recommending four different shortlists of products?
So in August 2026 we ran the study again, and this time we logged the brand names in each answer alongside the sources. Seven of the questions were open-ended recommendation queries, the ones where a buyer is genuinely asking the engine to pick for them, and all four engines answered every one.
The result flips the headline on its head.
| Question | Brands all four named | Brand overlap | Sources all four cited | Source overlap |
|---|---|---|---|---|
| best CRM for startups | HubSpot, Pipedrive | 22% | none | 0% |
| best project management software | Asana, ClickUp, Trello | 38% | zapier.com, project-management.com | 6% |
| best email marketing platform | ActiveCampaign, Kit | 25% | none | 0% |
| best AI writing tools | ChatGPT, Claude | 25% | none | 0% |
| Salesforce alternatives | HubSpot, Pipedrive, Zoho, MS Dynamics | 36% | none | 0% |
| cheaper alternatives to HubSpot | Zoho, Pipedrive, EngageBay | 30% | none | 0% |
| best free project management tools | Trello, ClickUp, Asana | 50% | none | 0% |
| Average | 32% | 1% |
The four engines agreed on 32% of the brands they recommended and 1% of the sources they cited. Every one of the seven questions produced at least one brand all four engines named. Six of the seven produced no shared source at all.
That is the finding: the engines are not reading the same pages, but they are arriving at the same shortlist. Overlap on brands runs more than thirty times higher than overlap on sources (32% against 1%).
The practical reading (and we want to be careful here, because it is an inference, not a measurement) is that a recommendation does not appear to come from any single page an engine happened to read. If it did, four engines reading four disjoint source lists would produce four disjoint shortlists, and they do not. What the shortlist tracks is something closer to a brand's overall weight across the web: how consistently it is named, everywhere, by everyone.
That is worse news than it sounds for anyone hoping to find the one page to get onto. There is no such page. It is also why "we got mentioned in a roundup" so rarely moves the needle on its own.
One small site did get into three engines
There was one exception worth staring at. On "cheaper alternatives to HubSpot," a site called Searchlab was cited by ChatGPT, Perplexity and Gemini. Established sites reach three or four engines on the popular questions (HubSpot on "best CRM," Zapier and project-management.com on "best project management software"), but Searchlab is the only site with no established footprint to do it. A second small site, AI Topia, made it into two.
Neither is a major publisher. Both are small sites running detailed, honest, specific comparison lists on a narrow topic. The big names that dominate individual engines (TechRadar, Forbes, G2) did not manage the same trick on that question.
We are not going to build a rule out of one query. But it is the clearest counter-example we have to the assumption that cross-engine reach requires a big masthead, and it is the kind of thing we will be testing deliberately in the next round.
What this corrects
Three things in our own earlier reporting need adjusting, and we would rather say so than quietly restate them:
We overstated a range, and the overstatement travelled. An earlier version of this page said the other three engines cited "between four and fifteen sources each" on the comparison questions. Checked against our own table, the true range is four to thirteen: the highest single count is Perplexity's 13 on "HubSpot vs Salesforce." The sentence has been replaced with the full count table below, because summarising counts as a range is what produced the error in the first place. The wrong figure also went out in a pitch email and a newsletter submission on 9 August, where we cannot correct it.
"Not one domain was cited by all three engines" was true for June's six questions, and is not true in general. In this round, two domains (zapier.com and project-management.com) were cited by all four engines on the project-management question, and Searchlab reached three engines on another. The honest version is: shared sources are rare, not absent.
We had no brand data before this round. When this study was cited elsewhere, the specific brands named in that citation were not ours. They could not have been, because we had not recorded any. Now we have, and the trio people tend to assume does not hold up. Across the seven questions, all four engines named HubSpot together twice and Zoho twice; they never once all named Salesforce. The brand all four agreed on more often than any other, three times out of seven, is Pipedrive, which is precisely the name these lists leave out.
How long does a citation last? We checked seven weeks later
Running the study twice gave us something the first round could not: the same twelve questions, the same engine, seven weeks apart. So we asked what happens to a source list over time.
Perplexity is the only engine where the comparison is clean (in June it was the one engine we put all twelve questions to), so this is a Perplexity measurement, not a claim about all four.
Across those twelve questions, Perplexity cited 80 distinct sources in June. Seven weeks later, 42 of them were still being cited. The other 48% were gone. Overlap between the pooled June and August lists was 0.29 (Jaccard). Those figures pool all twelve questions into one set per round. Question by question the turnover is higher (58% of June's source slots, pooled across the twelve questions) because a source can persist somewhere on the site while dropping off the specific question it used to answer. Both are recomputable from the two published files: sources are counted as distinct domain strings exactly as the CSV records them, with no subdomain collapsing (blog.hubspot.com and hubspot.com count separately).
But the churn is not total, and the exception is the useful part: every one of the twelve questions kept at least one source. Not a single list turned over completely. The names that survived tend to be the same kind of thing (zapier.com, thedigitalprojectmanager.com, project-management.com, techradar.com, youtube.com): deep topic-specific resources and large reference sites.
The spread is worth noting too. The most volatile question was "best AI writing tools," where 90% of the June sources were replaced, a category that genuinely churns. The most stable was "HubSpot vs Salesforce" at 25%: a mature comparison whose sources have settled.
Two things follow, and they pull in opposite directions:
- Being cited once is a snapshot, not a status. Half the list turned over in two months. If you earned a citation in the spring and have not checked since, you may well have lost it without anything visibly changing on your end.
- There is a durable core, and it is reachable. Roughly half the sources persisted, and they are not all major publishers. Being the thorough, specific resource on a narrow question is what appears to keep a source in rotation.
How to get cited when every engine reads something different
The headline finding has a direct, slightly uncomfortable implication: "get visible in AI" is not a single goal you can complete. You are running four campaigns, not one. Here is how we would prioritize them:
- Stop optimizing for "AI" in the abstract. Pick the engines your buyers actually use and treat them separately. Check where you stand in each. There is no shortcut, and the manual method is free; see how to measure AI search visibility. Once four engines a week stops fitting in an afternoon, the tools that log which page each engine cited are compared in best AI citation tracking tools.
- Keep doing real SEO for the AI Overview. It is the one engine that rewards classic ranking. If you are on Google's page one, you are most of the way to the AI Overview.
- Earn third-party validation, not just on-page work. Three of the four engines lean heavily on sites you do not own. A page that only says good things about itself is invisible to them. The mechanics of earning those mentions are in how to get cited by ChatGPT and our broader guide to improving brand visibility in AI search.
- Be present where people argue and compare: Reddit threads, comparison roundups, and an honest comparison page on your own site. Those are the placements that travel across the most engines.
- Accept that Gemini is the hard one. It cites little and cites unpredictably. Do not build your plan around it; treat any Gemini citation as a bonus, not a baseline.
The engines do not agree on who to trust. Until they do, your job is not to win "AI"; it is to be one of the few names that keeps showing up no matter which corner of the web a given engine happens to pull from.
Methodology
This page reports the Cross-Engine Citation Study, our own measurement of which sources ChatGPT, Perplexity, Gemini and Google's AI Overview cite when asked the same B2B-SaaS buyer questions. Every number on this page comes from that study, and both rounds are published in full under CC BY 4.0.
We ran ordinary B2B-SaaS buyer questions through four AI engines and logged every source each cited in its answer. There have been two rounds.
Read this before quoting a number from this page. The two rounds have different coverage, and mixing them will produce a number we did not measure.
Round 1 — 19 June 2026 (sources only)
The engines were not all asked the same number of questions, so the headline finding is scoped deliberately: six questions were put to ChatGPT, Perplexity and Gemini, and it is across those six that no domain was cited by all three. Google's AI Overview was sampled on three questions; where we say a source appeared in "all four" engines, we mean on those three questions specifically. Quoting this as "twelve questions across four engines" would overstate what we did.
- Questions: twelve buyer queries across four intents: recommendation ("best CRM for startups," "best email marketing platform"), comparison ("HubSpot vs Salesforce," "Mailchimp vs Klaviyo"), how-to ("how to reduce customer churn," "how to choose a CRM"), and alternatives ("Salesforce alternatives," "cheaper alternatives to HubSpot").
- Coverage: all twelve questions on Perplexity; a representative six (one to two per intent) on ChatGPT and Gemini; three (one per main intent) on Google's AI Overview. Single run per engine-question pair.
- Setup: Perplexity in default web-search mode; ChatGPT on the free tier with web search enabled (we note separately that with search off it cited no third-party sources); Gemini (Flash model); Google AI Overview as shown in standard search results.
- What we recorded: the domains/sources each engine displayed as citations or in its sources panel, in order.
Round 2 — 6–7 August 2026 (sources and brands)
Round 1 recorded which pages each engine read but not which products it recommended, which left the more interesting question unanswered. Round 2 fixed that.
- Questions: the same twelve queries, same wording, same intents.
- Coverage: all twelve questions on all four engines. Single run per pair.
- What we recorded: the brand names in each answer, in the order they appeared, plus every source, each tagged by type (vendor's own site / review media / independent listicle / Reddit / forum / YouTube / reference).
- The brand-overlap figures come from seven of the twelve questions: the open-ended recommendation and alternatives queries. We deliberately excluded the "X vs Y" comparisons, because the question itself names the brands: all four engines "agree" on Mailchimp and Klaviyo when you ask about Mailchimp versus Klaviyo, which would push the average to a meaningless 100%. We also excluded "how to reduce customer churn," where all four engines recommended no brands at all. And we excluded "how to choose a CRM," where Perplexity named no brands while the other three did; a four-way overlap cannot be computed when one engine contributes an empty set. Twelve questions minus three comparisons, minus those two, is the seven.
- Overlap is measured as Jaccard similarity: brands (or sources) named by all four engines, divided by the total distinct set named by any of them. Light normalization was applied to brand names only ("HubSpot CRM" and "HubSpot" count as one brand; "Salesforce Starter," "Salesforce Essentials" and "Salesforce Sales Cloud" count as Salesforce).
- The seven-week churn figure compares June's Perplexity source lists with August's, question by question, and is Perplexity-only because June's full twelve-question coverage exists for that engine alone. Several June panels were truncated, so it measures churn among visible sources.
- One behavioural note that affects the source figures: on "HubSpot vs Salesforce," "Notion vs Asana" and "Mailchimp vs Klaviyo," ChatGPT cited nothing (it answered from memory), while the other three engines did retrieve, with two exceptions: Gemini cited nothing on "Mailchimp vs Klaviyo" or on "how to choose a CRM"; on the latter it named nine brands without a single source. ChatGPT behaved the same way on both how-to questions. Rather than summarise that as a range, here is every count:
| Question | ChatGPT | Perplexity | Gemini | AI Overview |
|---|---|---|---|---|
| HubSpot vs Salesforce | 0 | 13 | 6 | 8 |
| Notion vs Asana | 0 | 10 | 4 | 11 |
| Mailchimp vs Klaviyo | 0 | 10 | 0* | 9 |
| how to choose a CRM | 0 | 9 | 0 | 7 |
| how to reduce customer churn | 0 | 10 | 1 | 6 |
Perplexity and the AI Overview cite consistently on every one of these. ChatGPT cites on none of them. Gemini sits between the two and is the least predictable of the four, which is the same character it showed in the June round.
* Gemini's "Mailchimp vs Klaviyo" answer is truncated at source. It stops mid-sentence and never reaches the point of citing. We re-opened the stored conversation rather than assuming our own load error, and it is genuinely cut off. Read that zero as "no sources in an incomplete answer," not as a clean measurement.
Limits, stated plainly
AI answers vary between runs, and every figure here comes from a single run per engine-question pair. We report only what we observed (nothing is estimated or modeled), but a single run cannot separate a stable pattern from a lucky draw. The 32%-versus-1% gap is large enough that we doubt it is noise; the individual per-question percentages are not, and should not be treated as precise. Both rounds are one market (US/English), one sector (B2B SaaS) and one moment in time.
This is a deliberately small, hand-run study; the value is in the side-by-side comparison, not the sample size. The next round will run each question multiple times so we can say something about stability rather than just position.
Which rows you can re-open yourself. A study that asks you to verify it should say where verification runs out. Of the twelve runs per engine in round two: Gemini, all twelve have a saved conversation that re-opens, so every Gemini row can be checked against the original answer. Perplexity, two of twelve; the rest were not saved at the time. ChatGPT, none: every question ran in a temporary chat, which the method requires so that no memory carries between questions, and temporary chats leave no transcript. That is the cost of a clean session rather than an oversight, but the effect is the same: the ChatGPT rows rest on our logging alone. Google's AI Overview, none, because AI Overviews have no permanent URL and can differ between runs.
Source counts are a floor, not a total. ChatGPT and the AI Overview both hide some sources behind a "+N" badge, so what the file records is what was visible.
From the next round on, every run gets a screenshot saved at collection time. We would rather publish the gap than quietly leave it out.
Cite this study
This is original, hand-run research. You are welcome to cite or share it (CC BY 4.0). Suggested attribution:
Is My Brand in AI (2026). Cross-engine AI citation study: where ChatGPT, Perplexity, Gemini and Google get their information. https://ismybrandinai.com/do-ai-engines-cite-the-same-sources
Twelve B2B-SaaS buyer questions, four AI engines, run 19 June 2026 (sources) and again 6–7 August 2026 (sources and recommended brands). See the method above or our full methodology page.
The raw runs are published, not available on request. Every answer we logged, with the brands named in order and every source with its type code, is in these two files under CC BY 4.0:
- cross-engine-brands-sources-2026-08.csv: 48 rows, one per engine-question pair
- cross-engine-brands-sources-2026-08.json: same data with the study metadata and the known limits
- cross-engine-brands-sources-2026-06.csv: the June round, the one the seven-week churn figure is measured against
- cross-engine-brands-sources-2026-06.json: same, with that round's metadata and limits
Every percentage on this page can be recalculated from those files. If a count here does not match what you get, tell us and we will publish a correction saying what changed.