Google Search, AI Overviews, and Gemini cite mostly different sources for the same query: 17% average overlap between AI Overviews and traditional results, 11% between the two AI surfaces. AI Overviews appear on 52% of real-user queries, favor less-popular domains than classic search, and rarely cite sites that block Google's AI crawler. Visibility in AI search has to be earned and measured per engine.
This matters because the assumption underneath most search strategies is that Google is one system: rank well and every Google surface follows. Our data shows three engines making three separate editorial decisions about who is credible on a topic. The rest of this report covers how often AI Overviews appear, how far the source lists diverge, which kinds of websites gain and lose citations, what happens when publishers block AI crawlers, and how stable these answers are from one run to the next.
What we tested
We collected results for 12,000 queries in nine categories: representative real-user queries from search logs, retail keyword queries, product-comparison and product-question variants, localized "near me" searches, debate topics, complex explainer questions, and natural questions in both question and keyword form. For every query we captured the first page of traditional results, the AI Overview when one was generated, and Gemini's response with search grounding enabled, all collected on the same days under identical device and location conditions.
Source lists were compared at the URL level with two standard measures: Jaccard similarity, which asks whether the same sources appear at all, and rank-biased overlap, which weights agreement near the top of the list more heavily. We also classified every cited domain by popularity rank and content category, and re-ran stratified samples across devices and cities to confirm the findings hold. Differences reported here passed statistical-significance testing.
How often AI Overviews appear
Rises to 90% when the query is phrased as a question. The AIO renders above the organic results every time.
Fewer than one in five cited sources also appears in the traditional results for the same query.
Across the full benchmark, 66% of queries produced an AI Overview. On representative real-user queries the rate is 52%. The variation across query types is large and systematic: complex explainer questions trigger an AIO 94% of the time, while narrow retail keyword queries sit at 18%. Informational queries are far more likely to produce one than navigational or transactional queries.
Phrasing is the strongest single factor we measured. Interrogative queries (ending in a question mark or starting with what, how, why) received an AIO 90% of the time; the keyword reformulation of the same intent received one 52% of the time. Prevalence also climbs steadily with query length. Users who type full questions, which is exactly the behavior AI chat interfaces train, see the most AI-composed answers.
Notably, product-comparison and product-question queries trigger AIOs at 88% and 92% even though plain retail queries rarely do. Generative search inserts itself into the research and consideration stage of buying, then steps back at the final purchase query.
How different are the sources?
| Engine pair | Jaccard similarity | Rank-biased overlap |
|---|---|---|
| AI Overview vs traditional search | 0.17 | 0.23 |
| Gemini vs traditional search | 0.15 | 0.21 |
| AI Overview vs Gemini | 0.11 | 0.15 |
A Jaccard of 0.17 means fewer than one in five sources is shared between the two lists. No engine pairing exceeded a rank-biased overlap of 0.27 on any query category, and the two AI surfaces agree with each other least of all, despite AI Overviews being generated by a lightweight Gemini model. Product and local queries showed the lowest similarity of any category.
The divergence is not explained by list length. The engines return similar source counts per query: 9.7 for Gemini, 9.2 for AI Overviews, 8.8 for traditional search. It also is not randomness, because the difference between engines is far larger than the run-to-run variation within any single engine. Each system retrieves differently: AI Overviews decompose the query with a lightweight model, Gemini grounds against its own retrieval stack, and the classic ranker does neither.
Which websites win and lose in AI search
We classified every cited domain by popularity (using independent web-traffic rankings) and by content category. Three patterns are consistent across the dataset.
First, generative search reaches past the head of the web. Traditional search takes 38% of its sources from the 1,000 most popular domains; AI Overviews take 36%, and Gemini 28%. The gap widens sharply at the top of the results:
Second, the biggest individual losers are community and reference platforms. Reddit shows the largest prevalence drop of any domain, cited in roughly 28 percentage points fewer queries by AI Overviews than by traditional search, and 30 points fewer by Gemini. Facebook, Amazon, Quora, and Wikipedia also lose ground. Government and education domains are cited significantly less by both generative engines, a shift with real consequences for informational-quality debates.
Third, Google favors Google. YouTube is the single biggest gainer inside AI Overviews, and Google-owned properties gain prevalence overall while nearly every other large platform declines. For brands, a YouTube presence is now a direct AI-citation asset, not just a social channel.
Blocking AI crawlers costs citations
Publishers that block Google's AI training crawler (Google-Extended) are betting their content still surfaces in search products. The data says otherwise. Websites blocking the bot are significantly less likely to be cited by AI Overviews, even though Google states that AIOs retain access to that content. The effect is starker on Gemini: more than 20 of the web's most popular publishers, spanning national news, broadcast networks, science publishers, and review platforms, received zero Gemini citations across our entire benchmark. Every one of them blocks Google-Extended in robots.txt.
Read plainly: the block is being honored, and the visibility loss is self-inflicted. With AI Overviews rendering above the organic results on half of real queries, the block-or-allow decision is no longer a licensing posture. It is a distribution decision, and it deserves the same scrutiny as any channel a brand turns off.
How consistent are AI answers?
Traditional search is stable: run the same query twice and the source lists match closely. Generative engines are measurably less stable, and they are fragile in ways traditional search is not.
| Engine | Same query, two runs | After cosmetic edits |
|---|---|---|
| Traditional search | 0.79 | ▼ 14% |
| AI Overview | 0.65 | ▼ 29% |
| Gemini | 0.47 | ▼ 4% |
Three findings stand out. Changing device type moves AI Overview sources far more than changing location does, and more than it moves traditional results; AIOs also retrieve about one more source per query on mobile than on desktop. Cosmetic query edits that leave intent untouched, like contracting "what is" to "what's", cut AI Overview source overlap by 29%, twice the effect on traditional search. And the instability propagates to the text: across the dataset, the more the retrieved sources changed, the more the generated answer changed with them, a strong and statistically significant correlation.
For measurement, one lookup is not a data point. A brand's presence in AI answers is a rate that has to be sampled across runs, phrasings, and devices before it means anything. This is why Erlin reports citation share over thousands of prompt runs rather than single checks.
High-stakes and trending queries
Debate topics, political figures, and trending events are where errors carry the most consequence, so we examined them separately.
| Behavior | AI Overview | Gemini | Traditional search |
|---|---|---|---|
| Takes a stance on debate queries | 33% | 6% | n/a (list of links) |
| Generated for political-figure queries | 94% | 97% | n/a |
| Questionable-credibility sources (political) | 11% | 15% | 10% |
One third of AI Overview summaries on debate-style queries open with an affirmative or negative position on the question rather than a neutral summary. Both generative engines also pull measurably more content from sources with questionable credibility ratings on political queries, while citing official government domains less often than traditional search does.
Trending topics show the inverse guardrail: AI Overviews appear on only 5% of queries that began trending within the past five days, then nearly triple to 13% once the topic ages past five days. The guardrail is imperfect. During our collection window, a trending sports query was answered with the wrong winner, and one of the cited sources was a satirical fan post. The response carried the same authoritative presentation as every other AI Overview.
Errors like this accumulate quietly. An engine that answers confidently from thin sources can repeat a mistake about your pricing, your product, or your people for weeks, which is why brand monitoring is now a core GEO function.
Implications for brands and publishers
- Track each engine separately. With 11 to 17% source overlap, visibility in one engine tells you nothing about the others. A single-engine view of AI search is a sampling error.
- Compete on the sub-question, not the head term.Generative engines cite niche domains where classic search cites famous ones. A specific, well-structured answer can now out-cite a top-10 domain.
- Revisit your crawler policy with data. Blocking Google-Extended measurably reduces citations in both AIOs and Gemini. If the block is deliberate, price the traffic loss; if it is a Cloudflare default, fix it.
- Invest in YouTube as a citation source. It is the biggest prevalence gainer inside AI Overviews, and video content is cited persistently regardless of age.
- Measure rates, not snapshots. Answers change across runs, devices, and phrasings. Citation share over many sampled runs is the reliable metric; a screenshot is not.
Want this analysis for your own category? Check your AI visibility to see how each engine answers the queries that matter to your brand, and whether they agree.
Frequently asked questions
Sources
- AI Overview prevalence on real-user queries (Erlin Research, 2026)
- Source overlap, AI Overviews vs traditional search (Erlin Research, 2026)
- Top-result share from top-1,000 domains (Erlin Research, 2026)
- Reddit citation drop in generative engines (Erlin Research, 2026)
- Gemini citations for publishers blocking Google-Extended (Erlin Research, 2026)
- Run-to-run source stability by engine (Erlin Research, 2026)


