AI OverviewsAI VisibilityAEOGEOGoogle SearchResearch Review

Being Cited by an AI Overview Does Not Mean It Said What You Said: How a 98,020-Claim Study Measured It

· 9 min read
Table of Contents
  1. Who did the study and how
  2. The numbers
  3. 11% does not mean 11% wrong
  4. The 60% is a different thing
  5. Not just Google
  6. What it means for site owners
  7. Limitations
  8. Sources

The short answer: an AI Overview citing your page does not mean the sentence next to your link is something you wrote. A claim-by-claim study found that roughly one sentence in nine could not be supported by the page the Overview cited. But the number is easy to overstate, so its method and limits need to be read together.

This post also untangles something else. When the study circulates online it is often paired with a figure saying AI Overviews appear on 60% of searches. Those are two different studies with entirely different denominators.

Who did the study and how

Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact, by Haofei Xu, Umar Iqbal and Jacob M. Montgomery of Washington University in St. Louis, was submitted to arXiv on 13 May 2026. It is marked "Under Review" and has not yet been peer reviewed.

The method:

  • Queries: every 24 hours, trending queries in 19 categories from US-localised Google Trends, 13 March to 21 April 2026, 55,393 queries over 40 days.
  • Environment: servers in Northern Virginia, a fresh cookie-less browser each time.
  • Sample: 7,583 AI Overviews captured, 7,491 of them verifiable.
  • Claim extraction: each Overview split into one-fact-per-sentence claims, pronouns replaced with full names so each claim stands alone, headings and duplicates removed, giving 98,020 claims.
  • Verification: each claim checked against the pages the Overview itself cited, sorted into five categories. A model (Grok 4.1 Fast Reasoning) made the judgements. On 100 human double-annotated claims the model agreed on 98, with a frequency-weighted accuracy of 95.6%.

The numbers

Category Definition Share
Clearly supported The cited page states it directly 84.6%
Vaguely supported The cited page has it, but less specifically 4.4%
Self-contradicting Different parts of the cited page disagree 1.4%
Contradicted The cited page says the opposite 2.7%
Not mentioned The claim cannot be found on the cited page 7.0%

The last three add up to 11.0%, which the paper calls inconsistent. Per Overview, the median share of supported claims is 93.33%, and 41.9% of the 7,491 Overviews had every claim supported.

The paper's own one-line conclusion: AI Overviews cite credible sources, yet nearly one claim in nine is not supported.

11% does not mean 11% wrong

This is where the number is most often overstated. The authors list several limits, and each one pushes the figure down:

  • Only 2.7% were contradicted. The biggest slice, 7.0%, is "not found in the cited-page content the study could retrieve", not "the cited page said the opposite".
  • Social platforms were excluded. The study did not retrieve Facebook, Instagram, Reddit or YouTube pages. Under the most generous assumption the authors estimate the inconsistency rate would fall from 11.0% to about 5.3%.
  • Paywalls. Content behind a paywall could not be retrieved and counts as unsupported.
  • Timing on live data. For values like weather, the study fetched the cited page later than the Overview was generated, so a changed number counts as inconsistent. The climate category was only 48.23% consistent, which the authors call a measurement artefact.
  • A single snapshot. 40 days, the US, trending queries. No other regions, personalised states or AI Mode.

The authors suggest treating 9.64% (not mentioned plus contradicted) as an upper bound, not a precise estimate.

There is noise in the other direction too. The study separately validated the claim-extraction step and found a recall of 83.23%, meaning about one fact in six that was there did not get extracted. Measurement error runs both ways.

So the more accurate statement is: on US trending queries, roughly one AI Overview claim in nine does not line up with the page it cites, and only a minority are outright contradicted. That is a real problem on its own. There is no need to turn it into "AI Overviews are 10% wrong".

The 60% is a different thing

A common line online reads: "AI Overviews now appear on 60% of US searches, and 11% of their claims are unsupported."

That sentence stitches two studies together:

60.32% 13.7%
Who measured it SEO tool vendor Advanced Web Ranking The academic study above
Where queries came from The vendor's own keyword tracking set (a secondary article says 8,000) 55,393 Google Trends queries
Method documentation Not on the tool page Fully written up in the paper
When November 2025 March to April 2026

Within the same academic study, the trigger rate was 64.7% for question queries and only 9.5% for non-questions, and by category it ranged from 3.5% in beauty and fashion to 46.1% in hobbies and leisure. The trigger rate depends heavily on which queries you pick, so "what share of searches show an AI Overview" has no single answer. Two numbers with different denominators cannot be multiplied, nor written as "11% of the 60%".

The same vendor demonstrates the point itself. In July 2024 Advanced Web Ranking published a study with a documented method: 8,000 keywords, 500 from each of 16 industries, US desktop, and at the time only 12.4% showed an AI Overview. The tool page does not say whether the 60.32% uses the same keyword set. Same vendor, a little over a year apart, from 12% to 60%: the time changed, and the query set may have too.

We call this out because we have tripped over the same kind of thing. While tracing sources for our AI visibility tools post, we found that a widely repeated "AI Visibility Score" claim did not appear in the original author's own article. Pass a number along three times and it grows meanings its author never gave it.

Not just Google

Other primary studies point the same way, but their scopes differ and should not be compared directly:

  • Medical questions: a Stanford team in Nature Communications (2025), with 800 medical questions and 58,000 statement-source pairs, found that even for GPT-4o with web search, about 30% of individual statements were not supported by the sources it cited.
  • News attribution: Columbia's Tow Center (2025) ran 1,600 queries through 8 AI search tools asking them to identify a news excerpt's original article, and collectively they answered more than 60% incorrectly. That measures finding the right source, not whether statements are supported.
  • Debate questions: Salesforce researchers' DeepTRACE (preprint) tested several generative search and deep research tools and found citation accuracy between 40% and 80%.

Google's official position is that AI Overviews only show information supported by top web results, so they generally do not hallucinate the way other language models might (2024 official blog, 2025 official document). Google has not published a specific citation-support rate.

What it means for site owners

The old worry was "AI does not cite me". This study points to a different risk: AI cites you, but hangs something you did not say under your URL. Readers see your brand and your link, and next to them a sentence you did not write.

What we can do is limited, but five things help:

  1. Write key statements as complete sentences. Use names instead of pronouns, and make each sentence understandable without context. The reasoning is intuitive: half-finished sentences and unclear references are harder for any machine summary to carry over intact. This is a common-sense judgement, and the study did not test it directly.
  2. Put dates and numbers in the same sentence. Time-sensitive values were the most misjudged in the study. "Revenue grew this quarter" is weaker than "Revenue grew 12% in Q3 2026 over the previous quarter."
  3. Keep key facts out from behind paywalls. When AI cannot read something, it does not stay silent. It pieces the answer together from elsewhere.
  4. Check manually from time to time. Pick the few questions you care about most, ask AI how it describes your pages, and see whether it matches what you wrote. Google's APIs cannot measure this, as our paid citation post found in testing.
  5. Know what a technical score measures. Our AVS measures whether a site can be read, which is a different thing from being repeated correctly. A higher score makes it easier for AI to read you, but it cannot guarantee AI gets your words right.

Limitations

This is a review of someone else's research, not our own experiment. We did not rerun the study, and we have not measured Traditional Chinese queries or AI Overviews in Taiwan. The study covers only US trending queries, so its results do not transfer directly to other markets. It is still under review, and its numbers may change.

These are lab notes, not a product commitment.

Sources

  • Xu, Iqbal, Montgomery, Measuring Google AI Overviews, arXiv 2605.14021, 2026-05-13 (under review)
  • Advanced Web Ranking, Google AI Overview tracking tool page (60.32%, dated 2025-11 per secondary citation); AI Overview Study, 2024-07-01
  • Wu et al., Nature Communications, 2025, doi:10.1038/s41467-025-58551-6
  • Jaźwińska, Chandrasekar, Tow Center for Digital Journalism, 2025-03-06
  • Venkit et al., DeepTRACE, arXiv 2509.04499
  • Google, AI Overviews: About last week (current page title: What happened with AI Overviews and next steps), 2024-05-30; AI Overviews and AI Mode in Search, 2025-05

FAQ

Do the pages cited by Google AI Overviews actually support what the Overview says?

Mostly, but not always. A May 2026 preprint from Washington University in St. Louis split 7,491 verifiable AI Overviews into 98,020 claims and checked each against the cited pages: 84.6% were clearly supported, 4.4% vaguely supported, and 11.0% inconsistent. Of that 11.0%, 7.0% were not mentioned on the cited page, 2.7% were contradicted, and 1.4% met a page that contradicted itself. The study is still under review and covers US trending queries.

Does 11% mean AI Overviews are 11% wrong?

No. The paper's category is 'not supported by the cited page', and only 2.7% were actually contradicted by it. The larger 7.0% could not be found in the cited-page content the study could retrieve, which paywalls, timing gaps on live data, or excluded social platforms can cause. The authors estimate a lower bound of about 5.3% after accounting for social platforms and suggest treating the figure as an upper bound.

Do AI Overviews appear on 60% of searches?

That is a different number. 60.32% comes from the SEO tool vendor Advanced Web Ranking's tracking of a US keyword set, and the tool page does not document its method. The academic study measured an overall trigger rate of 13.7% on Google Trends queries, 64.7% for question queries and 9.5% for non-questions. The denominators differ, so the numbers cannot be combined.

What does Google say?

Google's official position is that AI Overviews only show information supported by top web results, so they generally do not 'hallucinate' the way other language model products might. Google has not published a specific citation-support rate.

What can site owners do?

Write key statements as complete sentences that stand without context, put dates and numbers in the same sentence, keep key facts out from behind paywalls, and periodically ask AI tools how they describe your pages. A technical visibility score such as AVS measures whether AI can read your site, not whether AI repeats you correctly.

Weekly AI Automation Playbook

No fluff — just templates, SOPs, and technical breakdowns you can use right away.

Join the Solo Lab Community

Free resource packs, daily build logs, and AI agents you can talk to. A community for solo devs who build with AI.

Want to try it yourself?

UltraProbe is free and needs no sign-up. One scan tells you whether Google and AI engines can find your site.