โ† Blog Statistics Library Aug 2026 11 min read

AI Hallucination Statistics 2026: How Often AI Gets Facts Wrong

A data-led look at AI hallucination rates in 2026: benchmark leaderboards, newsroom audits, legal and medical studies and the labs' own system cards, showing how error rates swing with the task, with every stat cited.

Hero graphic reading AI Hallucination Statistics 2026, How often AI gets facts wrong and why the rate depends entirely on the task, in Tom Riley's pink and cream brand style

AI systems are now answering a meaningful share of the world's questions, and a measurable share of those answers are wrong. Not wrong in a vague, philosophical sense. Wrong in ways researchers have counted, benchmarked and published.

The numbers below come from hallucination leaderboards, newsroom audits, legal studies and the AI labs' own system cards. They tell a more complicated story than "AI lies" or "AI is fixed". Error rates depend enormously on the task, and some of the most capable models of 2025 hallucinated more than their predecessors, not less.

For marketers there is a commercial edge to all of this. If an assistant misattributes news 31% of the time, it can misdescribe your product, your pricing or your founder with the same confidence. Understanding where the failure rates sit is the starting point for doing anything about it.

45%of over 3,000 AI news answers misrepresented the content, across 18 countries and 14 languages [3]
60%+of 1,600 news citation queries were answered incorrectly by AI search engines, with Grok 3 wrong 94% of the time [2]
1.8%grounded-summary hallucination rate for the best model, versus 24.2% for the worst [1]
79%of SimpleQA questions produced a hallucinated answer from o4-mini [7]
1,847court decisions involving AI-hallucinated content logged by 6 August 2026 [10]
69% to 88%legal-query hallucination rate for general chatbots ChatGPT 3.5 and Llama 2 [5]

Key AI hallucination statistics for 2026

How often do AI models hallucinate?

It depends entirely on the task. On grounded summarisation the best models now err in under 2% of outputs. On open-ended factual recall about obscure topics, error rates for some models exceed 70%. A single "hallucination rate" for AI does not exist.

The Vectara hallucination leaderboard, updated 11 May 2026, measures how often models introduce facts not present in a document they were asked to summarise. Rates span 1.8% for the best-scoring model to 24.2% for the worst, with frontier models such as a GPT-5.4 nano variant at 3.1% and Gemini 2.5 Flash Lite at 3.3% [1].

That is the easy setting, because the model has the correct information in front of it. OpenAI's own SimpleQA benchmark, which asks short factual questions from memory, is far harsher. The o3 and o4-mini system card reports a 51% hallucination rate for o3 and 79% for o4-mini on SimpleQA [7].

Bar chart: Hallucination rate by task, recall vs grounded summary. o4-mini on SimpleQA (recall) 79%, o3 on SimpleQA (recall) 51%, Worst leaderboard model (grounded) 24.2%, Best leaderboard model (grounded) 1.8%. Source: OpenAI o3 and o4-mini system card [7]; Vectara leaderboard, May 2026 [1].

The gap between those two numbers is the single most useful thing to understand about this topic. When an AI is summarising a page it retrieved, it is mostly reliable. When it is recalling facts from training data, including facts about your brand, it is guessing far more often than its confident tone suggests.

How accurate are AI search engines at citing sources?

Badly, on the best available evidence. The Tow Center at Columbia tested eight AI search tools on 1,600 queries asking them to identify real news articles. Collectively they answered incorrectly more than 60% of the time.

The Tow Center for Digital Journalism study, published in March 2025, gave each chatbot an excerpt from a real article and asked for the headline, publisher, date and URL [2]. Perplexity was the best performer and still got 37% wrong. Grok 3 was wrong 94% of the time, and 154 of its 200 citations led to error pages.

Statistic callout: 60%+ of 1,600 news citation queries were answered incorrectly by AI search engines, with Grok 3 wrong 94% of the time [2]

The confidence problem was as striking as the error rate. ChatGPT answered all 200 prompts, misidentified 134 articles, and signalled uncertainty only 15 times. Premium tiers were in some ways worse than free ones, because they delivered definitive wrong answers rather than declining [2]. DeepSeek credited the wrong publisher in 115 of 200 responses.

For brands, misattribution is the detail to sit with. These systems do not just get facts wrong, they routinely attach real information to the wrong source, and fabricate links that look plausible. If an assistant can credit a scoop to the wrong newspaper, it can credit a competitor's product feature, award or price point to you, or yours to them.

How often do AI assistants get the news wrong?

Roughly half of AI news answers contain at least one significant problem. The largest study to date, run by 22 public service broadcasters, found 45% of responses had a significant issue and a fifth contained major accuracy failures including hallucinated details.

Statistic callout: 45% of more than 3,000 AI news answers misrepresented the content, across 18 countries and 14 languages [3]

The European Broadcasting Union and BBC coordinated journalists in 18 countries, working in 14 languages, to assess more than 3,000 responses from ChatGPT, Copilot, Gemini and Perplexity in October 2025 [3]. Beyond the 45% headline figure, 31% of responses showed serious sourcing problems and 20% had major accuracy issues.

Gemini fared worst, with significant issues in 76% of its responses, driven largely by sourcing failures [3]. The Register's coverage of the study notes that when minor problems were counted too, 81% of responses contained some form of mistake, and documents howlers such as ChatGPT stating Pope Francis was still alive weeks after his death [4].

Bar chart: AI news answers with significant sourcing inaccuracies. Gemini 72%, ChatGPT 24%, Perplexity 15%, Copilot 15%. Source: EBU and BBC, October 2025, as reported by The Register [4].
AssistantResponses with significant sourcing inaccuracies
Gemini72%
ChatGPT24%
Perplexity15%
Copilot15%

Per-assistant sourcing error rates from the EBU/BBC study, October 2025, as reported by The Register [4].

The consistency across languages and territories matters. This is not a quirk of English-language coverage or one bad product cycle. It is systemic behaviour, and there is no reason to think descriptions of companies are handled more carefully than descriptions of news events.

Do newer AI models hallucinate more or less?

Both, awkwardly. OpenAI's 2025 reasoning models hallucinated substantially more than their predecessors, then GPT-5 reversed the trend. Progress on hallucination has not been a straight line, and vendor claims deserve scrutiny.

The OpenAI system card for o3 and o4-mini, published April 2025, showed o3 hallucinating on 33% of PersonQA questions against 16% for the older o1, with o4-mini at 48% [7]. OpenAI's explanation was that o3 "makes more claims overall", producing more correct answers and more fabricated ones simultaneously.

Bar chart: PersonQA hallucination rate across OpenAI models. o4-mini 48%, o3 33%, o1 (older) 16%. Source: OpenAI o3 and o4-mini system card, April 2025 [7].

At GPT-5's launch in August 2025, OpenAI claimed responses were around 45% less likely to contain a factual error than GPT-4o, and around 80% less likely than o3 when reasoning, with measured deception rates falling from 4.8% to 2.1% on real-world chat data [9]. Those are OpenAI's own numbers, marking its own homework, but the direction matches the Vectara leaderboard, where recent frontier models cluster at the low single-digit end on grounded tasks [1].

The honest reading is that hallucination is being managed rather than solved. A capability jump can make it worse before the next round of training makes it better, which is a reason to re-test any AI-dependent workflow at every model upgrade rather than assuming improvement.

General chatbots are alarmingly unreliable on law, specialist tools are better but far from clean, and medicine shows what tight constraints can achieve. The spread runs from 88% error rates down to under 2%, on the same underlying technology.

Stanford RegLab research published in the Journal of Legal Analysis found hallucination rates of 69% for ChatGPT 3.5 and 88% for Llama 2 on verifiable questions about federal court cases, and found models often failed to correct false legal premises embedded in questions [5]. A follow-up study by Stanford HAI tested paid legal research tools on over 200 pre-registered queries: Lexis+ AI and Ask Practical Law AI produced incorrect information more than 17% of the time, Westlaw AI-Assisted Research more than 34% [6].

Medicine offers the counterpoint. A 2025 study in npj Digital Medicine had GPT-4 summarise medical consultation transcripts under an iteratively refined clinical safety protocol and recorded hallucinations in just 1.47% of 12,999 sentences, with omissions at 3.45% [11]. The authors note their best configurations produced fewer errors per note than the rates reported for human-written clinical notes.

Bar chart: Hallucination rate by setting and tool. Llama 2 legal Q&A 88%, ChatGPT 3.5 legal Q&A 69%, Westlaw AI research tool >34%, Lexis+ AI research tool >17%, GPT-4 medical summary 1.47%. Sources: Stanford RegLab [5], Stanford HAI [6], npj Digital Medicine [11].
SettingTool or modelHallucination rate
Legal Q&A, open chatbotLlama 288% [5]
Legal Q&A, open chatbotChatGPT 3.569% [5]
Legal research, specialist RAG toolWestlaw AI-Assisted Research>34% [6]
Legal research, specialist RAG toolLexis+ AI>17% [6]
Grounded document summarisationWorst leaderboard model24.2% [1]
Grounded document summarisationBest leaderboard model1.8% [1]
Medical summarisation, tight protocolGPT-41.47% of sentences [11]

The lesson is that constraint, retrieval and workflow design move the number by an order of magnitude or more. The technology is the same. The scaffolding around it is what varies.

What happens when hallucinations reach the real world?

Courts are now sanctioning people over them at scale. A public database of judicial decisions involving fabricated AI citations passed 1,800 cases in August 2026, and it only counts incidents a judge formally documented.

The database maintained by legal researcher Damien Charlotin recorded 1,847 cases as of 6 August 2026, including 1,278 in the United States and 61 in the United Kingdom [10]. Lawyers were responsible in 719 cases and self-represented litigants in 1,080. Twenty-seven cases involved judges themselves.

Statistic callout: 1,847 court decisions involving AI-hallucinated content had been logged by 6 August 2026 [10]

The catalogued failures include 1,538 instances of fabricated case citations and 504 false quotes, with documented penalties running to $15,000 fines, adverse cost orders and bar referrals [10]. This all traces back to the pattern in the benchmark data: models that rarely say "I don't know" will invent authority that looks exactly like the real thing.

For businesses the courtroom is simply the best-documented venue, because judges write everything down. The same fabrication happens silently in purchase research, comparison queries and due diligence, where nobody publishes a sanctions order when an assistant invents a returns policy or a safety certification you do not have.

Can hallucinations actually be fixed?

Not eliminated, on current evidence, but substantially reduced. The two levers that demonstrably work are grounding answers in retrieved documents and training models to abstain when uncertain. Both come with trade-offs that vendors tend to underplay.

OpenAI's September 2025 paper Why Language Models Hallucinate argues that standard benchmarks reward confident guessing, because accuracy-only scoring gives no credit for saying "I don't know" [8]. Its comparison is stark: on SimpleQA, gpt-5-thinking-mini abstained 52% of the time and kept its error rate to 26%, whilst o4-mini abstained just 1% of the time and got 75% wrong, despite marginally higher raw accuracy.

Bar chart: Abstention versus error rate on SimpleQA. o4-mini got wrong 75%, gpt-5-thinking-mini abstained 52%, gpt-5-thinking-mini got wrong 26%, o4-mini abstained 1%. Source: OpenAI, Why Language Models Hallucinate, September 2025 [8].

Retrieval helps but does not cure. The Stanford legal tools study exists precisely because vendors marketed RAG systems as hallucination-free, and the measured rates of 17% to 34% falsified that claim [6]. The medical summarisation result shows the ceiling: with retrieval, narrow scope and iterative prompt engineering, error rates can be pushed below typical human levels [11].

For marketers the practical mitigation is different, because you do not control the model. You control what it retrieves. Clear, current, machine-readable facts on your own domain, consistent third-party coverage, and periodic auditing of what each assistant actually says about your brand are the available levers, given that these systems demonstrably fill gaps with invention.

How to read these numbers

Hallucination statistics are unusually sensitive to definitions. The Vectara leaderboard measures unfaithfulness to a supplied document, SimpleQA measures wrong answers from memory, and the EBU study counts sourcing failures as well as factual ones. A model can score 3% on one measure and 50% on another without contradiction.

Vendor-reported figures, including OpenAI's GPT-5 claims, are marking their own homework and rarely arrive with full methodology. Independent audits like the Tow Center's use adversarial framings the average user may not replicate, so real-world experience can be better than the headline rate.

Sample sizes vary widely, from 200 legal queries to 3,000 news responses, and most studies are snapshots of specific model versions now superseded. Treat every figure here as dated evidence about a moving target, not a permanent property of a product.

Finally, court case counts are a floor, not a ceiling. They capture only hallucinations that were caught, litigated and written up.

What this means going into 2027

Expect the gap between grounded and ungrounded performance to keep widening. Summarisation-style hallucination is heading towards negligible for frontier models, whilst open-ended factual recall remains stubbornly unreliable, so the retrieval layer, including your website, matters more every quarter.

Abstention is becoming a design choice. OpenAI's own research pushes evaluation towards rewarding "I don't know", which should mean fewer confident fabrications but more refusals. Brands with thin, ambiguous public information may increasingly get no answer rather than a wrong one, which is its own visibility problem.

Regulated sectors will formalise tolerance levels. The legal sanctions database growing past 1,847 cases makes professional liability concrete, and insurers and regulators tend to follow documented harm [10].

Brand-fact auditing will become routine marketing hygiene. The measured error rates in news attribution, 31% serious sourcing problems in the EBU data, apply just as readily to product and company facts, and the only way to know what assistants say about you is to ask them, systematically, on a schedule [3].

And treat every model upgrade as a fresh risk event. The o1-to-o3 regression showed that newer does not automatically mean more truthful [7]. If you want to know what the AI engines currently claim about your brand, and to fix the gaps they are filling with invention, you can book a call with me.

Sources

  1. Vectara, "Hallucination Leaderboard", updated May 2026. https://github.com/vectara/hallucination-leaderboard/
  2. Columbia Journalism Review, Tow Center for Digital Journalism, "AI Search Has a Citation Problem", March 2025. https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php
  3. European Broadcasting Union / BBC, "Largest study of its kind shows AI assistants misrepresent news content 45% of the time", October 2025. https://www.ebu.ch/news/2025/10/ai-s-systemic-distortion-of-news-is-consistent-across-languages-and-territories-international-study-by-public-service-broadcaste
  4. The Register, "AI chatbots flub news nearly half the time, BBC study finds", October 2025. https://www.theregister.com/2025/10/24/bbc_probe_ai_news/
  5. Stanford RegLab, "Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models", Journal of Legal Analysis, 2024. https://reglab.stanford.edu/publications/hlarge-legal-fictions-profiling-legal-hallucinations-in-large-language-models/
  6. Stanford HAI, "AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queries", May 2024. https://hai.stanford.edu/news/ai-trial-legal-models-hallucinate-1-out-6-or-more-benchmarking-queries
  7. OpenAI, "OpenAI o3 and o4-mini System Card", April 2025. https://cdn.openai.com/pdf/2221c875-02dc-4789-800b-e7758f3722c1/o3-and-o4-mini-system-card.pdf
  8. OpenAI, "Why language models hallucinate", September 2025. https://openai.com/index/why-language-models-hallucinate/
  9. The Register, "OpenAI's GPT-5 is here with up to 80% fewer hallucinations", August 2025. https://www.theregister.com/2025/08/07/openai_gpt_5/
  10. Damien Charlotin, "AI Hallucination Cases Database", August 2026. https://www.damiencharlotin.com/hallucinations/
  11. npj Digital Medicine, "A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation", May 2025. https://www.nature.com/articles/s41746-025-01670-7

Key facts about this post

What this article is about A sourced overview of AI hallucination statistics for 2026, covering benchmark leaderboards, news audits, legal and medical studies, model-by-model changes, real-world court cases and mitigation
Type Statistics / research roundup
Author Tom Riley, AI SEO consultant in London
What an AI hallucination is A confident but false or fabricated output from an AI model, from invented facts and citations to misattributed sources
Key stat 1 AI assistants misrepresented news content in 45% of over 3,000 answers across 18 countries and 14 languages (EBU and BBC)
Key stat 2 AI search engines answered more than 60% of 1,600 news citation queries incorrectly, with Grok 3 wrong 94% of the time (Tow Center)
Key stat 3 Grounded-summary hallucination ranges from 1.8% to 24.2%, while o4-mini hallucinated on 79% of SimpleQA recall questions (Vectara; OpenAI)
Key stat 4 A public database logged 1,847 court decisions involving AI-hallucinated content by 6 August 2026 (Damien Charlotin)
Sources cited 11 (including Vectara, Tow Center, EBU and BBC, OpenAI, Stanford RegLab, Stanford HAI, npj Digital Medicine and Damien Charlotin)
Why it matters Error rates swing by task, so brands must supply clean retrievable facts and audit what assistants say about them on a schedule

Using these stats? Please credit this page with a link back to AI Hallucination Statistics 2026. It keeps research like this free.

About the author: AI SEO consultant

Written by Tom Riley, an AI SEO and AI search consultant in London. He helps brands get recommended by ChatGPT, Google AI Overviews, Perplexity and Gemini, using the same AI SEO playbook he runs on his own site. Read his author profile or connect on LinkedIn.

Want results like this for your brand?

Book a call to discuss your growth and visibility. We'll go through your SEO and AI search together. No pitch deck, no hard sell. Just what I'd do if it were my site.

๐Ÿ‘‰ Book a call with Tom