THE ARTICLE · 11 MIN
AI chatbots can write a fluent, confident answer that is simply wrong: a court case that never existed, a quotation nobody said, a reference to a paper that was never published. This is usually called hallucination. This page explains what the word means, why it happens, why nobody can give you a single error rate, and the checks that catch it. It is a snapshot as of September 2026.
What “hallucination” means
OpenAI describes it simply: “By this we mean instances where a model confidently generates an answer that isn’t true.”
Researchers often define it differently. A 2022 survey describes hallucination as “the generated content that is nonsensical or unfaithful to the provided source content”. It separates two kinds:
- Intrinsic: “The generated output that contradicts the source content.”
- Extrinsic: “The generated output that cannot be verified from the source content”. The survey adds: “Notably, the extrinsic hallucination is not always erroneous because it could be from factually correct external information”.
A later survey of large language models draws a similar line between being wrong about the world and being wrong about what you gave the model. “Factuality hallucination emphasizes the discrepancy between generated content and verifiable real-world facts”, while “faithfulness hallucination captures the divergence of generated content from user input or the lack of self-consistency within the generated content”.
In plain words: the model may be wrong about the world, wrong about your document, or contradict itself.
Not everyone likes the word
“Hallucination” is a metaphor, and several researchers think it is a misleading one.
- An opinion article in PLOS Digital Health (2023) argued: “The model is not “seeing” something that is not there, but it is making things up.” It proposed a term from psychiatry instead: “More accurate terminology is found in the psychiatric concept of confabulation”.
- A philosophy paper in Ethics and Information Technology (2024) argued that “the models are in an important way indifferent to the truth of their outputs”. It used the term “bullshit” in the technical sense given to it by the philosopher Harry Frankfurt: speech produced without regard to whether it is true.
- A research team writing in Nature (2024) used “confabulations” for a narrower group of errors — “arbitrary and incorrect generations” — and wrote: “We believe that combining these distinct mechanisms in the broad category hallucination is unhelpful.”
- A review of how the term is used found “a lack of consistency in how the term is used”.
Even OpenAI’s researchers note that the word, borrowed from human experience, is imperfect: hallucination in language models “differs fundamentally from the human perceptual experience”.
Why it happens
Rare facts cannot be predicted from patterns
Language models learn from huge amounts of text. OpenAI’s explanation starts there: pretraining is “a process of predicting the next word in huge amounts of text”. Some things follow patterns, such as spelling. Others do not: “But arbitrary low-frequency facts, like a pet’s birthday, cannot be predicted from patterns alone and hence lead to hallucinations.”
A 2023 survey makes the same point about knowledge in general: models struggle to “memorize all factual knowledge encountered during pre-training, especially the less frequent long-tail knowledge”, and training data “does not include rapidly evolving world knowledge or content restricted by copyright laws”.
Tests reward guessing
A paper by researchers at OpenAI and Georgia Tech makes a sharper argument. It was first posted in September 2025 as “Why Language Models Hallucinate” and published in Nature in April 2026 as “Evaluating large language models for accuracy incentivizes hallucinations”. The published version says that “dominant headline metrics such as accuracy systematically reward guessing over admitting uncertainty”. The 2025 preprint put it more bluntly: “language models are optimized to be good test-takers, and guessing when uncertain improves test performance.”
The comparison is a multiple-choice exam with no penalty for wrong answers. If a model does not know someone’s birthday and says “I don’t know”, it scores zero. OpenAI’s summary: “If it guesses “September 10,” it has a 1-in-365 chance of being right.” The 2025 preprint states it formally: “Under binary grading, abstaining is strictly sub-optimal.”
One detail is easy to miss: rare facts set a floor. As an illustration, the published paper says that if 20% of birthday facts appear only once in the training data, then “pretrained models should hallucinate on at least 20% of birthday facts”. Pretrained models are models before the extra training that turns them into assistants.
This is the authors’ argument, from researchers mostly employed by one AI company, not a settled consensus.
Disputed Whether hallucination can be avoided at all is argued both ways. OpenAI’s researchers say it can in principle: responding to the claim that hallucinations are inevitable, they write “They are not, because language models can abstain when uncertain.” Their preprint acknowledges that “Many have argued that hallucinations are inevitable”. A 2024 paper, “Hallucination is Inevitable”, argues that under its formal definition “it is impossible to eliminate hallucination in LLMs”. The two sides partly define the problem differently: one counts saying “I don’t know” as a way out.
How often does it happen? It depends on the test
There is no single “hallucination rate”. Every published figure is a rate on a particular task, graded in a particular way — and figures from different tests cannot be compared. Four examples:
| Test | What it measures | Result |
|---|---|---|
| SimpleQA (November 2024) | Short fact questions; “Each answer in SimpleQA is graded as either correct, incorrect, or not attempted.” | GPT-4o answered almost every question and was wrong 60.8% of the time; Claude 3.5 Sonnet declined 35.0% and was wrong 36.1%. The questions were “adversarially collected against GPT-4 responses”. |
| SimpleQA, newer models (OpenAI, 2025) | The same test, as reported for two OpenAI models | gpt-5-thinking-mini declined 52% and was wrong 26%; OpenAI o4-mini declined 1% and was wrong 75% — with accuracy of 22% and 24%. |
| PersonQA (OpenAI o3 system card, April 2025) | OpenAI’s questions about “publicly available facts about people” | Hallucination rate of 0.33 for o3 against 0.16 for the older o1, while accuracy was 0.59 against 0.47. The card said: “More research is needed to understand the cause of these results.” |
| Vectara leaderboard (updated May 2026) | “This evaluates how often an LLM introduces hallucinations when summarizing a document.” | From 1.8% to 24.2% across the models listed, as judged by Vectara’s own evaluation model. |
The second SimpleQA row shows the pattern the “reward guessing” argument predicts: the model that declined far more often made about a third as many errors at almost the same accuracy (22% against 24%). In the first row, the model that declined more often also got fewer answers right (28.9% against 38.2%), though these are two different models, so the table alone does not show why. The PersonQA row shows that a newer model can hallucinate more: o3 answered more questions correctly than o1 but also made more false claims, which the card linked to o3 making “more claims overall”.
When it has mattered
A court filing with invented cases
In Mata v. Avianca, a US federal court in New York found in June 2023 that lawyers “submitted non-existent judicial opinions with fake quotes and citations created by the artificial intelligence tool ChatGPT, then continued to stand by the fake opinions after judicial orders called their existence into question”. When one lawyer asked the chatbot whether the cases were real, “ChatGPT responded that it had supplied “real” authorities that could be found through Westlaw, LexisNexis and the Federal Reporter.”
The court was clear that the tool itself was not the problem: “Technological advances are commonplace and there is nothing inherently improper about using a reliable artificial intelligence tool for assistance. But existing rules impose a gatekeeping role on attorneys to ensure the accuracy of their filings.” Finding bad faith based on “acts of conscious avoidance and false and misleading statements to the Court”, it imposed “A penalty of $5,000”. Our article on spotting fake quotes and invented sources covers the case in more detail.
It was not a one-off. A public database of decisions by courts and tribunals involving AI-generated false material listed “2041 cases identified so far” when it was updated on 14 September 2026 — by one researcher’s count, and it notes that it “does not track the (necessarily wider) universe of all fake citations or use of AI in court filings.”
References to papers that do not exist
A 2023 study in Scientific Reports checked citations produced by ChatGPT and found that “55% of the GPT-3.5 citations but just 18% of the GPT-4 citations are fabricated”. It added: “Even with GPT-4, however, 70% of the cited book chapters are fabricated.” The most useful finding for readers is how convincing the fakes were: “most of the fabricated article, book, and website citations include the names of real journals, publishers, and organizations”. A real journal name is not evidence that a paper exists.
A 2024 study on references for systematic reviews found similar problems: “Hallucination rates stood at 39.6% (55/139) for GPT-3.5, 28.6% (34/119) for GPT-4, and 91.4% (95/104) for Bard”, using its own rule for what counted as a hallucinated paper.
A customer-service chatbot
In Moffatt v. Air Canada (British Columbia Civil Resolution Tribunal, February 2024), a chatbot on the airline’s website told a customer they could apply for a bereavement fare after travelling. The airline’s own policy page said otherwise. The tribunal held the airline responsible: “It makes no difference whether the information comes from a static page or a chatbot.” It found that “Air Canada did not take reasonable care to ensure its chatbot was accurate” and awarded “$650.88 in damages”. The decision does not say what technology the chatbot used, so it is an example of an automated answer being wrong, not proof of a language-model hallucination.
What reduces it — and the limits
Looking things up first. The technique usually called retrieval-augmented generation (RAG) has the model search a collection of documents before answering. Its original 2020 paper reported: “Qualitatively, we find that RAG models hallucinate less and generate factually correct text more often than BART”, the same kind of model without retrieval. But retrieval “can be easily impacted by irrelevant retrievals”, and the authors of the “Why Language Models Hallucinate” preprint note that scoring “still rewards guessing whenever search fails to yield a confident answer”.
Specialist tools still err. A study (2024 preprint; Journal of Empirical Legal Studies, 2025) tested commercial legal research tools from providers that had described their methods as “eliminating” or “avoid[ing]” hallucinations, or had promised “hallucination-free” citations. It found they “each hallucinate between 17% and 33% of the time” — fewer errors than a general chatbot, “While hallucinations are reduced relative to general-purpose chatbots (GPT-4)”, but not none: “AI tools for legal research have not eliminated hallucinations.”
Letting the model say “I don’t know”. Anthropic’s guidance for developers recommends: “Explicitly give Claude permission to admit uncertainty.” It also suggests requiring supporting quotes — “If it can’t find a quote, it must retract the claim” — and adds its own caveat: the techniques reduce hallucinations but “they don’t eliminate them entirely”.
Asking more than once. Answers that change each time can be a warning sign. But the Nature team noted that its detection method “does not guarantee factuality because it does not help when LLM outputs are systematically bad” — a model can repeat the same wrong answer every time.
Changing how tests are scored. OpenAI’s researchers propose to “Penalize confident errors more than you penalize uncertainty, and give partial credit for appropriate expressions of uncertainty.” That is a proposal, not yet a standard.
Seven checks that catch it
- Open every citation yourself. Search for the paper, case or book in a library catalogue, a court database or the publisher’s site. Do not ask the chatbot whether it is real — in Mata v. Avianca, it said yes.
- Check quotations against the source. If you cannot find the words in the original, do not use them.
- Take extra care with rare, specific facts: birthdays, obscure names, one-off details — the kind of rare, specific detail the research points to.
- Be careful with anything recent. A model’s built-in knowledge stops at its training cutoff unless it searches.
- Ask again, but do not treat agreement as proof. Changing answers are a red flag; consistent answers can still be consistently wrong.
- Prefer “I don’t know” to a confident guess — and tell the tool it is allowed to say it.
- If an organisation’s chatbot tells you something important, check the organisation’s written policy and keep a copy of what the chatbot said.
Our reading: hallucination is not a rare glitch that is likely to vanish with the next model. It follows from how these systems are trained and tested, and even the tools that search and cite still get things wrong. The practical answer is the one we try to apply everywhere: a claim is only as good as the source you can open and read yourself.
Sources
- OpenAI, “Why language models hallucinate” (5 September 2025); Kalai, Nachum, Vempala and Zhang, “Evaluating large language models for accuracy incentivizes hallucinations”, Nature (22 April 2026); preprint “Why Language Models Hallucinate”, arXiv:2509.04664.
- “Hallucination is Inevitable: An Innate Limitation of Large Language Models”, arXiv:2401.11817.
- Ji et al., “Survey of Hallucination in Natural Language Generation”, ACM Computing Surveys; arXiv:2202.03629.
- Huang et al., “A Survey on Hallucination in Large Language Models”, arXiv:2311.05232.
- Smith, Greaves and Panch, “Hallucination or Confabulation? Neuroanatomy as metaphor in Large Language Models”, PLOS Digital Health (2023).
- Hicks, Humphries and Slater, “ChatGPT is bullshit”, Ethics and Information Technology (2024).
- Farquhar, Kossen, Kuhn and Gal, “Detecting hallucinations in large language models using semantic entropy”, Nature (2024).
- Maleki, Padmanabhan and Dutta, “AI Hallucinations: A Misnomer Worth Clarifying”, arXiv:2401.06796.
- Wei et al., “Measuring short-form factuality in large language models” (SimpleQA), arXiv:2411.04368.
- OpenAI, “OpenAI o3 and o4-mini System Card” (April 2025).
- Vectara, Hallucination Leaderboard, GitHub (updated 11 May 2026).
- Mata v. Avianca, Inc., S.D.N.Y., Opinion and Order on Sanctions (22 June 2023).
- AI Hallucination Cases Database, Damien Charlotin (updated 14 September 2026).
- Walters and Wilder, “Fabrication and errors in the bibliographic citations generated by ChatGPT”, Scientific Reports (2023).
- Chelli et al., “Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews”, Journal of Medical Internet Research (2024).
- Moffatt v. Air Canada, 2024 BCCRT 149.
- Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”, arXiv:2005.11401.
- Magesh et al., “Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools”, arXiv:2405.20362; Journal of Empirical Legal Studies 22(2) (2025).
- Anthropic, Claude Platform Docs, “Reduce hallucinations”.
Checked September 2026.
Related: How large language models work · AI myths checked · How to spot a fake quote or an invented source
- artificial intelligence
- ai models
- hallucination
- fact check
- explainer
