What Is AI Hallucination? Why Chatbots Make Things Up
Confident, fluent, and sometimes completely wrong: why AI chatbots fabricate facts, how researchers trace it inside the model, and how to catch it before it catches you.

AI hallucination is when a chatbot or other generative AI system states something false, fabricated, or unsupported by its sources as if it were a confirmed fact, often in the same confident, fluent tone it uses for things that are actually true. It is not a bug that occasional patches can quietly erase. According to a September 2025 paper by OpenAI researchers and a Georgia Tech computer scientist, hallucination is a predictable byproduct of how large language models are trained and graded, which is why it shows up across virtually every major chatbot rather than just one vendor's product.
- What it is: confident, fluent AI output that is factually wrong, fabricated, or unsupported by the source material.
- NIST's term: the U.S. National Institute of Standards and Technology calls the same phenomenon "confabulation" in its generative AI risk guidance.
- Why it happens: OpenAI researchers argue that training and benchmark scoring reward confident guessing over admitting "I don't know."
- Inside the model: Anthropic's interpretability research found a specific neural circuit that can misfire and cause Claude to invent facts about people it barely recognizes.
- Real stakes: fabricated legal citations, invented customer-service policies, and false financial claims have already triggered lawsuits, sanctions, and binding tribunal rulings.
What Is AI Hallucination?
"Hallucination" is the term the AI industry settled on for outputs that sound authoritative but are not grounded in fact, in the model's source documents, or in reality. A chatbot might invent a court case that never existed, misstate a historical date, summarize an article that was never actually loaded, or cite a scientific paper with a real-looking but fake DOI. The term is a metaphor borrowed loosely from human psychology, and some researchers prefer more literal alternatives. The U.S. National Institute of Standards and Technology, for instance, uses "confabulation" in its official generative AI risk guidance, defining it as a phenomenon in which "GAI systems generate and confidently present erroneous or false content in response to prompts," adding that such outputs can also diverge from the prompt or contradict earlier statements in the same conversation.
Whatever it is called, the core feature is the same: the system is not lying in the human sense of intending to deceive, and it is not distinguishing "I know this" from "I am guessing" the way a careful person would. It is predicting plausible-sounding text, and plausible is not the same as true.
Why Do AI Models Hallucinate?
The most detailed public explanation comes from a paper titled "Why Language Models Hallucinate," published by OpenAI researchers Adam Tauman Kalai, Ofir Nachum, and Edwin Zhang together with Georgia Tech's Santosh Vempala in September 2025 (also posted to arXiv). Their central claim is that hallucination behaves like an ordinary statistical classification error that gets baked in during pretraining, and that it then survives because of how models are evaluated afterward.
The paper's illustrative example is blunt. When asked for the birthday of one of the paper's own co-authors and told to answer only if confident, a widely used chatbot (DeepSeek-V3, tested through its official app) gave three different wrong dates across three attempts, despite being told it could decline. The authors note that if a fact like a particular person's birthday appears only once in a model's training data, roughly 20% of such facts should be expected to be hallucinated by a base model using nothing but statistical prediction, simply because there is not enough signal to pin down the right answer. The same paper found that models including DeepSeek-V3, Meta AI, and Claude 3.7 Sonnet gave inconsistent, sometimes wildly wrong answers to a task that should require no world knowledge at all: counting how many times the letter "D" appears in the word "DEEPSEEK."
The paper's second argument is about incentives, not just data. Most benchmarks and leaderboards score answers as simply right or wrong, with no credit for saying "I don't know." That turns every evaluation into something like a multiple-choice exam where guessing is free and abstaining guarantees zero points. A model that always attempts an answer will outscore a more honest model that hedges on questions it cannot actually answer, so training that chases higher benchmark scores keeps nudging models toward confident guessing. The authors argue the fix is socio-technical: rescore the mainstream benchmarks that dominate leaderboards so that confident wrong answers are penalized more than admitted uncertainty, rather than bolting on yet another narrow hallucination-only test. OpenAI's companion post on the paper notes that newer reasoning-focused models such as GPT-5 hallucinate markedly less than their predecessors, especially on tasks that involve explicit reasoning, but that the behavior has not been eliminated.
Inside the Model: How a Hallucination Actually Forms
A separate line of work, Anthropic's interpretability research described in "Tracing the Thoughts of a Large Language Model" (March 2025), looked inside a model's internals rather than just its outputs, using circuit-tracing tools on Claude 3.5 Haiku. The researchers found that refusing to answer is Claude's computational default: a circuit that pushes the model toward "I don't have enough information" is active unless something turns it off. A separate "known entity" feature can suppress that default when the model recognizes a famous name, such as basketball player Michael Jordan, letting it answer normally.
Hallucination, in this account, is what happens when that suppression misfires. If the model recognizes a name as familiar but actually knows little else about the person, the known-entity feature can still fire and wrongly switch off the "I don't know" circuit. The model then commits to answering anyway and produces something fluent but false. Anthropic's researchers demonstrated this directly: when they artificially activated the known-answer features for an obscure, little-known name in their test set, the model confidently and consistently claimed the person played chess, a detail it had no real basis for. It is a mechanistic look at the same failure OpenAI's paper describes from the outside: a system built to produce a confident-sounding answer, doing exactly that even when it shouldn't.
How Common and How Serious Is Hallucination?
There is no single, agreed-upon "hallucination rate" for AI, and claims that reduce it to one number should be treated with caution. Reported rates swing enormously depending on what is being measured: a model asked to summarize a document it was actually given tends to hallucinate far less than the same model asked an open-ended factual question from memory, and different research groups use very different test sets and scoring rules. What is consistent across the research is the direction of the risk, not a fixed magnitude: the NIST Generative AI Profile (NIST AI 600-1) lists confabulation as one of a dozen named risk categories unique to or amplified by generative AI, alongside risks like harmful bias and data privacy, and specifically flags that users tend to believe false content "often due to the confident nature of the response," which can lead them to act on it. NIST singles out domains with real consequences, warning that a confabulated summary of patient information "could cause doctors to make incorrect diagnoses and/or recommend the wrong treatments."
The same risk shows up when comparing AI systems head to head: two frontier chatbots can produce very different answers to the same factual question with equal confidence, which is one reason outlets like this one compare models side by side on benchmarks rather than taking either vendor's accuracy claims at face value. NIST's guidance also notes a subtler version of the problem: models can generate confabulated "logic" or citations that appear to justify an answer, which can make an already-wrong response look more trustworthy, not less.
Real Hallucination Examples
Hallucination stopped being a theoretical concern once it started showing up in courtrooms, newsrooms, and customer-service logs. A sample of documented incidents:
| Year | Incident | What happened |
|---|---|---|
| 2022 | Meta's Galactica model | Meta withdrew the public demo of its science-focused model within days after it generated inaccurate content, including a fabricated academic paper. |
| 2023 | Mata v. Avianca | A lawyer submitted a legal brief containing six fake case citations generated by ChatGPT; the judge dismissed the matter and sanctioned the attorneys involved. |
| 2023–2025 | Walters v. OpenAI | A radio host sued after ChatGPT falsely claimed he had embezzled funds; a court ultimately ruled for OpenAI in 2025, finding no defamation claim was established. |
| 2024 | Moffatt v. Air Canada | A Canadian tribunal held Air Canada to a bereavement-fare discount its support chatbot had invented, rejecting the airline's argument that the bot was a separate legal entity. |
| 2025 | Deloitte government reports | Consulting reports delivered to Australian and Newfoundland and Labrador government bodies were found to contain fabricated citations. |
These are not isolated flukes. A tracking database of court filings that cites AI-generated hallucinations began logging cases in April 2025 and had recorded more than 1,300 instances in legal decisions within its first year, with courts in Canada and Israel separately logging large numbers of their own filings containing fictitious citations. Academic citation studies tell a similar story on a smaller scale: one widely cited 2023 analysis of medical-literature references produced by GPT-3 found that of 178 references, 69 had incorrect or nonexistent DOIs and 28 more could not be located at all, while a separate study of GPT-3.5-generated references found 47% were entirely fabricated.
How to Spot a Hallucinated Answer
There is no foolproof tell, because hallucinated text is specifically the kind of text that is designed by the model's training process to look just as fluent and confident as accurate text. A few practical habits reduce the risk of being fooled:
- Treat specific, checkable claims as unverified until checked. Names, dates, statistics, case citations, DOIs, and quotes are exactly the details models are most likely to fabricate, because they require pinpoint recall rather than general pattern-matching.
- Ask for sources, then verify the sources exist. A fabricated citation can look completely normal; the giveaway is usually that the paper, case, or URL it points to does not actually exist or says something different.
- Be extra skeptical of long-form, open-ended answers. Both OpenAI's and NIST's research note that confabulation is more common in open-ended, long-form responses than in answers grounded in a document the model was actually given to read.
- Watch for unwarranted confidence on obscure topics. Anthropic's circuit research suggests hallucination is especially likely to strike on people, products, or events the model has only glancing familiarity with, rather than topics it is deeply trained on.
- Don't assume a correction means the model "knows" it was wrong. A model can contradict its own previous answer when pushed, which reflects regenerated plausible text, not a verified correction.
How Companies Are Trying to Reduce It
No released technique eliminates hallucination outright, but several approaches measurably reduce how often it occurs, according to the research summary compiled on Wikipedia and the primary papers above:
- Retrieval-augmented generation (RAG): giving the model relevant documents to consult at answer time, rather than relying purely on memorized training data, generally lowers hallucination rates because the model has something concrete to ground its answer in.
- Benchmark and scoring reform: OpenAI's researchers argue the highest-leverage fix is changing how mainstream leaderboards grade answers, so that confident wrong guesses score lower than an honest "I don't know," rather than rewarding guessing as they do today.
- Guardrails and confidence scoring: some deployments hard-code certain high-stakes responses or add a confidence-estimation layer, though NIST's guidance notes this is computationally expensive and users often prefer a fast, confident-sounding answer over a hedged one anyway.
- Interpretability-driven fixes: Anthropic's circuit-level findings open the door to more targeted interventions, such as detecting when a "known entity" feature has fired without real supporting knowledge behind it, though this remains research-stage rather than a shipped product feature.
Hallucination and AI Agents
The stakes rise further once a chatbot stops just talking and starts acting. An AI agent that can browse the web, edit files, or call outside tools on a person's behalf can turn a hallucinated "fact" into a hallucinated action, such as sending an email based on a fabricated price, booking something based on an invented policy, or editing code based on a library function that was never real to begin with. NIST's generative AI profile flags exactly this compounding effect, warning that automation bias, where people over-trust an automated system's output, "can exacerbate other risks of GAI, such as risks of confabulation." The more autonomy a system is given, the less a human is positioned to catch a hallucination before it has consequences, which is why agentic deployments generally need tighter grounding and verification than a simple question-and-answer chatbot.
Bottom Line
AI hallucination is not a rare glitch confined to older or lower-quality models; it is a structural consequence of how today's language models are trained and scored, present in some form across every major chatbot family. Research from OpenAI and Georgia Tech traces the behavior to benchmarks that reward confident guessing, Anthropic's interpretability work shows a specific internal circuit that can misfire to produce it, and NIST's official risk guidance treats it as one of generative AI's defining hazards rather than a footnote. Newer, reasoning-focused models have measurably cut hallucination rates, and techniques like retrieval-augmented generation help further, but no current system has eliminated the problem. Until one does, the realistic approach is to treat confident-sounding AI output the way a careful editor treats an anonymous tip: useful as a starting point, not as a verified fact on its own.
Frequently asked questions
What is AI hallucination in simple terms?
It is when an AI chatbot states something false, invented, or unsupported as if it were a verified fact, usually in the same confident tone it uses for correct information. The term covers everything from a wrong date to a fabricated court citation or a fake scientific reference.
Why do chatbots like ChatGPT make things up?
Researchers at OpenAI and Georgia Tech argue it is because training and benchmark scoring reward confident guessing over admitting uncertainty, so models learn to always produce an answer rather than say "I don't know," even when they are effectively guessing.
How common is AI hallucination?
There is no single agreed-upon rate; it varies enormously depending on the task. Models hallucinate far less when summarizing a document they were actually given than when answering open-ended factual questions from memory, which is why NIST and researchers avoid citing one universal number.
Can AI hallucination be fixed completely?
Not with current techniques. Methods like retrieval-augmented generation, benchmark scoring reform, and guardrails reduce how often it happens, but no released system has eliminated it, and OpenAI's own researchers say even its newest reasoning models still hallucinate at times.
What are some real examples of AI hallucination?
A lawyer was sanctioned after ChatGPT fabricated six legal citations in the Mata v. Avianca case, a Canadian tribunal held Air Canada to a bereavement-fare policy its chatbot invented, Meta pulled its Galactica model after days over fabricated content, and Deloitte reports delivered to government clients were found to contain fabricated citations.
How can I tell if an AI answer is hallucinated?
Treat specific, checkable details like names, dates, citations, and statistics as unverified until you confirm them, ask the model for sources and check that those sources actually exist, and be most skeptical of long, open-ended answers on obscure topics, where hallucination is most likely to occur.
Sources
- OpenAI: Why Language Models Hallucinateopenai.com
- Kalai, Nachum, Vempala, Zhang: Why Language Models Hallucinate (arXiv)arxiv.org
- Anthropic: Tracing the Thoughts of a Large Language Modelanthropic.com
- NIST AI 600-1: Artificial Intelligence Risk Management Framework - Generative AI Profiledoi.org
- Wikipedia: Hallucination (artificial intelligence)en.wikipedia.org
Theo Park runs the AI desk at Pandromeda. He follows model launches from the frontier labs and the open-weight community, tracks the assistants and developer tools built on them, and explains what each release changes on pricing, capability and safety. His reporting leans on primary sources: model cards, technical reports, API documentation and the companies' own announcements.


