Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability
Hallucination
A hallucination is a fluent, confident statement that is false or unsupported, and it follows from how language models are trained and graded: they learn to produce plausible text, and most tests reward a guess over 'I don't know'.
The problem
Users cannot tell a correct answer from a made-up one by its tone: invented citations, column names, API functions and dates read exactly like real ones.
The solution
Understand the causes (facts seen rarely or never, imitation of human errors, training and benchmarks that reward guessing), then reduce and catch hallucinations: ground answers in retrieved sources, grade abstention fairly, measure factuality claim by claim, and verify what can be checked.
The consequence
Hallucination is a rate to measure and reduce for each use, not a bug that has been fixed; systems are designed so that a confident wrong answer is caught before it does harm.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- Feed-Forward Sublayer (MLP)
- Parametric vs Retrieved Knowledge
- Hallucination
Plausible is not the same as true
A language model is trained to continue text the way its training data would. For a well-known fact, the most plausible continuation is the true one. For an obscure one, the most plausible continuation is something that looks like the right kind of answer: a date in the right format, a citation with real-sounding authors, a database column named the way columns usually are. The model has no separate signal that says "this one I actually know", unless training gives it one.
Three causes
Facts seen rarely. Kalai and colleagues argue that hallucinations are ordinary errors of a classification problem the model cannot solve from its data: for arbitrary facts such as birthdays, the error rate of a base model is bounded below by roughly the fraction of facts that appear exactly once in training (the singleton rate). If 20% of birthday facts appear once, expect hallucinations on at least about 20% of birthday questions. Established Facts you saw once are like facts you never saw: there is no pattern to generalize from. Interpretation
Imitating human errors. TruthfulQA asked 817 questions that some people answer falsely because of misconceptions; the best model was truthful on 58% against 94% for humans, and the largest models were generally the least truthful. Established Imitation transmits whatever the web believes.
Grading that rewards guessing. Kalai and colleagues show that under binary grading, where "I don't know" scores the same as a wrong answer, abstaining is never optimal, and most popular benchmarks grade this way. Established A model tuned to score well on such tests learns to always answer. Fine-tuning on facts the model did not already know also increased its tendency to hallucinate in a controlled study (Gekhman et al.). Established
Measuring it
- Short answers: accuracy, plus the rate of wrong answers and abstentions, reported separately.
- Long answers: FActScore breaks a generation into atomic facts and reports the share supported by a knowledge source; ChatGPT's biographies scored 58%. Established
- Grounded answers: whether each claim is supported by the retrieved passages it cites (RAG evaluation).
- Images: whether the model describes objects that are absent (multimodal evaluation).
Tiny example
A chat assistant is asked, "What does the signup_source column contain?" in a warehouse where no such column exists. Answering "It stores the marketing channel of each signup" is fluent and well-formed: the name invites it. A grader that counts only right answers scores this the same as "I can't find that column." The calibration lab puts numbers on the incentive: with binary grading, answering every question is always the best policy.
Reducing it
No single fix removes hallucination; each one lowers the rate for a class of questions. Interpretation- Ground answers in retrieved sources and show them (RAG).
- Check what is checkable: run the SQL, validate against the schema, resolve the citation.
- Reward abstention: grade "I don't know" above a wrong answer, in training and in evaluation (abstention).
- Use confidence where it is calibrated, and test that it is.
Mini experiment
Ask a chat model for three papers on a narrow topic you know well, with authors, years and venues. Check each one. Then ask again with "If you are not sure a paper exists, say so." Did the number of invented papers change? Did the number of real ones?
Why should I care?
As a researcher
Hallucination connects pretraining statistics, post-training incentives and benchmark design; the 2025 statistical account makes it a question about scoring rules as much as models.
As an engineer
Every LLM feature that states facts needs a plan for wrong answers: sources, validation, abstention and a way for users to check.
Modern systems that depend on it
- retrieval-augmented generation
- citation and grounding checks
- abstention and confidence thresholds
- factuality benchmarks
Historical context
Before
Fluent output was taken as a sign of knowledge, and 'the model said so' was treated as a source.
After
Factuality is measured claim by claim, answers are grounded and cited where possible, and evaluations increasingly give credit for honest uncertainty.
Used today
RAG with citations, FActScore-style checking of long answers, abstention policies, structured outputs validated against a schema or database, and benchmarks that penalize confident errors.
What to remember
- Hallucination = fluent and confident but false or unsupported. It is not lying; there is no intent.
- Causes: rare facts (seen once or never), imitating human misconceptions, and training or grading that rewards guessing.
- Kalai et al. 2025: for arbitrary facts, base models should hallucinate on at least about the share seen exactly once in training (the singleton rate).
- Under right-or-wrong grading, guessing always beats abstaining, so test-taking models learn to guess.
- TruthfulQA: best model truthful on 58% vs humans 94%; the largest models were generally the least truthful on these misconception questions.
- FActScore: split a long answer into atomic facts and check each; ChatGPT biographies scored 58%.
- Fixes reduce, not remove: retrieval, abstention credit, verification of checkable claims.
Key papers
Evaluating Object Hallucination in Large Vision-Language Models
Yifan Li, Yifan Du et al. · 2023
A systematic look at vision-language models describing objects that are not in the image, and POPE, a simple yes/no probe for it.
How to read it: The three negative-sampling settings are the clever part: an adversarial absent object is one that often appears alongside what is really there.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Stephanie Lin, Jacob Hilton, Owain Evans · 2021
A benchmark built around a failure that scaling can make worse: repeating popular misconceptions learned from human text.
How to read it: The inverse-scaling result is specific to questions designed around misconceptions; read it as 'imitation transmits errors', not 'bigger models are less truthful in general'.
Why Language Models Hallucinate
Adam Tauman Kalai, Ofir Nachum et al. · 2025
A clear statistical account of why hallucinations arise in pretraining and why benchmarks graded right-or-wrong keep rewarding them.
How to read it: Section 4 (how evaluations reinforce hallucination) is short and needs no maths; the pretraining bounds in Section 3 are where the theory lives.
FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
Sewon Min, Kalpesh Krishna et al. · 2023
A practical recipe for measuring factuality in long answers: split into atomic facts, check each against a source.
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
Zorik Gekhman, Gal Yona et al. · 2024
A controlled experiment linking one training choice to hallucination: teaching facts the model did not already know.