Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability
Abstention and Selective Prediction
Abstention means answering only when confidence clears a threshold and otherwise saying 'I don't know' or handing off, which trades how many questions are answered (coverage) for how often the answers are right.
The problem
A system that must answer everything is wrong on its hardest questions, and when a wrong answer costs more than no answer, answering everything is the wrong policy.
The solution
Score answers so that a wrong answer costs more than abstaining, then answer only when the (calibrated) probability of being right exceeds the break-even threshold set by those costs.
The consequence
Abstention turns calibration into fewer confident errors; it works only if confidence is calibrated and only if the evaluation or product actually rewards saying 'I don't know'.
You should understand first
- Probability and Distributions
- Softmax
- Entropy
- Loss Functions
- Cross-Entropy Loss
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Vectors
- Features, Labels and Tasks
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Logistic Regression
- Evaluation Metrics for Classifiers
- Calibration
- Abstention and Selective Prediction
When is an answer worth giving?
Score each response: +1 for a right answer, 0 for "I don't know", and for a wrong answer. If the model's probability of being right is , answering is worth on average and abstaining is worth 0. Answer only when
Kalai and colleagues propose putting such confidence targets in evaluation instructions: answer only if more than confident, since mistakes are penalized points, right answers get 1 point and "I don't know" gets 0. Established With , which is how most benchmarks grade, : any answer beats abstaining, so a model optimized for the benchmark should never abstain.
Coverage versus risk
Lowering confidence thresholds answers more questions (higher coverage) at a higher error rate among them. Plotting error against coverage gives a risk–coverage curve, a standard picture in selective prediction. Geifman and El-Yaniv showed how to pick a threshold that guarantees a target risk with high probability; for example, a 2% top-5 error on ImageNet at almost 60% coverage, with probability 99.9%. Established
Tiny example (lab numbers)
In the lab, the toy model answers 1,000 held-out questions and is right on 59.2% of them. With a wrong answer costing 3 points, the threshold is 0.75.
- The calibrated "pretrained" model answers 373 questions at that threshold, gets 48 wrong, and scores +0.18 per question. Answering everything would score −0.63.
- The overconfident "chat-tuned" model at the same threshold answers 532 and gets 101 wrong: +0.13.
- With a penalty of 9 (threshold 0.9) the overconfident model scores −0.14, worse than abstaining on everything, while the calibrated one scores +0.07. After temperature scaling, the chat-tuned model scores +0.07 too.
So abstention depends on calibration: the threshold is right only if the confidences are.
In real systems
Abstaining does not have to mean silence. A support assistant can ask a clarifying question, run another retrieval, show sources and flag uncertainty, or hand the case to a person (human in the loop). Each option has a cost, and that cost sets the threshold. Language models can be trained to predict whether they know an answer (P(IK)), and that prediction rose when relevant documents were in the context (Kadavath et al.). Established
Mini experiment
In the lab, set the penalty to 0 and drag the threshold up from 0. Does the score ever go up? Now set the penalty to 3 and find the best threshold by hand. How close is it to 0.75 for the calibrated model, and for the overconfident one before and after scaling?
What to remember
- Score: +1 right, 0 abstain, −k wrong. Answer only if P(right) > k / (1 + k).
- k = 0 (binary grading): guessing always wins, so abstention is never rewarded.
- k = 1 → threshold 0.5; k = 3 → 0.75; k = 9 → 0.9.
- Coverage = share answered; selective accuracy = accuracy on those. Plot one against the other (risk–coverage curve).
- An overconfident model following the right threshold still answers too often: calibrate first.
- In products, 'abstain' can mean asking a clarifying question, retrieving more, or escalating to a person.
Key papers
Language Models (Mostly) Know What They Know
Saurav Kadavath, Tom Conerly et al. · 2022
Evidence that a language model's probabilities carry usable information about whether it is right, which is what abstention needs.
Why Language Models Hallucinate
Adam Tauman Kalai, Ofir Nachum et al. · 2025
A clear statistical account of why hallucinations arise in pretraining and why benchmarks graded right-or-wrong keep rewarding them.
How to read it: Section 4 (how evaluations reinforce hallucination) is short and needs no maths; the pretraining bounds in Section 3 are where the theory lives.
Selective Classification for Deep Neural Networks
Yonatan Geifman, Ran El-Yaniv · 2017
The formal version of 'answer only when confident': trade coverage for risk with a guarantee.