Skip to content
Road to Intelligence

Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability

Abstention and Selective Prediction

Must knowUnderstand10 minDifficulty

Abstention means answering only when confidence clears a threshold and otherwise saying 'I don't know' or handing off, which trades how many questions are answered (coverage) for how often the answers are right.

The problem

A system that must answer everything is wrong on its hardest questions, and when a wrong answer costs more than no answer, answering everything is the wrong policy.

The solution

Score answers so that a wrong answer costs more than abstaining, then answer only when the (calibrated) probability of being right exceeds the break-even threshold set by those costs.

The consequence

Abstention turns calibration into fewer confident errors; it works only if confidence is calibrated and only if the evaluation or product actually rewards saying 'I don't know'.

When is an answer worth giving?

Score each response: +1 for a right answer, 0 for "I don't know", and −k-k for a wrong answer. If the model's probability of being right is pp, answering is worth p−k(1−p)p - k(1-p) on average and abstaining is worth 0. Answer only when

p>t=k1+k.p > t = \frac{k}{1+k} .

Kalai and colleagues propose putting such confidence targets in evaluation instructions: answer only if more than tt confident, since mistakes are penalized t/(1−t)t/(1-t) points, right answers get 1 point and "I don't know" gets 0. Established With k=0k = 0, which is how most benchmarks grade, t=0t = 0: any answer beats abstaining, so a model optimized for the benchmark should never abstain.

Coverage versus risk

Lowering confidence thresholds answers more questions (higher coverage) at a higher error rate among them. Plotting error against coverage gives a risk–coverage curve, a standard picture in selective prediction. Geifman and El-Yaniv showed how to pick a threshold that guarantees a target risk with high probability; for example, a 2% top-5 error on ImageNet at almost 60% coverage, with probability 99.9%. Established

Tiny example (lab numbers)

In the lab, the toy model answers 1,000 held-out questions and is right on 59.2% of them. With a wrong answer costing 3 points, the threshold is 0.75.

  • The calibrated "pretrained" model answers 373 questions at that threshold, gets 48 wrong, and scores +0.18 per question. Answering everything would score −0.63.
  • The overconfident "chat-tuned" model at the same threshold answers 532 and gets 101 wrong: +0.13.
  • With a penalty of 9 (threshold 0.9) the overconfident model scores −0.14, worse than abstaining on everything, while the calibrated one scores +0.07. After temperature scaling, the chat-tuned model scores +0.07 too.

So abstention depends on calibration: the threshold is right only if the confidences are.

In real systems

Abstaining does not have to mean silence. A support assistant can ask a clarifying question, run another retrieval, show sources and flag uncertainty, or hand the case to a person (human in the loop). Each option has a cost, and that cost sets the threshold. Language models can be trained to predict whether they know an answer (P(IK)), and that prediction rose when relevant documents were in the context (Kadavath et al.). Established

Mini experiment

In the lab, set the penalty to 0 and drag the threshold up from 0. Does the score ever go up? Now set the penalty to 3 and find the best threshold by hand. How close is it to 0.75 for the calibrated model, and for the overconfident one before and after scaling?

What to remember

  • Score: +1 right, 0 abstain, −k wrong. Answer only if P(right) > k / (1 + k).
  • k = 0 (binary grading): guessing always wins, so abstention is never rewarded.
  • k = 1 → threshold 0.5; k = 3 → 0.75; k = 9 → 0.9.
  • Coverage = share answered; selective accuracy = accuracy on those. Plot one against the other (risk–coverage curve).
  • An overconfident model following the right threshold still answers too often: calibrate first.
  • In products, 'abstain' can mean asking a clarifying question, retrieving more, or escalating to a person.

Key papers

Important

Language Models (Mostly) Know What They Know

Saurav Kadavath, Tom Conerly et al. · 2022

Evidence that a language model's probabilities carry usable information about whether it is right, which is what abstention needs.

~45 min readarXiv:2207.05221✓ verified 2026-10-07
Essential

Why Language Models Hallucinate

Adam Tauman Kalai, Ofir Nachum et al. · 2025

A clear statistical account of why hallucinations arise in pretraining and why benchmarks graded right-or-wrong keep rewarding them.

How to read it: Section 4 (how evaluations reinforce hallucination) is short and needs no maths; the pretraining bounds in Section 3 are where the theory lives.

~40 min readarXiv:2509.04664✓ verified 2026-10-07
Optional

Selective Classification for Deep Neural Networks

Yonatan Geifman, Ran El-Yaniv · 2017

The formal version of 'answer only when confident': trade coverage for risk with a guarantee.

~25 min readarXiv:1705.08500✓ verified 2026-10-07