Skip to content
Road to Intelligence

Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability

Calibration

Must knowKnow well14 minDifficulty

A model is calibrated when its confidence matches its accuracy: of all the answers it gives with 80% confidence, about 80% are right.

The problem

A confidence score is only useful for deciding when to trust, check or abstain if it means what it says, and modern networks, especially after post-training, are often overconfident.

The solution

Measure calibration with a reliability diagram and expected calibration error on held-out data, and fix systematic miscalibration with a simple recalibration such as temperature scaling fitted on a separate set.

The consequence

Calibrated confidence lets a system route uncertain cases to people, retrieve more, or say 'I don't know'; uncalibrated confidence makes those decisions wrong in a predictable direction.

What confidence should mean

Suppose a model attaches a confidence to each answer. Collect every answer it gave with confidence near 0.8. If about 80% of them are right, and the same holds at every level, the model is calibrated. Calibration says nothing about how often the model is right; it says whether its confidence is honest.

Measuring it

A reliability diagram sorts answers into bins by confidence (0–0.1, 0.1–0.2, …) and plots each bin's accuracy against its average confidence. A calibrated model's bars sit on the diagonal; bars below it mean overconfidence. Expected calibration error summarises the gaps:

ECE=∑bnbn ∣ acc(b)−conf(b) ∣.\mathrm{ECE} = \sum_b \frac{n_b}{n}\,\big|\,\mathrm{acc}(b) - \mathrm{conf}(b)\,\big| .

ECE depends on the binning and on sample size: even a perfectly calibrated model shows a few points of ECE on a thousand answers, from noise alone.

Overconfidence and its cheap fix

Guo and colleagues found that a 1998 LeNet's confidence closely matched its accuracy while a 110-layer ResNet's average confidence was substantially higher than its accuracy, and that temperature scaling, dividing the logits by a single number TT fitted on held-out data, was surprisingly effective. Established p^=σ ⁣(zT)(or softmax(z/T) for many classes).\hat p = \sigma\!\left(\frac{z}{T}\right) \quad\text{(or softmax}(z/T)\text{ for many classes)}.

T>1T > 1 softens overconfident predictions. Because dividing by a positive number keeps the order of the logits, the predicted class, and so the accuracy, does not change.

Language models

Kadavath and colleagues found larger models well calibrated on multiple-choice and true/false questions in the right format, and that models could estimate the probability that their own proposed answer was true. Established The GPT-4 technical report shows the pretrained model highly calibrated on a subset of MMLU (ECE 0.007) and the post-trained model much less so (ECE 0.074). Established Training toward answers that people prefer rewards sounding sure, which pushes stated confidence toward the extremes. Interpretation On the expert questions of Humanity's Last Exam, all models tested had RMS calibration errors above 70%. Established

For a chat model, "confidence" can come from token probabilities (when available), a number the model states, or agreement among several sampled answers. Each needs its own check.

Tiny example (lab numbers)

The lab gives a toy model 2,000 questions with known true probabilities of being right, fits on 1,000 and reports on the other 1,000. The "pretrained" model states its true probability and has an ECE of 0.031, which is just sampling noise. The "chat-tuned" model ranks questions identically but pushes its confidence toward 0 and 1: ECE 0.114, with 399 of 1,000 answers claiming over 90% confidence. Temperature scaling fits T = 2.17 and brings ECE back to 0.032. Accuracy is 59.2% in all three cases.

Mini experiment

In the lab, choose the chat-tuned model and look at the top bin of the reliability diagram: what fraction of its "over 90% sure" answers are right? Then turn on temperature scaling. Why does one number fix it here, and what kind of miscalibration could a single temperature not fix?

Why should I care?

As a researcher

Calibration is evidence about whether a model represents its own uncertainty, and post-training can damage it; it is reported by HELM and by recent hard benchmarks alongside accuracy.

As an engineer

Every threshold you set on a model score (auto-approve, escalate, abstain) assumes calibration. Check it on your data before trusting the threshold.

Modern systems that depend on it

  • abstention and selective prediction
  • confidence-based routing to humans
  • ensembles and uncertainty estimates
  • risk scores in classifiers

Historical context

Before

Softmax outputs were read as probabilities without checking, and more accurate networks were assumed to be better in every way.

After

Calibration is measured on held-out data and fixed post hoc where needed; language-model confidence is extracted and tested before it drives decisions.

Used today

Temperature scaling of classifiers, confidence thresholds in extraction and moderation systems, abstention in question answering, and calibration columns in benchmark reports.

What to remember

  • Calibrated: among answers with confidence c, the fraction right is about c.
  • Reliability diagram: bin by confidence, plot accuracy per bin against the diagonal. Below the diagonal = overconfident.
  • ECE: the average gap between confidence and accuracy, weighted by how many answers fall in each bin.
  • Guo et al. 2017: deep networks became overconfident; temperature scaling (logits ÷ T, one number fitted on held-out data) largely fixes it without changing accuracy.
  • GPT-4's report: the pretrained model was highly calibrated on MMLU (ECE 0.007); after post-training ECE was 0.074.
  • Calibration and accuracy are separate: a model can be accurate and overconfident, or weak and honest about it.

Key papers

Important

GPT-4 Technical Report

OpenAI et al. · 2023

Documented a large jump in capability, including human-level scores on many professional and academic exams, and marked the point where frontier labs stopped disclosing model size, data and training details.

How to read it: Note what the report does not contain: architecture, parameter count, data and compute are all withheld.

~45 min readarXiv:2303.08774✓ verified 2026-09-26
Important

Holistic Evaluation of Language Models

Percy Liang, Rishi Bommasani et al. · 2022

HELM argued that a single accuracy hides most of what matters and evaluated 30 models on the same scenarios with seven metrics each.

How to read it: The paper is very long; read the introduction's figure of scenarios × metrics and the summary of 25 findings, then dip into a scenario you care about.

~1 h readarXiv:2211.09110✓ verified 2026-10-07
Optional

Humanity's Last Exam

Long Phan, Alice Gatti et al. · 2025

A large, expert-written benchmark built explicitly because MMLU-style tests had saturated, and one that reports calibration error alongside accuracy.

How to read it: Note the selection effect: questions were kept only if models of the time failed them, so early low scores partly reflect how the test was built.

~20 min readarXiv:2501.14249✓ verified 2026-10-07
Essential

On Calibration of Modern Neural Networks

Chuan Guo, Geoff Pleiss et al. · 2017

Showed that more accurate deep networks had become overconfident, and that a single number, a temperature, largely fixes it.

How to read it: Figure 1 is the whole story in one picture; the rest explains which training choices (depth, batch norm, weight decay) correlate with miscalibration.

~25 min readarXiv:1706.04599✓ verified 2026-10-07
Important

Language Models (Mostly) Know What They Know

Saurav Kadavath, Tom Conerly et al. · 2022

Evidence that a language model's probabilities carry usable information about whether it is right, which is what abstention needs.

~45 min readarXiv:2207.05221✓ verified 2026-10-07