Skip to content
Road to Intelligence

Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability

Error Bars on Evals

Must knowKnow well12 minDifficulty

A benchmark score is an estimate from a sample of questions, so it has a standard error, and comparing two models on the same questions question by question (a paired comparison) separates real gaps from noise with far fewer questions.

The problem

Leaderboards print scores to a decimal place and bold the winner, but a gap of a few points on a few hundred questions is often smaller than the noise from which questions happened to be chosen.

The solution

Report a standard error with every score, compare models on the same questions using the per-question differences, account for grouped questions, and check in advance that the benchmark is big enough to detect the gap you care about.

The consequence

Many published 'wins' are within noise; with paired analysis and enough questions, small real gaps become detectable and fake ones stop being reported.

A score is an estimate

Think of a benchmark as a sample from a much larger population of questions you could have asked. Miller frames evals exactly this way and recommends reporting the standard error of the mean, computed from the central limit theorem. Established For an accuracy pp on nn questions:

SE=p(1−p)n,95% interval≈p±1.96 SE.\mathrm{SE} = \sqrt{\frac{p(1-p)}{n}}, \qquad \text{95\% interval} \approx p \pm 1.96\,\mathrm{SE}.

Tiny example. A model gets 216 of 300 right: p=0.72p = 0.72, SE=0.72×0.28/300=0.026\mathrm{SE} = \sqrt{0.72 \times 0.28 / 300} = 0.026. The interval is 72% ± 5 points. A rival at 75% on another 300 questions is not clearly better.

Compare on the same questions

When two models answer the same questions, use that. For each question compute the difference di=sB,i−sA,id_i = s_{B,i} - s_{A,i}, which is +1+1 (only B right), −1-1 (only A right) or 00 (they agree). Miller recommends inference on these question-level paired differences rather than on the two summary scores; the paired variance subtracts twice the covariance between the models' scores. Established

SEpaired=sd⁡(di)n.\mathrm{SE}_{\text{paired}} = \frac{\operatorname{sd}(d_i)}{\sqrt{n}} .

Models find the same questions hard, so their scores are positively correlated and the paired error is smaller. The intuition: questions both models get right, or both get wrong, say nothing about which is better. Only the disagreements count.

Tiny example (lab numbers)

The lab simulates two models whose true accuracies over the whole question pool are 72% and 75%. On the default 300 questions, A scores 72.0% and B 77.7%. They disagree on 53 questions: 18 only A got right, 35 only B. The unpaired 95% interval for the gap is −1.3 to +12.6 points; the paired interval is +0.9 to +10.4, which excludes zero. The measured gap of 5.7 is also nearly twice the true gap of 3: small samples overshoot as often as they undershoot.

Rerun that benchmark 200 times on fresh questions and B fails to come out ahead in 22 runs (11%); at 100 questions, in 60 runs (30%). At 1,000 questions the paired test detects the 3-point gap in 130 of 200 reruns, the unpaired test in only 54.

Other sources of noise

  • Clusters. When questions come in related groups, such as several questions about one passage, Miller recommends clustered standard errors and reports real cases where they were up to about three times the naive ones. Established
  • Sampling. With temperature above zero the same model gives different answers; average several samples per question, or use token probabilities where possible.
  • Format. Equivalent prompt formats moved one model's accuracy by up to 76 points (Sclar et al.). Established Fix the format for comparisons, and report a range across formats when claiming general ability.

Plan before you run

Power analysis turns the question around: how many questions do I need to detect a gap of size δ\delta? Roughly, you need the gap to be about three paired standard errors. If your benchmark has 200 questions and you care about 2-point differences, no amount of analysis will save the comparison; you need more questions.

Mini experiment

In the lab, set the true gap to 0 and click "Rerun 200 times" at each benchmark size. About 1 run in 20 still shows a "significant" gap at the 95% level. Now imagine 20 teams each trying a tweak and publishing only the ones that won.

What to remember

  • Standard error of an accuracy p on n questions: √(p(1−p)/n). At 72% and 300 questions that is 2.6 points, so the 95% interval is about ±5.
  • Unpaired: SE of a difference = √(SE_A² + SE_B²). Paired: the standard deviation of per-question differences divided by √n, usually smaller.
  • Only questions the two models disagree on carry information about which is better.
  • Grouped questions (several about one passage) are not independent: cluster the standard errors, which can be several times larger.
  • Prompt format and sampling add more variance; report a range or average over them.
  • Decide how many questions you need before running (power analysis).

Key papers

Essential

Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Evan Miller · 2024

A short, practical guide to treating an eval as an experiment: standard errors, clustered questions, paired comparisons and power analysis.

How to read it: Section 4 on paired differences is the most useful page for anyone comparing two models on the same benchmark.

~25 min readarXiv:2411.00640✓ verified 2026-10-07

Watch