Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability
Benchmarks and What They Measure
A benchmark is a fixed set of tasks, a way of asking the model and a scoring rule, and its number measures only what those three together capture, on questions like those, under that setup.
The problem
A single headline score is used to compare models and decide what to deploy, yet it says nothing by itself about which skills were tested, how the model was prompted, what the score's uncertainty is, or whether the test still discriminates.
The solution
Read a benchmark as an experiment: what population of questions it samples, how the model is called, how answers are scored, what a chance and a human score are, and whether it is saturated or contaminated.
The consequence
Benchmarks are indispensable for tracking progress and catching regressions, but a score is evidence about a specific question, not a verdict on a model's ability.
You should understand first
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Vectors
- Features, Labels and Tasks
- Loss Functions
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Probability and Distributions
- Entropy
- Softmax
- Cross-Entropy Loss
- Logistic Regression
- Evaluation Metrics for Classifiers
- Expected Value and Variance
- Sampling and Uncertainty
- Generalization, Overfitting and Underfitting
- Text as Data
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Dot Product
- Embeddings
- Attention
- Self-Attention
- Causal Masking
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- One-Hot Encoding
- Tokenization
- Building a Pretraining Dataset
- Data Leakage
- Benchmark Contamination
- Benchmarks and What They Measure
Three parts, not one number
A benchmark score looks like a property of the model. It is really the result of an experiment with three parts:
- The questions. A finite sample of tasks, chosen by someone, from some population (exam questions, GitHub issues, user prompts).
- The call. How the model is asked: prompt wording and format, number of examples in the prompt, whether chain-of-thought, tools or retrieval are allowed, sampling temperature, how many attempts.
- The scoring. Exact match, multiple-choice letter, unit tests passing, a model judge, a human rating, and whether "I don't know" counts as wrong.
Change any part and the number changes. A study of prompt formats found accuracy differences of up to 76 points for LLaMA-2-13B between formats that differ only in meaningless ways such as separators and casing (Sclar et al.). Established Two papers reporting "MMLU" may have run different experiments.
Reading a score
MMLU tests 57 subjects with four-option multiple choice; at release most models were near random chance and the largest GPT-3 reached 43.9%. Established Four options means 25% for guessing, so 43.9% is real but modest. Questions to ask of any score:
| Question | Why |
|---|---|
| What is chance? What do people score? | Places the number on a scale. |
| How many questions? | Sets the uncertainty: ±5 points at 300 questions (error bars). |
| How was the model called? | Prompts, shots, tools and attempts can move scores by many points. |
| What counts as right? | Exact match, a judge, partial credit, abstentions. |
| Could it have seen the test? | Contamination inflates scores. |
| Is it saturated? | Near the ceiling, differences stop meaning much. |
Saturation: benchmarks wear out
Kiela and colleagues documented benchmarks saturating ever faster; GLUE, introduced as beyond current methods, saturated within a year (Dynabench). Established By 2025, the authors of Humanity's Last Exam reported frontier models above 90% on MMLU, and built 2,500 expert questions, rejecting any that frontier models could already answer. Established That selection step matters when you read early scores on such a test: the questions were chosen to be failed.
Responses to saturation include harder expert-written tests (GPQA: experts 65%, skilled non-experts with the web 34%), adversarially collected data, private held-out sets, and tests drawn from fresh, real tasks.
Beyond accuracy
HELM measured seven metrics (accuracy, calibration, robustness, fairness, bias, toxicity and efficiency) for 30 models on 16 core scenarios, and found that before it, models had been evaluated on only 17.9% of those scenarios on average. Established The design lesson outlasts the leaderboard: decide what you need to know (Is it right? Does it know when it is wrong? Does it hold up under rephrasing? What does it cost?) and measure each.
Construct validity
The deepest question is whether the benchmark measures what its name says. A model that scores well on bar-exam questions has shown it can answer bar-exam questions in that format; whether it can do a lawyer's work is a separate, untested claim. Interpretation For your own use, the most valid benchmark is usually a few hundred real, anonymized inputs from your task, scored the way your users would score them.
Mini experiment
Pick a public leaderboard. For the top two models, find: the number of questions, the prompt format, whether chain-of-thought or tools were used, and the reported uncertainty. How many of these can you find? Then use the error-bars lab to decide whether their gap could be noise.
Why should I care?
As a researcher
Every claim of progress in the papers you will read rests on benchmark numbers; knowing how a benchmark is built tells you how much a two-point gain can mean.
As an engineer
Public scores are rarely about your task. The skill that transfers is building a small benchmark of your own: real inputs, a fixed calling setup, a scoring rule and a baseline.
Modern systems that depend on it
- model leaderboards
- regression tests for prompts and models
- scaling-law fits on downstream tasks
- model cards and system cards
Historical context
Before
Each research group reported its own tasks and prompts; models were rarely tested on the same things, so they could not be compared.
After
Shared benchmarks (MMLU, HELM's scenarios, GPQA, coding and agent suites) make comparisons possible, and their weaknesses (saturation, contamination, format sensitivity) are themselves measured.
Used today
Model release tables, open leaderboards, internal regression suites run on every model or prompt change, and procurement decisions between vendors.
What to remember
- A benchmark = questions (a sample) + how the model is called (prompt, shots, tools) + how answers are scored.
- Always compare against chance (25% for four-way multiple choice) and, if available, a human or expert score.
- MMLU (2020): 57 subjects; most models then were near chance, GPT-3 reached 43.9%. Frontier models later passed 90%.
- Saturation: once top models cluster near the ceiling, the benchmark stops separating them.
- Contamination: test questions in the training data turn a score into partly a memory test (Chapter 9).
- HELM's point: accuracy is one of several things to measure (calibration, robustness, fairness, bias, toxicity, efficiency).
- Construct validity: does doing well on these questions mean the ability you care about?
Key papers
Measuring Massive Multitask Language Understanding
Dan Hendrycks, Collin Burns et al. · 2020
MMLU became the most quoted single number for language-model knowledge for several years, and its rise from near chance to above 90% is the textbook case of a benchmark saturating.
How to read it: Look at the calibration section as well as the accuracy table: the authors already noticed that GPT-3's confidence did not track its accuracy.
Holistic Evaluation of Language Models
Percy Liang, Rishi Bommasani et al. · 2022
HELM argued that a single accuracy hides most of what matters and evaluated 30 models on the same scenarios with seven metrics each.
How to read it: The paper is very long; read the introduction's figure of scenarios × metrics and the summary of 25 findings, then dip into a scenario you care about.
GPQA: A Graduate-Level Google-Proof Q&A Benchmark
David Rein, Betty Li Hou et al. · 2023
A response to saturation: questions hard even for skilled people with the web, built to study how humans can supervise systems that know more than they do.
Dynabench: Rethinking Benchmarking in NLP
Douwe Kiela, Max Bartolo et al. · 2021
Documented how fast benchmarks saturate and proposed collecting test data adversarially, against the current best models.
Humanity's Last Exam
Long Phan, Alice Gatti et al. · 2025
A large, expert-written benchmark built explicitly because MMLU-style tests had saturated, and one that reports calibration error alongside accuracy.
How to read it: Note the selection effect: questions were kept only if models of the time failed them, so early low scores partly reflect how the test was built.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 12: Evaluation
Percy Liang's lecture is the best single overview of how language models are actually evaluated, and of why 'which benchmark?' is a question about your goal, not a lookup.