Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability
Human Preference Evaluation and Arenas
For open-ended tasks with no single right answer, people compare two answers side by side, and a Bradley–Terry model turns many such votes into ratings with uncertainty.
The problem
Chat answers, summaries and code explanations have no answer key, absolute 1–10 ratings are inconsistent between raters, and fixed test sets miss how people really use models.
The solution
Ask raters which of two anonymous answers is better, which people do more consistently than scoring one answer, and fit a model of win probability from rating differences across many comparisons.
The consequence
Arenas and side-by-side studies measure what users prefer on their own prompts, including style; they are only as representative as the raters and prompts, and preference is not the same as correctness.
You should understand first
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Vectors
- Features, Labels and Tasks
- Loss Functions
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Probability and Distributions
- Entropy
- Softmax
- Cross-Entropy Loss
- Logistic Regression
- Evaluation Metrics for Classifiers
- Expected Value and Variance
- Sampling and Uncertainty
- Generalization, Overfitting and Underfitting
- Text as Data
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Dot Product
- Embeddings
- Attention
- Self-Attention
- Causal Masking
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- One-Hot Encoding
- Tokenization
- Building a Pretraining Dataset
- Data Leakage
- Benchmark Contamination
- Benchmarks and What They Measure
- Decoding: Greedy, Temperature, Top-k, Top-p
- Learning from Comparisons
- Human Preference Evaluation and Arenas
Compare, don't score
Ask ten people to rate an answer from 1 to 10 and you get ten scales. Ask them which of two answers is better and they agree far more often. This is the same reason preference data for RLHF is collected as comparisons.
A rating model turns comparisons into a leaderboard. The Bradley–Terry model gives each model a strength and says
Fitting is a logistic regression on the votes. Only differences matter, so one model is pinned at zero, and scores are often rescaled to look like chess Elo ratings. The reward models of Chapter 10 use the same form for answers instead of models. Established
Arenas
Chatbot Arena lets users chat with two anonymous models and vote for the better answer; ratings are Bradley–Terry coefficients with confidence intervals, and by March 2024 it had collected over 240K votes from about 90K users. Established Its strengths: real prompts that nobody can train on in advance, and a live ranking. Its limits follow from the same design. The voters are self-selected, prompts skew toward what curious visitors ask, and a vote rewards whatever the voter notices, which includes length, formatting and tone.
Tiny example
Three models, 100 votes each way. Model X beats Y 60 times out of 100, so . Y beats Z 60 times: . The model then predicts X beats Z with , without any X–Z votes. That transitivity is an assumption; if the real X–Z record is far from 69%, the single-number ranking is hiding something (for example, Z is better at code and X at writing).
Win rate against a reference
A cheaper variant compares each model only with one fixed reference answer per prompt, and reports the share of prompts it wins. That is how AlpacaEval works, and how the judge lab is built. It needs fewer comparisons, but the number depends on the reference: everyone beats a weak reference.
Mini experiment
Write down three prompts from your own work. For each, imagine two answers: one correct but terse, one long, friendly and subtly wrong. Which would a quick voter pick? What would you add to the rating instructions to change that?
What to remember
- Pairwise: 'Which is better, A or B?' is easier to answer consistently than 'Rate this from 1 to 10'.
- Bradley–Terry: P(i beats j) = σ(θᵢ − θⱼ). Fit θ from many votes; report intervals.
- Chatbot Arena: anonymous models, the user's own prompt, a vote; over 240K votes by early 2024.
- Win rate against a fixed reference answer is another common pairwise metric (AlpacaEval).
- Preferences reward style (length, formatting, confidence) as well as substance.
- The ranking describes the people and prompts that produced the votes.
Key papers
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Lianmin Zheng, Wei-Lin Chiang et al. · 2023
Made 'LLM-as-a-judge' a standard method, and in the same paper measured the biases that make it risky.
How to read it: Section 3.3's limitations and Table 2 (position bias) are the parts to remember; the agreement numbers in Section 4 are averages over easy and hard comparisons.
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Wei-Lin Chiang, Lianmin Zheng et al. · 2024
Turned anonymous side-by-side votes from the public into a leaderboard with confidence intervals, the best-known human-preference evaluation of chat models.
How to read it: Read the statistics section for how Bradley–Terry scores and their intervals are computed, then ask what population of users and prompts the ranking represents.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 12: Evaluation
Percy Liang's lecture is the best single overview of how language models are actually evaluated, and of why 'which benchmark?' is a question about your goal, not a lookup.