Skip to content
Road to Intelligence

Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability

Human Preference Evaluation and Arenas

Must knowUnderstand10 minDifficulty

For open-ended tasks with no single right answer, people compare two answers side by side, and a Bradley–Terry model turns many such votes into ratings with uncertainty.

The problem

Chat answers, summaries and code explanations have no answer key, absolute 1–10 ratings are inconsistent between raters, and fixed test sets miss how people really use models.

The solution

Ask raters which of two anonymous answers is better, which people do more consistently than scoring one answer, and fit a model of win probability from rating differences across many comparisons.

The consequence

Arenas and side-by-side studies measure what users prefer on their own prompts, including style; they are only as representative as the raters and prompts, and preference is not the same as correctness.

Compare, don't score

Ask ten people to rate an answer from 1 to 10 and you get ten scales. Ask them which of two answers is better and they agree far more often. This is the same reason preference data for RLHF is collected as comparisons.

A rating model turns comparisons into a leaderboard. The Bradley–Terry model gives each model a strength θ\theta and says

P(i beats j)=σ(θi−θj)=11+e−(θi−θj).P(i \text{ beats } j) = \sigma(\theta_i - \theta_j) = \frac{1}{1 + e^{-(\theta_i - \theta_j)}} .

Fitting θ\theta is a logistic regression on the votes. Only differences matter, so one model is pinned at zero, and scores are often rescaled to look like chess Elo ratings. The reward models of Chapter 10 use the same form for answers instead of models. Established

Arenas

Chatbot Arena lets users chat with two anonymous models and vote for the better answer; ratings are Bradley–Terry coefficients with confidence intervals, and by March 2024 it had collected over 240K votes from about 90K users. Established Its strengths: real prompts that nobody can train on in advance, and a live ranking. Its limits follow from the same design. The voters are self-selected, prompts skew toward what curious visitors ask, and a vote rewards whatever the voter notices, which includes length, formatting and tone.

Tiny example

Three models, 100 votes each way. Model X beats Y 60 times out of 100, so θX−θY=logit⁡(0.6)=0.41\theta_X - \theta_Y = \operatorname{logit}(0.6) = 0.41. Y beats Z 60 times: θY−θZ=0.41\theta_Y - \theta_Z = 0.41. The model then predicts X beats Z with σ(0.82)=0.69\sigma(0.82) = 0.69, without any X–Z votes. That transitivity is an assumption; if the real X–Z record is far from 69%, the single-number ranking is hiding something (for example, Z is better at code and X at writing).

Win rate against a reference

A cheaper variant compares each model only with one fixed reference answer per prompt, and reports the share of prompts it wins. That is how AlpacaEval works, and how the judge lab is built. It needs fewer comparisons, but the number depends on the reference: everyone beats a weak reference.

Mini experiment

Write down three prompts from your own work. For each, imagine two answers: one correct but terse, one long, friendly and subtly wrong. Which would a quick voter pick? What would you add to the rating instructions to change that?

What to remember

  • Pairwise: 'Which is better, A or B?' is easier to answer consistently than 'Rate this from 1 to 10'.
  • Bradley–Terry: P(i beats j) = σ(θᵢ − θⱼ). Fit θ from many votes; report intervals.
  • Chatbot Arena: anonymous models, the user's own prompt, a vote; over 240K votes by early 2024.
  • Win rate against a fixed reference answer is another common pairwise metric (AlpacaEval).
  • Preferences reward style (length, formatting, confidence) as well as substance.
  • The ranking describes the people and prompts that produced the votes.

Key papers

Essential

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Lianmin Zheng, Wei-Lin Chiang et al. · 2023

Made 'LLM-as-a-judge' a standard method, and in the same paper measured the biases that make it risky.

How to read it: Section 3.3's limitations and Table 2 (position bias) are the parts to remember; the agreement numbers in Section 4 are averages over easy and hard comparisons.

~35 min readarXiv:2306.05685✓ verified 2026-10-07
Important

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Wei-Lin Chiang, Lianmin Zheng et al. · 2024

Turned anonymous side-by-side votes from the public into a leaderboard with confidence intervals, the best-known human-preference evaluation of chat models.

How to read it: Read the statistics section for how Bradley–Terry scores and their intervals are computed, then ask what population of users and prompts the ranking represents.

~35 min readarXiv:2403.04132✓ verified 2026-10-07

Watch