Skip to content
Road to Intelligence

Concept · Chapter 16: Evaluation, Reliability, Safety & Interpretability

LLM-as-a-Judge

Must knowKnow well14 minDifficulty

An LLM judge reads a question and one or two answers and outputs a verdict, which makes open-ended evaluation cheap and fast, but the judge has measurable biases (answer order, length, its own style) that must be tested and corrected.

The problem

Human preference ratings are the best signal for open-ended answers but are slow and expensive, so they cannot run on every prompt change or model checkpoint.

The solution

Prompt a strong model to grade or compare answers with a clear rubric, then validate it: check agreement with human ratings on a sample, judge both answer orders, control for length, and give it reference answers for checkable questions.

The consequence

Model-graded evals now power most automatic leaderboards and internal test suites, and their numbers are trustworthy only to the extent their biases were measured and removed.

A grader that never gets tired

The idea is simple: give a strong model the question, the answer (or two answers) and a rubric, and ask for a verdict. Zheng and colleagues found that GPT-4 as a judge agreed with human preferences over 80% of the time, the same level as agreement between humans. Established That makes model judges good enough to be useful, and cheap enough to run on every change.

The same paper is the best catalogue of what goes wrong.

Measured biases

Position. Judges tended to prefer the answer shown first. Asked to compare two nearly identical answers in both orders, GPT-4 gave consistent verdicts in 65% of cases, GPT-3.5 in 46% and Claude-v1 in 24%. Established Wang and colleagues showed that by choosing the order, Vicuna-13B could be made to beat ChatGPT on 66 of 80 queries with ChatGPT as the judge. Established Verbosity. In a "repetitive list" test, where an answer was padded with a rephrased copy of its own list, Claude-v1 and GPT-3.5 preferred the padded answer 91.3% of the time and GPT-4 8.7%. Established Self-preference. In the same study GPT-4 gave itself a win rate about 10% higher than humans did and Claude-v1 25% higher, but the authors could not determine whether this was a real self-enhancement bias. Interpretation Limited grading of reasoning. Judges marked wrong math answers as right even on problems they could solve when asked separately, because they were misled by the answers in front of them; giving the judge its own independently written reference answer cut failures on a small math set from 14 of 20 to 3 of 20. Established

Corrections

  • Both orders. Judge twice with the answers swapped; count a win only if it wins both times, otherwise a tie.
  • Length control. Length-controlled AlpacaEval fits a regression of the judge's preferences with a length term and reports the win rate as if both answers were the same length, which raised its correlation with Chatbot Arena rankings from 0.94 to 0.98. Established
  • References and rubrics. For checkable questions, give the judge the answer key; for open ones, a specific rubric beats "which is better?".
  • Validation. Have people rate a few hundred items and measure the judge's agreement with them, per category.

Tiny example (lab numbers)

The judge lab builds an AlpacaEval-style leaderboard: four assistants, each compared with a reference answer on 400 prompts. Atlas is the best by construction (it wins 68% against the reference by true preference) but writes short answers; Birch is second (61%) and pads its answers to about 520 words.

With the default judge (it favours the first answer and longer answers) and the candidate always shown first, Birch tops the board at 91% and Atlas gets 74%. Judging both orders removes the first-position bonus but Birch still leads, 78% to 62%. Length control without swapping restores the order but leaves every rate inflated (Atlas 83%, Dune 66% though its true rate is 45%). With both fixes the board reads Atlas 70%, Birch 59%, Cedar 55%, Dune 49%: the true order, close to the true rates.

Mini experiment

In the lab, set the position bias to 0 and the verbosity bias high, and keep "candidate first". Does judging both orders help? Now do the opposite. Which fix addresses which bias, and why can't one fix do both jobs?

Why should I care?

As a researcher

Many papers' headline results are judged by another model; you need to know which biases could produce the reported gain.

As an engineer

A judge prompt is the cheapest regression test for chat, summarization and RAG outputs, as long as you have checked it against people on your data.

Modern systems that depend on it

  • MT-Bench and AlpacaEval
  • RAG faithfulness scoring (Ragas)
  • Constitutional AI's AI feedback
  • automatic red-teaming classifiers

Historical context

Before

Open-ended outputs were scored by n-gram overlap metrics that correlated poorly with quality, or by slow human studies.

After

A judge model grades thousands of answers in minutes; papers report its agreement with humans and the corrections applied (order swapping, length control, reference answers).

Used today

Automatic chat leaderboards, eval suites in model development, RAG and agent pipelines, synthetic preference data, content classifiers.

What to remember

  • Two formats: single-answer grading (score this answer) and pairwise (which is better?).
  • Zheng et al. 2023: GPT-4 agreed with human preferences over 80% of the time, about as often as humans agreed with each other.
  • Position bias: verdicts flip when the answers swap places. GPT-4 was consistent in 65% of swapped cases in one test.
  • Verbosity bias: a padded, repetitive answer fooled Claude-v1 and GPT-3.5 in 91.3% of cases, GPT-4 in 8.7%.
  • Fixes: judge both orders, length-controlled win rates, reference answers for checkable questions, a rubric.
  • Always measure the judge's agreement with people on a sample of your own data.
  • A judge cannot reliably grade what it cannot solve, and can be misled by the answers it reads.

Key papers

Essential

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Lianmin Zheng, Wei-Lin Chiang et al. · 2023

Made 'LLM-as-a-judge' a standard method, and in the same paper measured the biases that make it risky.

How to read it: Section 3.3's limitations and Table 2 (position bias) are the parts to remember; the agreement numbers in Section 4 are averages over easy and hard comparisons.

~35 min readarXiv:2306.05685✓ verified 2026-10-07
Important

Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Yann Dubois, Balázs Galambosi et al. · 2024

Shows how to remove a known judge bias (preferring longer answers) with a simple regression, and that doing so makes the automatic leaderboard track human preferences better.

How to read it: The idea is counterfactual: 'what would the judge have preferred if the two answers were the same length?' Look at how the regression answers it.

~20 min readarXiv:2404.04475✓ verified 2026-10-07
Optional

Large Language Models are not Fair Evaluators

Peiyi Wang, Lei Li et al. · 2023

A vivid demonstration that the order of two answers in the judge's prompt can decide the verdict.

~20 min readarXiv:2305.17926✓ verified 2026-10-07

Watch