Skip to content
Road to Intelligence

Concept · Chapter 10: From Base Model to Assistant

Learning from Comparisons

Must knowKnow well10 minDifficulty

Preference data records which of two responses a judge favors for the same prompt under a stated rubric.

The problem

There are many acceptable answers, and writing the ideal response can be harder than comparing alternatives.

The solution

Generate candidate responses, compare them under explicit criteria, and retain the prompt, responses and judgment together.

The consequence

Comparisons support reward fitting or direct preference training, while carrying the judge's uncertainty and biases.

Better by which rule?

Two summaries may differ in accuracy, length and style. A judge asked for the “best” answer must trade those dimensions somehow. Write the rubric first: preserve the paragraph's facts, keep its uncertainty, obey the requested length, then prefer clarity.

Record the exact prompt with both candidates. A response cannot be judged independently of the request: a concise answer can be excellent for one user and incomplete for another.

From rankings to pairs

Four candidates have 4×3/2=64\times3/2=6 unordered pairs. A full ranking can supply all six comparisons, but the pairs share answers and a prompt. Splitting some into training and the rest into test leaks much of the task. Split by prompt, and sometimes by task family, before expanding comparisons.

Human ties and disagreement are useful observations. A binary chosen/rejected dataset may encode one resolution of them; it does not make the uncertainty disappear. The lab's explicit tie label uses a 0.5 target, a teaching choice that must not be silently substituted into another dataset's format.

Two source examples

Anthropic HH-RLHF includes chosen/rejected conversations. UltraFeedback includes model responses and AI-generated feedback along several dimensions. Original multidimensional annotations and a later binarized derivative are different datasets.

Judge-generated labels are not independent factual verification. Useful checks include response-order swaps, length/style controls and review by a separate evaluator. Retain the rubric and judge version so that a future change in preference can be investigated.

What to remember

  • Chosen means preferred under a judgment, not guaranteed correct.
  • Four ranked responses yield six pairs, not six independent prompts.
  • Keep ties and disagreement visible; blind response order where possible.

Key papers

Essential

Deep reinforcement learning from human preferences

Paul Christiano, Jan Leike et al. · 2017 · NeurIPS 2017

Showed that agents can be trained from human comparisons between behaviours rather than a hand-written reward — the foundation of RLHF.

~45 min readarXiv:1706.03741✓ verified 2026-09-26
Important

Learning to summarize from human feedback

Nisan Stiennon, Long Ouyang et al. · 2020

Connects preference learning to generated language before instruction-following assistants.

How to read it: Read the comparison setup and human evaluation before the aggregate scores.

~30 min readarXiv:2009.01325✓ verified 2026-10-04
Important

UltraFeedback: Boosting Language Models with Scaled AI Feedback

Ganqu Cui, Lifan Yuan et al. · 2023

A concrete source of synthetic feedback used in open assistant research.

How to read it: Distinguish original annotations from a binarized derivative dataset.

~25 min readarXiv:2310.01377✓ verified 2026-10-04

Watch