Concept · Chapter 10: From Base Model to Assistant
Learning from Comparisons
Preference data records which of two responses a judge favors for the same prompt under a stated rubric.
The problem
There are many acceptable answers, and writing the ideal response can be harder than comparing alternatives.
The solution
Generate candidate responses, compare them under explicit criteria, and retain the prompt, responses and judgment together.
The consequence
Comparisons support reward fitting or direct preference training, while carrying the judge's uncertainty and biases.
You should understand first
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Decoding: Greedy, Temperature, Top-k, Top-p
- Learning from Comparisons
Better by which rule?
Two summaries may differ in accuracy, length and style. A judge asked for the “best” answer must trade those dimensions somehow. Write the rubric first: preserve the paragraph's facts, keep its uncertainty, obey the requested length, then prefer clarity.
Record the exact prompt with both candidates. A response cannot be judged independently of the request: a concise answer can be excellent for one user and incomplete for another.
From rankings to pairs
Four candidates have unordered pairs. A full ranking can supply all six comparisons, but the pairs share answers and a prompt. Splitting some into training and the rest into test leaks much of the task. Split by prompt, and sometimes by task family, before expanding comparisons.
Human ties and disagreement are useful observations. A binary chosen/rejected dataset may encode one resolution of them; it does not make the uncertainty disappear. The lab's explicit tie label uses a 0.5 target, a teaching choice that must not be silently substituted into another dataset's format.
Two source examples
Anthropic HH-RLHF includes chosen/rejected conversations. UltraFeedback includes model responses and AI-generated feedback along several dimensions. Original multidimensional annotations and a later binarized derivative are different datasets.
Judge-generated labels are not independent factual verification. Useful checks include response-order swaps, length/style controls and review by a separate evaluator. Retain the rubric and judge version so that a future change in preference can be investigated.
What to remember
- Chosen means preferred under a judgment, not guaranteed correct.
- Four ranked responses yield six pairs, not six independent prompts.
- Keep ties and disagreement visible; blind response order where possible.
Key papers
Deep reinforcement learning from human preferences
Paul Christiano, Jan Leike et al. · 2017 · NeurIPS 2017
Showed that agents can be trained from human comparisons between behaviours rather than a hand-written reward — the foundation of RLHF.
Learning to summarize from human feedback
Nisan Stiennon, Long Ouyang et al. · 2020
Connects preference learning to generated language before instruction-following assistants.
How to read it: Read the comparison setup and human evaluation before the aggregate scores.
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Yuntao Bai, Andy Jones et al. · 2022
Makes feedback data and the helpfulness/harmlessness tension concrete.
How to read it: Inspect the data and evaluation categories; do not reduce safety to refusal rate.
UltraFeedback: Boosting Language Models with Scaled AI Feedback
Ganqu Cui, Lifan Yuan et al. · 2023
A concrete source of synthetic feedback used in open assistant research.
How to read it: Distinguish original annotations from a binarized derivative dataset.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 15: Alignment - SFT/RLHF
A technical companion to the chapter’s demonstration, preference and reinforcement-learning pipeline.