Concept · Chapter 10: From Base Model to Assistant
Reward Models
A reward model learns a scalar score that predicts which response a judge will prefer for a prompt.
The problem
A useful answer can be phrased many ways, and writing a reward function for quality is difficult.
The solution
Collect comparisons, score each prompt-response pair, and train score differences to predict the preferred response.
The consequence
A reusable judge can score new candidates, but its score remains a learned proxy for the feedback it received.
You should understand first
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Decoding: Greedy, Temperature, Top-k, Top-p
- Learning from Comparisons
- Features, Labels and Tasks
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Logistic Regression
- Reward Models
From a comparison to a number
A labeler sees two answers to the same prompt and picks one under a rubric. A reward model assigns a score to each answer. The training objective asks the score difference to agree with the label.
One common model of that judgment is
The chosen answer is , the rejected answer . Minimize the negative log of this probability. This is the logistic-regression idea from Chapter 3, applied to a difference in scores.
A tiny judge
Equal scores mean 50–50 and a loss of 0.693 nats. A score difference of 1 gives the chosen answer a 73.1% probability and loss 0.313. A difference of −1 gives 26.9% and loss 1.313. The model is penalized more for being confidently wrong.
The lab fits four free scores to your comparisons. It really minimizes a logistic objective, but cannot read new text: its four scores belong only to the four supplied responses. A language-model-based judge instead learns features that can generalize across prompts and answers.
Try it · toy model
Judge candidate summaries and fit a small reward model. Change the rubric and watch the ranking change.
A score is not a verdict
Scores 100 and 99 imply exactly the same pair probability as 1 and 0. The numbers do not mean “100 units of truth.” The judge is learning a conditional pattern of preferences, including biases and disagreement in the annotations.
Try Faithfulness first, then Confidence first. The words stay fixed; the labels change; the learned ranking changes. Optimization faithfully reproduces a bad rubric too.
What to measure
Hold out entire prompts, not merely one response from a ranked set. Inspect disagreement, length/style bias, and performance on newly generated policy responses. A reward model that selects polished hallucinations is useful evidence about the judge's weakness, not proof that the hallucinations became better answers.
Why should I care?
As a researcher
Preference prediction connects annotation noise, generalization and optimization: a good held-out classifier can still be exploited by a changing policy.
As an engineer
A high reward is not a factuality check; test the scoring model on the kinds of responses the policy actually produces.
Modern systems that depend on it
- RLHF
- candidate ranking
- rejection sampling
Historical context
Before
Human comparisons only tell us about the particular answers people reviewed.
After
A learned function predicts judgments for new prompt-response pairs, allowing repeated selection or optimization.
Used today
Reported assistant pipelines use reward models for reinforcement learning and to select high-quality candidates for later SFT.
What to remember
- Pairwise preferences constrain score differences, not an absolute truth scale.
- Adding the same constant to both scores leaves the preference probability unchanged.
- Reward-model errors become more important when another model actively maximizes the score.
Key papers
Deep reinforcement learning from human preferences
Paul Christiano, Jan Leike et al. · 2017 · NeurIPS 2017
Showed that agents can be trained from human comparisons between behaviours rather than a hand-written reward — the foundation of RLHF.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu et al. · 2022 · NeurIPS 2022
InstructGPT: the supervised fine-tuning + reward model + RL recipe that turned GPT-3 into an instruction-following assistant, and the template for ChatGPT.
How to read it: Figure 2 is the three-step RLHF pipeline you'll meet in Chapter 10.
Learning to summarize from human feedback
Nisan Stiennon, Long Ouyang et al. · 2020
Connects preference learning to generated language before instruction-following assistants.
How to read it: Read the comparison setup and human evaluation before the aggregate scores.
Scaling Laws for Reward Model Overoptimization
Leo Gao, John Schulman, Jacob Hilton · 2022
Explains why maximizing a learned judge can diverge from the intended outcome.
How to read it: Read the synthetic-gold-reward setup before interpreting the scaling curves.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 15: Alignment - SFT/RLHF
A technical companion to the chapter’s demonstration, preference and reinforcement-learning pipeline.