Skip to content
Road to Intelligence

Concept · Chapter 10: From Base Model to Assistant

Reward Models

Must knowImplement14 minDifficulty

A reward model learns a scalar score that predicts which response a judge will prefer for a prompt.

The problem

A useful answer can be phrased many ways, and writing a reward function for quality is difficult.

The solution

Collect comparisons, score each prompt-response pair, and train score differences to predict the preferred response.

The consequence

A reusable judge can score new candidates, but its score remains a learned proxy for the feedback it received.

From a comparison to a number

A labeler sees two answers to the same prompt and picks one under a rubric. A reward model assigns a score to each answer. The training objective asks the score difference to agree with the label.

One common model of that judgment is

P(yw≻yl∣x)=σ(rϕ(x,yw)−rϕ(x,yl)),σ(a)=11+e−a.P(y_w\succ y_l\mid x)=\sigma\big(r_\phi(x,y_w)-r_\phi(x,y_l)\big),\qquad \sigma(a)=\frac{1}{1+e^{-a}}.

The chosen answer is ywy_w, the rejected answer yly_l. Minimize the negative log of this probability. This is the logistic-regression idea from Chapter 3, applied to a difference in scores.

A tiny judge

Equal scores mean 50–50 and a loss of 0.693 nats. A score difference of 1 gives the chosen answer a 73.1% probability and loss 0.313. A difference of −1 gives 26.9% and loss 1.313. The model is penalized more for being confidently wrong.

The lab fits four free scores to your comparisons. It really minimizes a logistic objective, but cannot read new text: its four scores belong only to the four supplied responses. A language-model-based judge instead learns features that can generalize across prompts and answers.

Try it · toy model

Teach a Preference

Judge candidate summaries and fit a small reward model. Change the rubric and watch the ranking change.

Implement12 min

A score is not a verdict

Scores 100 and 99 imply exactly the same pair probability as 1 and 0. The numbers do not mean “100 units of truth.” The judge is learning a conditional pattern of preferences, including biases and disagreement in the annotations.

Try Faithfulness first, then Confidence first. The words stay fixed; the labels change; the learned ranking changes. Optimization faithfully reproduces a bad rubric too.

What to measure

Hold out entire prompts, not merely one response from a ranked set. Inspect disagreement, length/style bias, and performance on newly generated policy responses. A reward model that selects polished hallucinations is useful evidence about the judge's weakness, not proof that the hallucinations became better answers.

Why should I care?

As a researcher

Preference prediction connects annotation noise, generalization and optimization: a good held-out classifier can still be exploited by a changing policy.

As an engineer

A high reward is not a factuality check; test the scoring model on the kinds of responses the policy actually produces.

Modern systems that depend on it

  • RLHF
  • candidate ranking
  • rejection sampling

Historical context

Before

Human comparisons only tell us about the particular answers people reviewed.

After

A learned function predicts judgments for new prompt-response pairs, allowing repeated selection or optimization.

Used today

Reported assistant pipelines use reward models for reinforcement learning and to select high-quality candidates for later SFT.

What to remember

  • Pairwise preferences constrain score differences, not an absolute truth scale.
  • Adding the same constant to both scores leaves the preference probability unchanged.
  • Reward-model errors become more important when another model actively maximizes the score.

Key papers

Essential

Deep reinforcement learning from human preferences

Paul Christiano, Jan Leike et al. · 2017 · NeurIPS 2017

Showed that agents can be trained from human comparisons between behaviours rather than a hand-written reward — the foundation of RLHF.

~45 min readarXiv:1706.03741✓ verified 2026-09-26
Essential

Training language models to follow instructions with human feedback

Long Ouyang, Jeff Wu et al. · 2022 · NeurIPS 2022

InstructGPT: the supervised fine-tuning + reward model + RL recipe that turned GPT-3 into an instruction-following assistant, and the template for ChatGPT.

How to read it: Figure 2 is the three-step RLHF pipeline you'll meet in Chapter 10.

~1 h readarXiv:2203.02155✓ verified 2026-09-26
Important

Learning to summarize from human feedback

Nisan Stiennon, Long Ouyang et al. · 2020

Connects preference learning to generated language before instruction-following assistants.

How to read it: Read the comparison setup and human evaluation before the aggregate scores.

~30 min readarXiv:2009.01325✓ verified 2026-10-04
Important

Scaling Laws for Reward Model Overoptimization

Leo Gao, John Schulman, Jacob Hilton · 2022

Explains why maximizing a learned judge can diverge from the intended outcome.

How to read it: Read the synthetic-gold-reward setup before interpreting the scaling curves.

~30 min readarXiv:2210.10760✓ verified 2026-10-04

Watch