Skip to content
Road to Intelligence

Concept · Chapter 10: From Base Model to Assistant

Reinforcement Learning from Human Feedback

Must knowKnow well17 minDifficulty

RLHF improves a policy using a reward learned from human feedback, often with a penalty for departing from a reference model.

The problem

A fixed set of demonstrations does not tell a model which of its newly generated responses people would prefer.

The solution

Generate responses, score them with a preference-trained reward model, and optimize the policy while controlling drift.

The consequence

The model can improve beyond imitating fixed demonstrations, while also discovering ways to exploit mistakes in the reward.

Let the policy try

In Chapter 5, a policy chose actions and received rewards. For a causal language model, the state is the prompt plus tokens generated so far, and an action is the next token. A completed response can receive a score from a learned reward model. A policy-gradient algorithm uses that signal to change the probabilities of the sampled actions.

InstructGPT first trained on demonstrations, then learned a reward model from ranked responses, then used PPO for policy optimization. This is an influential recipe, not the definition of every assistant.

Two routes from demonstrations to preferences
  1. 01

    Show an answer

    Demonstrations train an SFT policy. Keep a frozen reference copy.

  2. 02

    Compare answers

    Judges rank candidates. A reward model learns to predict those comparisons.

  3. 03

    Sample and improve

    The policy generates responses. PPO uses reward and a penalty for drifting from the reference.

These are common teaching pipelines, not mandatory stages for every assistant. DPO-based systems can still use a reward model elsewhere, for example to select demonstrations.

Why keep a reference?

A judge is easiest to exploit far from examples it knows. One restraint is a KL penalty against a frozen reference, commonly the SFT checkpoint. A simplified sequence-level objective is

max⁡θ  Ey∼πθ(⋅∣x)[r(x,y)]−βDKL(πθ(⋅∣x)  ∥  πref(⋅∣x)).\max_\theta\;\mathbb E_{y\sim\pi_\theta(\cdot\mid x)}[r(x,y)] -\beta D_{\mathrm{KL}}\big(\pi_\theta(\cdot\mid x)\;\|\;\pi_{\mathrm{ref}}(\cdot\mid x)\big).

Read it as reward for the response, minus a cost for changing the distribution. Larger positive β makes deviation more expensive in this idealized objective. Real implementations estimate terms from samples and may add other objectives.

Three answers are enough to see the problem

Our invented reference chooses a brief correct answer 50% of the time, a detailed correct one 30%, and a confident false one 20%. The flawed judge assigns rewards 0, 1 and 2. At β=1, the optimal distribution is approximately 17.9%, 29.2%, 52.9%. The proxy improved and correctness fell from 80% to 47.1%.

Try it · toy model

Reward and Restraint

Optimize a three-answer policy and discover why more reward can mean less truth.

Know well12 min

This lab solves a three-action objective exactly; it does not run PPO. Reducing β makes the bad judge's favorite answer even more likely. Penalizing the unsupported claim changes the target and reverses the effect.

Deep diveThe finite-action optimumShould know

For actions with reference probabilities qiq_i, reward rir_i and positive β, the optimum is

pi∗=qieri/β∑jqjerj/β.p_i^*=\frac{q_i e^{r_i/\beta}}{\sum_j q_j e^{r_j/\beta}}.

The reference supplies a starting distribution; exponentiated reward tilts it. If the reference gives an action zero probability, finite forward KL keeps it outside the support. This analytic solution explains the lab; large neural policies require optimization instead.

Restraint is not truth

KL can limit a distribution change. It cannot determine whether a date was invented, a refusal was appropriate or a summary preserved uncertainty. Those require better feedback and separate evaluation. This is why monitoring reward alone is insufficient.

Keep the reference and the policy used to collect the latest rollout distinct. The latter appears in PPO's importance ratio; it need not be the frozen SFT model. The PPO concept follows those two roles explicitly.

Why should I care?

As a researcher

The learning signal is itself learned, so optimization and evaluation cannot be treated as the same problem.

As an engineer

Rollouts, reward scoring and reference log probabilities add cost; independent checks tell you whether that cost improved the assistant.

Modern systems that depend on it

  • feedback-tuned assistants
  • policy optimization
  • reward robustness research

Historical context

Before

SFT imitates answers already present in the demonstration dataset.

After

The policy is optimized using feedback on generated answers, usually with constraints on how far it moves.

Used today

InstructGPT and feedback-trained summarization are documented examples; later post-training recipes combine several objectives rather than all using one fixed pipeline.

What to remember

  • The reward model supplies a proxy, not a complete specification of good behavior.
  • The KL penalty is typically measured against a reference such as the SFT model.
  • PPO's old rollout policy has a different role from that reference.
  • A higher proxy reward can accompany worse factuality.

Key papers

Essential

Training language models to follow instructions with human feedback

Long Ouyang, Jeff Wu et al. · 2022 · NeurIPS 2022

InstructGPT: the supervised fine-tuning + reward model + RL recipe that turned GPT-3 into an instruction-following assistant, and the template for ChatGPT.

How to read it: Figure 2 is the three-step RLHF pipeline you'll meet in Chapter 10.

~1 h readarXiv:2203.02155✓ verified 2026-09-26
Essential

Proximal Policy Optimization Algorithms

John Schulman, Filip Wolski et al. · 2017

PPO: a simple, robust actor–critic policy-gradient method. It became the default RL algorithm in many labs and was the optimiser in InstructGPT-style RLHF.

~30 min readarXiv:1707.06347✓ verified 2026-09-26
Important

Learning to summarize from human feedback

Nisan Stiennon, Long Ouyang et al. · 2020

Connects preference learning to generated language before instruction-following assistants.

How to read it: Read the comparison setup and human evaluation before the aggregate scores.

~30 min readarXiv:2009.01325✓ verified 2026-10-04
Important

Scaling Laws for Reward Model Overoptimization

Leo Gao, John Schulman, Jacob Hilton · 2022

Explains why maximizing a learned judge can diverge from the intended outcome.

How to read it: Read the synthetic-gold-reward setup before interpreting the scaling curves.

~30 min readarXiv:2210.10760✓ verified 2026-10-04

Watch

3 h 31 min

Andrej Karpathy

Deep Dive into LLMs like ChatGPT

A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.

Should know