Concept · Chapter 10: From Base Model to Assistant
Reinforcement Learning from Human Feedback
RLHF improves a policy using a reward learned from human feedback, often with a penalty for departing from a reference model.
The problem
A fixed set of demonstrations does not tell a model which of its newly generated responses people would prefer.
The solution
Generate responses, score them with a preference-trained reward model, and optimize the policy while controlling drift.
The consequence
The model can improve beyond imitating fixed demonstrations, while also discovering ways to exploit mistakes in the reward.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Pretrain, Then Fine-Tune
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Supervised Fine-Tuning
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Autoregressive Next-Token Prediction
- Decoding: Greedy, Temperature, Top-k, Top-p
- Learning from Comparisons
- Features, Labels and Tasks
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Logistic Regression
- Reward Models
- Expected Value and Variance
- Reinforcement Learning
- MDPs, Policies and Value
- Policy Gradients
- KL Divergence
- Reinforcement Learning from Human Feedback
Let the policy try
In Chapter 5, a policy chose actions and received rewards. For a causal language model, the state is the prompt plus tokens generated so far, and an action is the next token. A completed response can receive a score from a learned reward model. A policy-gradient algorithm uses that signal to change the probabilities of the sampled actions.
InstructGPT first trained on demonstrations, then learned a reward model from ranked responses, then used PPO for policy optimization. This is an influential recipe, not the definition of every assistant.
- 01
Show an answer
Demonstrations train an SFT policy. Keep a frozen reference copy.
- 02
Compare answers
Judges rank candidates. A reward model learns to predict those comparisons.
- 03
Sample and improve
The policy generates responses. PPO uses reward and a penalty for drifting from the reference.
These are common teaching pipelines, not mandatory stages for every assistant. DPO-based systems can still use a reward model elsewhere, for example to select demonstrations.
Why keep a reference?
A judge is easiest to exploit far from examples it knows. One restraint is a KL penalty against a frozen reference, commonly the SFT checkpoint. A simplified sequence-level objective is
Read it as reward for the response, minus a cost for changing the distribution. Larger positive β makes deviation more expensive in this idealized objective. Real implementations estimate terms from samples and may add other objectives.
Three answers are enough to see the problem
Our invented reference chooses a brief correct answer 50% of the time, a detailed correct one 30%, and a confident false one 20%. The flawed judge assigns rewards 0, 1 and 2. At β=1, the optimal distribution is approximately 17.9%, 29.2%, 52.9%. The proxy improved and correctness fell from 80% to 47.1%.
Try it · toy model
Optimize a three-answer policy and discover why more reward can mean less truth.
This lab solves a three-action objective exactly; it does not run PPO. Reducing β makes the bad judge's favorite answer even more likely. Penalizing the unsupported claim changes the target and reverses the effect.
Deep diveThe finite-action optimumShould know
For actions with reference probabilities , reward and positive β, the optimum is
The reference supplies a starting distribution; exponentiated reward tilts it. If the reference gives an action zero probability, finite forward KL keeps it outside the support. This analytic solution explains the lab; large neural policies require optimization instead.
Restraint is not truth
KL can limit a distribution change. It cannot determine whether a date was invented, a refusal was appropriate or a summary preserved uncertainty. Those require better feedback and separate evaluation. This is why monitoring reward alone is insufficient.
Keep the reference and the policy used to collect the latest rollout distinct. The latter appears in PPO's importance ratio; it need not be the frozen SFT model. The PPO concept follows those two roles explicitly.
Why should I care?
As a researcher
The learning signal is itself learned, so optimization and evaluation cannot be treated as the same problem.
As an engineer
Rollouts, reward scoring and reference log probabilities add cost; independent checks tell you whether that cost improved the assistant.
Modern systems that depend on it
- feedback-tuned assistants
- policy optimization
- reward robustness research
Historical context
Before
SFT imitates answers already present in the demonstration dataset.
After
The policy is optimized using feedback on generated answers, usually with constraints on how far it moves.
Used today
InstructGPT and feedback-trained summarization are documented examples; later post-training recipes combine several objectives rather than all using one fixed pipeline.
What to remember
- The reward model supplies a proxy, not a complete specification of good behavior.
- The KL penalty is typically measured against a reference such as the SFT model.
- PPO's old rollout policy has a different role from that reference.
- A higher proxy reward can accompany worse factuality.
Key papers
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu et al. · 2022 · NeurIPS 2022
InstructGPT: the supervised fine-tuning + reward model + RL recipe that turned GPT-3 into an instruction-following assistant, and the template for ChatGPT.
How to read it: Figure 2 is the three-step RLHF pipeline you'll meet in Chapter 10.
Proximal Policy Optimization Algorithms
John Schulman, Filip Wolski et al. · 2017
PPO: a simple, robust actor–critic policy-gradient method. It became the default RL algorithm in many labs and was the optimiser in InstructGPT-style RLHF.
Learning to summarize from human feedback
Nisan Stiennon, Long Ouyang et al. · 2020
Connects preference learning to generated language before instruction-following assistants.
How to read it: Read the comparison setup and human evaluation before the aggregate scores.
Scaling Laws for Reward Model Overoptimization
Leo Gao, John Schulman, Jacob Hilton · 2022
Explains why maximizing a learned judge can diverge from the intended outcome.
How to read it: Read the synthetic-gold-reward setup before interpreting the scaling curves.
Watch
Andrej Karpathy
Deep Dive into LLMs like ChatGPT
A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 15: Alignment - SFT/RLHF
A technical companion to the chapter’s demonstration, preference and reinforcement-learning pipeline.