Concept · Chapter 10: From Base Model to Assistant
PPO for Language Models
PPO uses a clipped policy objective to limit the incentive for large changes relative to the policy that collected a rollout.
The problem
Repeatedly optimizing sampled actions can push a policy too far from the distribution that generated those samples.
The solution
Compare new and old action probabilities and clip the surrogate objective when further change would over-reward the update.
The consequence
Policy updates are easier to control, but clipping is neither a strict distance bound nor a guarantee of good behavior.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Pretrain, Then Fine-Tune
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Supervised Fine-Tuning
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Autoregressive Next-Token Prediction
- Decoding: Greedy, Temperature, Top-k, Top-p
- Learning from Comparisons
- Features, Labels and Tasks
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Logistic Regression
- Reward Models
- Expected Value and Variance
- Reinforcement Learning
- MDPs, Policies and Value
- Policy Gradients
- KL Divergence
- Reinforcement Learning from Human Feedback
- Q-Learning
- Actor–Critic Methods
- PPO for Language Models
The two comparisons
In PPO, compare the new probability of a sampled token with its probability under the old rollout policy. In a common RLHF objective, also compare the policy with a reference, often the frozen SFT checkpoint. These comparisons solve different problems.
The rollout policy can refresh as training proceeds. The reference need not. Collapsing both into one “old model” hides what the algorithm is controlling.
A clipped incentive
Let and be an advantage estimate. PPO maximizes
For a toy positive advantage of 2, ratio 1.4 and ε=0.2, the unclipped term is 2.8 and clipped term 2.4; the minimum is 2.4. Further increasing the ratio gets no extra benefit from this term. For negative advantage the other side of the range matters. See equation 7 and Figure 1 in the original paper.
In the language-model loop
Generate responses, collect token log probabilities, score completed responses, estimate returns/advantages and take optimizer steps. An actor-critic implementation also learns a value baseline. These components are training machinery, not extra modules required to produce every deployed answer.
What to remember
- The old policy collected the rollout; the reference anchors the KL penalty.
- Advantages compare a return with a baseline, rather than using raw reward alone.
- Clipping the objective does not forbid all large policy changes.
Key papers
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu et al. · 2022 · NeurIPS 2022
InstructGPT: the supervised fine-tuning + reward model + RL recipe that turned GPT-3 into an instruction-following assistant, and the template for ChatGPT.
How to read it: Figure 2 is the three-step RLHF pipeline you'll meet in Chapter 10.
Proximal Policy Optimization Algorithms
John Schulman, Filip Wolski et al. · 2017
PPO: a simple, robust actor–critic policy-gradient method. It became the default RL algorithm in many labs and was the optimiser in InstructGPT-style RLHF.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 15: Alignment - SFT/RLHF
A technical companion to the chapter’s demonstration, preference and reinforcement-learning pipeline.