Skip to content
Road to Intelligence

Concept · Chapter 10: From Base Model to Assistant

PPO for Language Models

Should knowKnow well13 minDifficulty

PPO uses a clipped policy objective to limit the incentive for large changes relative to the policy that collected a rollout.

The problem

Repeatedly optimizing sampled actions can push a policy too far from the distribution that generated those samples.

The solution

Compare new and old action probabilities and clip the surrogate objective when further change would over-reward the update.

The consequence

Policy updates are easier to control, but clipping is neither a strict distance bound nor a guarantee of good behavior.

The two comparisons

In PPO, compare the new probability of a sampled token with its probability under the old rollout policy. In a common RLHF objective, also compare the policy with a reference, often the frozen SFT checkpoint. These comparisons solve different problems.

The rollout policy can refresh as training proceeds. The reference need not. Collapsing both into one “old model” hides what the algorithm is controlling.

A clipped incentive

Let ρt=πθ(at∣st)/πold(at∣st)\rho_t=\pi_\theta(a_t\mid s_t)/\pi_{\mathrm{old}}(a_t\mid s_t) and A^t\hat A_t be an advantage estimate. PPO maximizes

Lclip=Et[min⁡(ρtA^t,clip⁡(ρt,1−ϵ,1+ϵ)A^t)].L^{\mathrm{clip}}=\mathbb E_t\left[\min\left(\rho_t\hat A_t, \operatorname{clip}(\rho_t,1-\epsilon,1+\epsilon)\hat A_t\right)\right].

For a toy positive advantage of 2, ratio 1.4 and ε=0.2, the unclipped term is 2.8 and clipped term 2.4; the minimum is 2.4. Further increasing the ratio gets no extra benefit from this term. For negative advantage the other side of the range matters. See equation 7 and Figure 1 in the original paper.

In the language-model loop

Generate responses, collect token log probabilities, score completed responses, estimate returns/advantages and take optimizer steps. An actor-critic implementation also learns a value baseline. These components are training machinery, not extra modules required to produce every deployed answer.

What to remember

  • The old policy collected the rollout; the reference anchors the KL penalty.
  • Advantages compare a return with a baseline, rather than using raw reward alone.
  • Clipping the objective does not forbid all large policy changes.

Key papers

Essential

Training language models to follow instructions with human feedback

Long Ouyang, Jeff Wu et al. · 2022 · NeurIPS 2022

InstructGPT: the supervised fine-tuning + reward model + RL recipe that turned GPT-3 into an instruction-following assistant, and the template for ChatGPT.

How to read it: Figure 2 is the three-step RLHF pipeline you'll meet in Chapter 10.

~1 h readarXiv:2203.02155✓ verified 2026-09-26
Essential

Proximal Policy Optimization Algorithms

John Schulman, Filip Wolski et al. · 2017

PPO: a simple, robust actor–critic policy-gradient method. It became the default RL algorithm in many labs and was the optimiser in InstructGPT-style RLHF.

~30 min readarXiv:1707.06347✓ verified 2026-09-26

Watch