Concept · Chapter 10: From Base Model to Assistant
Direct Preference Optimization
DPO trains a language model directly on preferred and rejected responses using their log-probability changes relative to a reference model.
The problem
An explicit reward-model and online reinforcement-learning pipeline has several moving parts and substantial rollout cost.
The solution
Use the connection between a KL-regularized reward objective and its optimal policy to form a direct pairwise training loss.
The consequence
Preference fitting becomes simpler, but still depends on data quality, reference choice and evaluation beyond the loss.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Pretrain, Then Fine-Tune
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Supervised Fine-Tuning
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Autoregressive Next-Token Prediction
- Decoding: Greedy, Temperature, Top-k, Top-p
- Learning from Comparisons
- KL Divergence
- Direct Preference Optimization
Compare two changes
For the same prompt, you have a chosen answer and a rejected answer. Ask how much the policy has increased each answer's probability relative to a fixed reference. DPO rewards a larger relative gain for the chosen answer.
Suppose reference log probabilities are −4 for the chosen response and −3 for the rejected one. The current policy assigns −3.5 and −3.2. The gains are +0.5 and −0.2, a difference of 0.7. At β=0.5, the scaled margin is 0.35, the implied pair preference is 58.7%, and the loss is 0.533 nats. At the reference, the margin is zero and loss is 0.693.
Try it · toy model
Follow chosen and rejected sequence probabilities through a reference-relative preference loss.
The equation follows the example
For preferred , rejected and prompt :
Average this over training pairs. A response log probability is a sum of its conditional token log probabilities. Length normalization would define a different variant. Prompt boundaries, role tokens and end tokens must be handled consistently in policy and reference scoring.
The original derivation links this objective to KL-regularized reward optimization. Vanilla DPO needs neither a separately trained reward model nor online sampling during preference fitting. That theoretical connection is not a guarantee that finite-data DPO and a particular PPO run produce identical models.
A counterexample worth remembering
Click Both probabilities fall. The policy now assigns −4.5 to both responses. The chosen response lost 0.5 log units relative to its reference; the rejected response lost 1.5. The pair margin improved even though both absolute probabilities decreased. DPO controls a comparison, not two independent promises about absolute likelihood.
What to evaluate
Track the pair objective, but also compare held-out outputs for task success, factuality, refusals and retained capability. Preference data can reward verbosity, confidence or agreement with a user's mistake. DPO does not remove those biases simply by avoiding an explicit reward head. The TRL documentation provides a practical companion to the paper's mathematics.
Why should I care?
As a researcher
It connects reward and policy parameterizations, while making its assumptions and finite-data behavior worth examining.
As an engineer
Training uses paired responses and reference log probabilities; formatting or masking errors change the objective.
Modern systems that depend on it
- preference-tuned checkpoints
- offline preference learning
- open assistant recipes
Historical context
Before
A common route learns an explicit reward model and repeatedly generates responses for reinforcement learning.
After
The language model itself is fitted directly to a reference-relative preference objective.
Used today
DPO is used in Zephyr and in the documented Llama 3 and Tulu 3 post-training pipelines.
What to remember
- Both chosen and rejected response probabilities enter the loss.
- Sequence log probabilities sum token log probabilities under consistent masks.
- The objective does not guarantee that every chosen response's absolute probability rises.
- A DPO-based pipeline can still use a reward model for candidate selection elsewhere.
Key papers
The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey et al. · 2024
The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.
How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafael Rafailov, Archit Sharma et al. · 2023
Provides a direct preference-training route with fewer moving parts.
How to read it: Start with equations 3 and 7; inspect the assumptions behind the policy/reward connection.
Zephyr: Direct Distillation of LM Alignment
Lewis Tunstall, Edward Beeching et al. · 2023
A practical application of DPO and synthetic data.
How to read it: Follow the data stages; do not attribute the entire outcome to one optimizer.
Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Nathan Lambert, Jacob Morrison et al. · 2024
Provides an open case study connecting data, objectives and evaluation.
How to read it: Read the evaluation split and training stages; leave reasoning RL detail for Chapter 14.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 15: Alignment - SFT/RLHF
A technical companion to the chapter’s demonstration, preference and reinforcement-learning pipeline.