Skip to content
Road to Intelligence

Concept · Chapter 10: From Base Model to Assistant

Direct Preference Optimization

Must knowImplement16 minDifficulty

DPO trains a language model directly on preferred and rejected responses using their log-probability changes relative to a reference model.

The problem

An explicit reward-model and online reinforcement-learning pipeline has several moving parts and substantial rollout cost.

The solution

Use the connection between a KL-regularized reward objective and its optimal policy to form a direct pairwise training loss.

The consequence

Preference fitting becomes simpler, but still depends on data quality, reference choice and evaluation beyond the loss.

Compare two changes

For the same prompt, you have a chosen answer and a rejected answer. Ask how much the policy has increased each answer's probability relative to a fixed reference. DPO rewards a larger relative gain for the chosen answer.

Suppose reference log probabilities are −4 for the chosen response and −3 for the rejected one. The current policy assigns −3.5 and −3.2. The gains are +0.5 and −0.2, a difference of 0.7. At β=0.5, the scaled margin is 0.35, the implied pair preference is 58.7%, and the loss is 0.533 nats. At the reference, the margin is zero and loss is 0.693.

Try it · toy model

The DPO Subtraction

Follow chosen and rejected sequence probabilities through a reference-relative preference loss.

Implement10 min

The equation follows the example

For preferred ywy_w, rejected yly_l and prompt xx:

u=β[log⁡πθ(yw∣x)πref(yw∣x)−log⁡πθ(yl∣x)πref(yl∣x)],Lpair=−log⁡σ(u).u=\beta\left[\log\frac{\pi_\theta(y_w\mid x)}{\pi_{\mathrm{ref}}(y_w\mid x)}-\log\frac{\pi_\theta(y_l\mid x)}{\pi_{\mathrm{ref}}(y_l\mid x)}\right],\qquad \mathcal L_{\mathrm{pair}}=-\log\sigma(u).

Average this over training pairs. A response log probability is a sum of its conditional token log probabilities. Length normalization would define a different variant. Prompt boundaries, role tokens and end tokens must be handled consistently in policy and reference scoring.

The original derivation links this objective to KL-regularized reward optimization. Vanilla DPO needs neither a separately trained reward model nor online sampling during preference fitting. That theoretical connection is not a guarantee that finite-data DPO and a particular PPO run produce identical models.

A counterexample worth remembering

Click Both probabilities fall. The policy now assigns −4.5 to both responses. The chosen response lost 0.5 log units relative to its reference; the rejected response lost 1.5. The pair margin improved even though both absolute probabilities decreased. DPO controls a comparison, not two independent promises about absolute likelihood.

What to evaluate

Track the pair objective, but also compare held-out outputs for task success, factuality, refusals and retained capability. Preference data can reward verbosity, confidence or agreement with a user's mistake. DPO does not remove those biases simply by avoiding an explicit reward head. The TRL documentation provides a practical companion to the paper's mathematics.

Why should I care?

As a researcher

It connects reward and policy parameterizations, while making its assumptions and finite-data behavior worth examining.

As an engineer

Training uses paired responses and reference log probabilities; formatting or masking errors change the objective.

Modern systems that depend on it

  • preference-tuned checkpoints
  • offline preference learning
  • open assistant recipes

Historical context

Before

A common route learns an explicit reward model and repeatedly generates responses for reinforcement learning.

After

The language model itself is fitted directly to a reference-relative preference objective.

Used today

DPO is used in Zephyr and in the documented Llama 3 and Tulu 3 post-training pipelines.

What to remember

  • Both chosen and rejected response probabilities enter the loss.
  • Sequence log probabilities sum token log probabilities under consistent masks.
  • The objective does not guarantee that every chosen response's absolute probability rises.
  • A DPO-based pipeline can still use a reward model for candidate selection elsewhere.

Key papers

Essential

The Llama 3 Herd of Models

Aaron Grattafiori, Abhimanyu Dubey et al. · 2024

The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.

How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.

~2 h readarXiv:2407.21783✓ verified 2026-10-04
Important

Zephyr: Direct Distillation of LM Alignment

Lewis Tunstall, Edward Beeching et al. · 2023

A practical application of DPO and synthetic data.

How to read it: Follow the data stages; do not attribute the entire outcome to one optimizer.

~25 min readarXiv:2310.16944✓ verified 2026-10-04
Important

Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Nathan Lambert, Jacob Morrison et al. · 2024

Provides an open case study connecting data, objectives and evaluation.

How to read it: Read the evaluation split and training stages; leave reasoning RL detail for Chapter 14.

~40 min readarXiv:2411.15124✓ verified 2026-10-04

Watch