Skip to content
Road to Intelligence

Part III · Large Language Models

Chapter 10

From Base Model to Assistant

Why a pretrained model isn't yet helpful — and how post-training changes that.

2 h core path12 concepts4 interactivesCore path · 3 optional concepts foldedDeep · all 12 concepts shown in full

In one sentencePost-training uses demonstrations and feedback to make a pretrained next-token model more reliably behave as an assistant.

The problem

The checkpoint is not the assistant

Chapter 9 ended with a checkpoint: a large collection of weights learned from trillions of tokens. It can continue code, stories, arguments and conversations. What it has not necessarily learned is that, when you ask a question, it should stop being a document and start being your assistant.

Give it a quiz question and one plausible continuation is another quiz question. Give it a carefully framed example conversation and it may answer well. The capability and the reliability of the behavior are different questions. A base model is not incapable of answering; its pretraining objective simply does not specify the conversational contract you have in mind.

Post-training works on that contract. The generation mechanism remains familiar: predict a token, append it, predict the next. The training signal changes what counts as a good continuation.

Training signalWhat the data saysWhat the model learns from
PretrainingHere is text that occurred.The next observed token.
Supervised fine-tuningHere is a response we want.Demonstrated target tokens.
Preference trainingThis response is better than that one.Comparisons, directly or through a learned reward.

The distinction has practical consequences. In InstructGPT's evaluation, people preferred outputs from a 1.3B-parameter tuned model to those from the 175B GPT-3 base model Established. That is a result on the paper's prompt distribution, not a claim that the smaller model acquired every ability of the larger one. The users were judging answers, not parameter counts.

Demonstrations

Show the answer

Suppose the task is to summarize a short transport report in two sentences, preserve its uncertainty and add no facts. A demonstration contains the source paragraph, that instruction and a suitable summary. Thousands of such examples can teach the model that this is the kind of continuation to produce.

This is supervised fine-tuning, or SFT. It is the same forward pass, loss, backward pass and weight update from earlier chapters, applied to a different collection of examples. Instruction tuning is the version in which those examples describe tasks as instructions. FLAN and T0 study how training on many tasks helps with instructions from held-out task families.

For a data engineer, the important change is the unit of data. A cleaned document was enough for pretraining. Now a record needs a task, a response, role boundaries and a reason to trust that response. Coverage matters: if every demonstration is verbose and certain, brevity and appropriate uncertainty will be underrepresented.

The InstructGPT paper's first version reports 12,725 SFT training prompts, 33,207 reward-model training prompts and 31,144 PPO training prompts Established. Those are prompts, not response pairs or all the tokens consumed. The three datasets supply different signals; “more training data” alone does not describe the recipe.

The objective

What the loss sees

Consider an invented training sequence: 2 + 2 ? Four . END. Prompt tokens remain available as context, while a loss mask can select only the three assistant targets. With probabilities 0.50, 0.25 and 0.80, their negative log probabilities average to about 0.768 nats.

Masking the prompt does not make it invisible. The answer still depends on it. It means the training objective does not directly reward predicting the user's words. Some recipes supervise all tokens; some supervise assistant turns or only the final completion. The choice must agree with the intended task.

Try it · toy model

What Does the Loss See?

Select training targets, change their probabilities and see exactly what a loss mask removes.

Implement10 min

Start with Assistant only and move the prompt slider. Nothing happens to the loss. Now select All tokens: the same probability affects the objective. The model did not change; the definition of what it is being asked to learn did.

During training, the prefix comes from the demonstration. When predicting the period, the model is given Four, even if it would have guessed something else. This is teacher forcing. At inference, it conditions on its own generated tokens instead. A low demonstration loss therefore does not guarantee that a long sampled conversation will stay on track.

The interface

Conversation has a format

The application holds messages with roles: system, user, assistant. A causal language model receives a token sequence. A chat template connects the two by choosing role markers, turn terminators and the prefix at which the assistant should continue.

The tokens are model-specific. A perfectly good demonstration can be damaged by an extra beginning token, a missing terminator or the wrong assistant marker. Inspect one serialized conversation and its mask before processing a million of them. This is a schema check with behavioral consequences.

Role markers are learned conventions, not a mathematical firewall. They help establish who is speaking; they do not, by themselves, solve prompt injection or application security. Those questions return in the systems chapters.

Comparisons

Which answer is better?

A demonstration gives one answer to imitate. It does not rank the many other answers a model might generate. Comparing candidates can be easier than writing an ideal response from scratch.

Here is our invented source: a town tested a bus route for four weeks; ridership rose; the trial was too short to tell whether the increase would last; no launch date was announced. One candidate preserves those facts in two sentences. Another announces a permanent launch date. A third is faithful but uses five sentences. A fourth refuses unnecessarily.

Which should win? “Sounds confident” and “preserves the source” do not choose the same answer. A preference label only becomes meaningful alongside the prompt and judging rubric.

Four ranked responses yield six pair comparisons, but not six independent prompts. Keep all variants of a prompt on the same side of the train/test split. Record ties and disagreement. “Chosen” in a dataset means preferred under a judgment, not guaranteed true.

A learned judge

Learning a judge

To reuse those judgments, train a reward model. It reads the prompt and response and returns a scalar. The training loss asks the chosen response to score higher than the rejected one. A common conversion from scores to a preference probability is

P(chosen wins)=σ(rchosen−rrejected).P(\text{chosen wins})=\sigma\big(r_{\mathrm{chosen}}-r_{\mathrm{rejected}}\big).

Equal scores give 50–50. A difference of one gives about 73.1%. Only the difference matters: scores 100 and 99 express the same pair probability as 1 and 0. The scale is not a meter of truth.

Try it · toy model

Teach a Preference

Judge candidate summaries and fit a small reward model. Change the rubric and watch the ranking change.

Implement12 min

Label a pair, then fit the judge. For a sharper demonstration, open the two rubric presets. Faithfulness first favors the grounded summary. Confidence first rewards the invented conclusion. The same fitting procedure succeeds at learning either set of labels. This tiny model fits four free scores; a real reward model must also generalize to unseen language.

Optimization

Reward with restraint

Now let the assistant generate its own responses. The reward model scores them; reinforcement learning changes the policy to favor higher-scoring responses. In the InstructGPT recipe, PPO performs those updates after SFT. This is a common form of reinforcement learning from human feedback, RLHF.

There is a catch. A model actively searching for high reward can discover mistakes that a passive validation set did not expose. One restraint is to penalize departure from a frozen reference, commonly the SFT model:

maximize expected reward  −  β DKL(πθ  ∥  πref).\text{maximize expected reward}\; -\;\beta\,D_{\mathrm{KL}}(\pi_\theta\;\|\;\pi_{\mathrm{ref}}).

The KL term says how much the distribution changed. It cannot say whether a claim is true. Our next lab makes the difference measurable: its flawed judge rewards the confident false response most highly. At β=1, optimization raises that response's probability from 20% to about 53%.

Try it · toy model

Reward and Restraint

Optimize a three-answer policy and discover why more reward can mean less truth.

Know well12 min

Lower the restraint. Reward rises while correctness falls. Then penalize the unsupported claim: changing the judge changes what optimization achieves. This is an exact three-action calculation, not a PPO simulation or an empirical LLM result.

Deep diveWhy PPO has another old modelShould know

PPO compares a token's new probability with its probability under the policy that collected the rollout. Its clipped objective limits the incentive for large updates to that sampled policy. This old rollout policy is distinct from the reference policy used for the KL penalty. The rollout policy refreshes; the reference can remain fixed. Clipping is not a guarantee of a maximum policy distance or safe behavior.

OptionalPPO for Language Models· folded on the core path. Open it, or switch to Deep to show it here.

A shorter route

The DPO shortcut

The reward-model-plus-RL route is not the only way to use preferences. Direct Preference Optimization, DPO, fits the language model directly to chosen/rejected pairs, using a fixed reference to measure how each response's probability changes.

The arithmetic fits on a napkin. The reference log probabilities are −4 and −3. The policy's are −3.5 and −3.2. The chosen answer gained 0.5 log units; the rejected answer lost 0.2. The relative difference is 0.7. Multiply by β=0.5, apply a sigmoid and take negative log: the pair loss is 0.533, down from 0.693 at the reference.

Try it · toy model

The DPO Subtraction

Follow chosen and rejected sequence probabilities through a reference-relative preference loss.

Implement10 min

Try Both probabilities fall. DPO can improve the comparison while making both answers absolutely less likely, because the rejected answer loses more relative to its reference. It does not independently promise to raise every chosen answer's probability.

Two routes from demonstrations to preferences
  1. 01

    Show an answer

    Demonstrations train an SFT policy. Keep a frozen reference copy.

  2. 02

    Compare answers

    Judges rank candidates. A reward model learns to predict those comparisons.

  3. 03

    Sample and improve

    The policy generates responses. PPO uses reward and a penalty for drifting from the reference.

These are common teaching pipelines, not mandatory stages for every assistant. DPO-based systems can still use a reward model elsewhere, for example to select demonstrations.

Vanilla DPO avoids a separately trained reward model and online sampling during preference fitting. A whole training pipeline can still use judges elsewhere. Llama 3's report describes reward-based candidate selection alongside SFT and DPO. Do not confuse one objective with the entire development process.

The source of judgment

Whose feedback?

Human comparisons take time. Models can help write demonstrations, critique drafts or compare responses. The important question is which signal they supply and how it is checked.

Constitutional AI uses explicit principles to guide critique and revision, and AI-generated preferences for a feedback stage. Humans still choose the principles and judging setup. AI feedback shifts annotation work; it does not eliminate human choices or automatically create reliable ground truth.

OptionalConstitutional AI and AI Feedback· folded on the core path. Open it, or switch to Deep to show it here.

Another operation is to generate several candidates, select good ones and train on the selected responses. This rejection sampling supplies an SFT dataset; it is different from optimizing a policy on sampled rewards. It also explains why inference efficiency matters while building a model: the training pipeline may generate and judge many answers before a user ever sees one.

OptionalGenerate, Rank, Then Train· folded on the core path. Open it, or switch to Deep to show it here.

“Pretraining gives knowledge; post-training only changes style” is a useful first approximation that becomes misleading when treated as a law Interpretation. Reported recipes such as Tulu 3 target multiple skills and combine SFT, DPO and verifiable rewards. Chapter 14 takes up reasoning-oriented reinforcement learning; here, measure the task changes instead of assuming they must be cosmetic.

Behavioral boundaries

Safety is more than refusal

A system that always says no would pass a test consisting only of requests it should decline. It would also fail as an assistant. Safety tuning needs examples of appropriate responses and appropriate refusals, including benign requests that resemble problematic ones.

Helpfulness, truthfulness and harmlessness are related but separate aims. A polite hallucination, an accurate but inappropriate response and an unnecessary refusal fail in different ways. Training examples and preferences can shape these behaviors; they do not replace the wider system checks covered in Chapter 16.

The exam

Test the assistant

After each stage, ask what actually improved. Keep a development set for choosing the recipe and an unseen test set for the eventual claim. Use prompt-level splits, blind model identity in comparisons, randomize answer order and report uncertainty rather than just a win percentage.

CheckA concrete failure
Instruction followingThe summary has five sentences instead of two.
FaithfulnessIt invents a launch date.
UncertaintyA short trial becomes proof of permanent demand.
Appropriate refusalIt refuses a harmless transport summary.
Retained abilityA familiar factual or coding task regresses after tuning.

Reward is useful training feedback, but it is not the exam. The reward-overoptimization paper studies divergence from a synthetic gold reward; work on sycophancy investigates feedback that favors agreement with a user's belief over a more accurate response. Neither finding justifies a universal claim that preference training always fails. They tell us what to test.

For an engineer, this chapter is a reminder that schemas, labels and evaluation splits are part of model behavior. For a researcher, it exposes a coupled learning problem: the policy changes, the judge is imperfect, and the metric can become a target. More optimization is only helpful when the signal continues to represent what you want.

From comparisons to assistant training, 2017–2024Open the full timeline →
Transformers and scaleLLMs, agents and reasoning (evolving)2017: Sparsely-gated mixture of expertsSparsely-gated mixture of experts20172017: The TransformerThe Transformer20172017: Learning from human preferencesLearning from human preferences20172018: GPT-1: generative pretrainingGPT-1: generative pretraining20182018: BERTBERT20182019: GPT-2GPT-220192019: ZeRO removes redundant copiesZeRO removes redundant copies20192020: Scaling lawsScaling laws20202020: GPT-3 and in-context learningGPT-3 and in-context learning20202020: Summaries learned from human feedbackSummaries learned from human feedback20202020: Vision TransformerVision Transformer20202020: AlphaFold 2AlphaFold 220202021: CLIP: images meet languageCLIP: images meet language20212021: LoRA: low-rank fine-tuningLoRA: low-rank fine-tuning20212022: Chain-of-thought promptingChain-of-thought prompting20222022: InstructGPT and RLHFInstructGPT and RLHF20222022: Chinchilla: compute-optimal trainingChinchilla: compute-optimal…20222022: FlashAttentionFlashAttention20222022: Open text-to-image modelsOpen text-to-image models20222022: ChatGPTChatGPT20222023: GPT-4GPT-420232023: Direct Preference OptimizationDirect Preference Optimization20232023: Capable open-weight LLMsCapable open-weight LLMs20232023: PagedAttention and vLLMPagedAttention and vLLM20232024: Mixtral: an open mixture-of-experts modelMixtral: an open mixture-of-experts model20242024: The Llama 3 report opens up a frontier-scale runThe Llama 3 report opens up…20242024: EU AI Act enters into forceEU AI Act enters into force20242024: Reasoning models: OpenAI o1Reasoning models: OpenAI o120242024: Nobel Prizes for AI researchNobel Prizes…20242024: DeepSeek-V3DeepSeek-V32024

What comes next

The assistant still has to run

The output is another checkpoint. It still predicts tokens, but demonstrations and feedback have shaped which continuations it favors. Deploying it creates a different set of problems: growing context memory, serial token generation, many simultaneous requests and the cost of adapting large weights. Chapter 11 goes inside the engineering that makes those systems practical.

Concepts in this chapter

Mark each one as you go. Must-know concepts are the core path.

What do I actually need to remember?

  • Base and assistant models can share the same next-token generation mechanism.
  • SFT updates weights using demonstrations; examples in a prompt do not.
  • A loss mask selects training targets without removing the prompt from context.
  • Preferences depend on a prompt, judge and rubric.
  • Reward-model scores are proxies, not truth.
  • RLHF can optimize sampled responses with a penalty for drifting from a reference.
  • PPO’s rollout policy and the KL reference serve different purposes.
  • DPO fits chosen/rejected pairs relative to a reference; the whole pipeline may still use judges.
  • AI feedback still depends on human choices about principles, models and data.
  • Evaluate task success, factuality, safety and retained capability separately.

You do not need to memorize everything else. This list is the revision sheet.

Key papers

Essential

Training language models to follow instructions with human feedback

Long Ouyang, Jeff Wu et al. · 2022 · NeurIPS 2022

InstructGPT: the supervised fine-tuning + reward model + RL recipe that turned GPT-3 into an instruction-following assistant, and the template for ChatGPT.

Problem
Pretrained language models continue text; they don't reliably follow instructions or behave helpfully.
What was new
Fine-tune on human demonstrations, train a reward model on human rankings, then optimize the model against it with PPO.

How to read it: Figure 2 is the three-step RLHF pipeline you'll meet in Chapter 10.

~1 h readarXiv:2203.02155✓ verified 2026-09-26
Essential

Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Rafael Rafailov, Archit Sharma et al. · 2023

Provides a direct preference-training route with fewer moving parts.

Problem
A separate reward-model and online RL loop is operationally complex.
What was new
Derive a reference-relative policy objective fitted directly to preference pairs.

How to read it: Start with equations 3 and 7; inspect the assumptions behind the policy/reward connection.

~40 min readarXiv:2305.18290✓ verified 2026-10-04
Essential

Deep reinforcement learning from human preferences

Paul Christiano, Jan Leike et al. · 2017 · NeurIPS 2017

Showed that agents can be trained from human comparisons between behaviours rather than a hand-written reward — the foundation of RLHF.

Problem
For many tasks, writing a reward function that captures what we want is hard or impossible.
What was new
Learn a reward model from pairwise human preferences, then optimize a policy against it with reinforcement learning.
~45 min readarXiv:1706.03741✓ verified 2026-09-26
Essential

Proximal Policy Optimization Algorithms

John Schulman, Filip Wolski et al. · 2017

PPO: a simple, robust actor–critic policy-gradient method. It became the default RL algorithm in many labs and was the optimiser in InstructGPT-style RLHF.

Problem
Policy-gradient updates that are too large can wreck a policy in a single step; the principled fixes were complicated.
What was new
A clipped objective that removes the incentive to move the policy too far from the one that collected the data, so the same batch can be reused for several updates.
~30 min readarXiv:1707.06347✓ verified 2026-09-26
Important

Finetuned Language Models Are Zero-Shot Learners

Jason Wei, Maarten Bosma et al. · 2021 · ICLR 2022

Instruction tuning: fine-tune on many tasks phrased as instructions and the model follows instructions for new tasks too. The 137B FLAN beat zero-shot GPT-3 on 20 of 25 tasks.

Problem
Base models were good at few-shot prompting but weak at simply following an instruction with no examples.
What was new
Fine-tune a pretrained model on over 60 datasets rewritten as natural-language instructions, then test on unseen task types.
~35 min readarXiv:2109.01652✓ verified 2026-09-26
Important

Learning to summarize from human feedback

Nisan Stiennon, Long Ouyang et al. · 2020

Connects preference learning to generated language before instruction-following assistants.

Problem
Automatic summary metrics miss qualities that people care about.
What was new
Learn a reward from human summary comparisons and optimize a summarization policy.

How to read it: Read the comparison setup and human evaluation before the aggregate scores.

~30 min readarXiv:2009.01325✓ verified 2026-10-04
Important

Multitask Prompted Training Enables Zero-Shot Task Generalization

Victor Sanh, Albert Webson et al. · 2021

Shows why task diversity and prompt design belong in the instruction-data story.

Problem
Supervised models often specialize in the tasks used for training.
What was new
Train on many tasks expressed through prompts and evaluate generalization to held-out tasks.

How to read it: Compare task splits and prompt formulations with FLAN.

~30 min readarXiv:2110.08207✓ verified 2026-10-04
Important

Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Yuntao Bai, Andy Jones et al. · 2022

Makes feedback data and the helpfulness/harmlessness tension concrete.

Problem
Helpfulness and harmlessness are not one easily specified objective.
What was new
Train assistants with preference feedback and examine behavior along these separate dimensions.

How to read it: Inspect the data and evaluation categories; do not reduce safety to refusal rate.

~35 min readarXiv:2204.05862✓ verified 2026-10-04
Important

Scaling Laws for Reward Model Overoptimization

Leo Gao, John Schulman, Jacob Hilton · 2022

Explains why maximizing a learned judge can diverge from the intended outcome.

Problem
A policy can exploit errors in a proxy reward model.
What was new
Study optimization against a proxy while measuring a synthetic gold reward.

How to read it: Read the synthetic-gold-reward setup before interpreting the scaling curves.

~30 min readarXiv:2210.10760✓ verified 2026-10-04
Important

Constitutional AI: Harmlessness from AI Feedback

Yuntao Bai, Saurav Kadavath et al. · 2022

A documented route for scaling some feedback while retaining human-selected principles.

Problem
Direct human feedback is costly and principles need consistent application.
What was new
Use principle-guided critique/revision and AI-generated preferences in separate training stages.

How to read it: Distinguish the supervised revision stage from the feedback/RL stage.

~30 min readarXiv:2212.08073✓ verified 2026-10-04
Important

Self-Instruct: Aligning Language Models with Self-Generated Instructions

Yizhong Wang, Yeganeh Kordi et al. · 2022

Makes synthetic instruction generation an explicit data pipeline to evaluate.

Problem
Writing enough diverse instruction examples by hand is expensive.
What was new
Generate and filter synthetic instructions and instances from a language model.

How to read it: Trace seed examples, generation, filtering and held-out evaluation separately.

~25 min readarXiv:2212.10560✓ verified 2026-10-04
Optional

LIMA: Less Is More for Alignment

Chunting Zhou, Pengfei Liu et al. · 2023

A bounded case study of data quality, not a universal sample-size rule.

Problem
How far can carefully curated supervised examples shape a strong base model?
What was new
Study a 65B model fine-tuned on 1,000 curated examples without reinforcement learning.

How to read it: Separate the reported results from broader interpretations about where capability comes from.

~25 min readarXiv:2305.11206✓ verified 2026-10-04
Important

UltraFeedback: Boosting Language Models with Scaled AI Feedback

Ganqu Cui, Lifan Yuan et al. · 2023

A concrete source of synthetic feedback used in open assistant research.

Problem
Feedback annotation needs scale and more than one quality dimension.
What was new
Collect model responses with AI-generated multidimensional feedback.

How to read it: Distinguish original annotations from a binarized derivative dataset.

~25 min readarXiv:2310.01377✓ verified 2026-10-04
Important

Towards Understanding Sycophancy in Language Models

Mrinank Sharma, Meg Tong et al. · 2023

Separates being preferred from being accurate.

Problem
A response can please a user by agreeing with a mistaken belief.
What was new
Investigate sycophancy and how preference judgments can reward it.

How to read it: Inspect task construction and judge behavior before generalizing to all feedback-trained models.

~25 min readarXiv:2310.13548✓ verified 2026-10-04
Important

Zephyr: Direct Distillation of LM Alignment

Lewis Tunstall, Edward Beeching et al. · 2023

A practical application of DPO and synthetic data.

Problem
Open assistant training needs an effective, reproducible preference route.
What was new
Combine distilled instruction data and direct preference optimization.

How to read it: Follow the data stages; do not attribute the entire outcome to one optimizer.

~25 min readarXiv:2310.16944✓ verified 2026-10-04
Essential

The Llama 3 Herd of Models

Aaron Grattafiori, Abhimanyu Dubey et al. · 2024

The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.

Problem
Frontier labs had largely stopped publishing how their models were built.
What was new
A 405B model trained on 15.6T tokens with 3.8 × 10²⁵ FLOPs, with detailed data cleaning, a data mix chosen by scaling experiments, 4D parallelism at 38–43% utilisation, and 466 job interruptions in a 54-day period.

How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.

~2 h readarXiv:2407.21783✓ verified 2026-10-04
Important

Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Nathan Lambert, Jacob Morrison et al. · 2024

Provides an open case study connecting data, objectives and evaluation.

Problem
Strong post-training recipes are often difficult to inspect or reproduce.
What was new
Release a pipeline combining SFT, DPO, verifiable rewards and explicit development/unseen evaluation.

How to read it: Read the evaluation split and training stages; leave reasoning RL detail for Chapter 14.

~40 min readarXiv:2411.15124✓ verified 2026-10-04

Watch

3 h 31 min

Andrej Karpathy

Deep Dive into LLMs like ChatGPT

A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.

Covers: The full training stack behind a chat model, from internet text and tokens to a pretrained base model and the post-training that turns it into an assistant.

Should know

What came next?

Chapter 11

Inside Modern LLMs →

The original Transformer is too slow and memory-hungry to serve at modern scale.