Part III · Large Language Models
Chapter 10
From Base Model to Assistant
Why a pretrained model isn't yet helpful — and how post-training changes that.
In one sentencePost-training uses demonstrations and feedback to make a pretrained next-token model more reliably behave as an assistant.
The problem
The checkpoint is not the assistant
Chapter 9 ended with a checkpoint: a large collection of weights learned from trillions of tokens. It can continue code, stories, arguments and conversations. What it has not necessarily learned is that, when you ask a question, it should stop being a document and start being your assistant.
Give it a quiz question and one plausible continuation is another quiz question. Give it a carefully framed example conversation and it may answer well. The capability and the reliability of the behavior are different questions. A base model is not incapable of answering; its pretraining objective simply does not specify the conversational contract you have in mind.
Post-training works on that contract. The generation mechanism remains familiar: predict a token, append it, predict the next. The training signal changes what counts as a good continuation.
| Training signal | What the data says | What the model learns from |
|---|---|---|
| Pretraining | Here is text that occurred. | The next observed token. |
| Supervised fine-tuning | Here is a response we want. | Demonstrated target tokens. |
| Preference training | This response is better than that one. | Comparisons, directly or through a learned reward. |
The distinction has practical consequences. In InstructGPT's evaluation, people preferred outputs from a 1.3B-parameter tuned model to those from the 175B GPT-3 base model Established. That is a result on the paper's prompt distribution, not a claim that the smaller model acquired every ability of the larger one. The users were judging answers, not parameter counts.
Demonstrations
Show the answer
Suppose the task is to summarize a short transport report in two sentences, preserve its uncertainty and add no facts. A demonstration contains the source paragraph, that instruction and a suitable summary. Thousands of such examples can teach the model that this is the kind of continuation to produce.
This is supervised fine-tuning, or SFT. It is the same forward pass, loss, backward pass and weight update from earlier chapters, applied to a different collection of examples. Instruction tuning is the version in which those examples describe tasks as instructions. FLAN and T0 study how training on many tasks helps with instructions from held-out task families.
For a data engineer, the important change is the unit of data. A cleaned document was enough for pretraining. Now a record needs a task, a response, role boundaries and a reason to trust that response. Coverage matters: if every demonstration is verbose and certain, brevity and appropriate uncertainty will be underrepresented.
The InstructGPT paper's first version reports 12,725 SFT training prompts, 33,207 reward-model training prompts and 31,144 PPO training prompts Established. Those are prompts, not response pairs or all the tokens consumed. The three datasets supply different signals; “more training data” alone does not describe the recipe.
The objective
What the loss sees
Consider an invented training sequence: 2 + 2 ? Four . END. Prompt tokens remain available as context, while a loss mask can select only the three assistant targets. With probabilities 0.50, 0.25 and 0.80, their negative log probabilities average to about 0.768 nats.
Masking the prompt does not make it invisible. The answer still depends on it. It means the training objective does not directly reward predicting the user's words. Some recipes supervise all tokens; some supervise assistant turns or only the final completion. The choice must agree with the intended task.
Try it · toy model
Select training targets, change their probabilities and see exactly what a loss mask removes.
Start with Assistant only and move the prompt slider. Nothing happens to the loss. Now select All tokens: the same probability affects the objective. The model did not change; the definition of what it is being asked to learn did.
During training, the prefix comes from the demonstration. When predicting the period, the model is given Four, even if it would have guessed something else. This is teacher forcing. At inference, it conditions on its own generated tokens instead. A low demonstration loss therefore does not guarantee that a long sampled conversation will stay on track.
The interface
Conversation has a format
The application holds messages with roles: system, user, assistant. A causal language model receives a token sequence. A chat template connects the two by choosing role markers, turn terminators and the prefix at which the assistant should continue.
The tokens are model-specific. A perfectly good demonstration can be damaged by an extra beginning token, a missing terminator or the wrong assistant marker. Inspect one serialized conversation and its mask before processing a million of them. This is a schema check with behavioral consequences.
Role markers are learned conventions, not a mathematical firewall. They help establish who is speaking; they do not, by themselves, solve prompt injection or application security. Those questions return in the systems chapters.
Comparisons
Which answer is better?
A demonstration gives one answer to imitate. It does not rank the many other answers a model might generate. Comparing candidates can be easier than writing an ideal response from scratch.
Here is our invented source: a town tested a bus route for four weeks; ridership rose; the trial was too short to tell whether the increase would last; no launch date was announced. One candidate preserves those facts in two sentences. Another announces a permanent launch date. A third is faithful but uses five sentences. A fourth refuses unnecessarily.
Which should win? “Sounds confident” and “preserves the source” do not choose the same answer. A preference label only becomes meaningful alongside the prompt and judging rubric.
Four ranked responses yield six pair comparisons, but not six independent prompts. Keep all variants of a prompt on the same side of the train/test split. Record ties and disagreement. “Chosen” in a dataset means preferred under a judgment, not guaranteed true.
A learned judge
Learning a judge
To reuse those judgments, train a reward model. It reads the prompt and response and returns a scalar. The training loss asks the chosen response to score higher than the rejected one. A common conversion from scores to a preference probability is
Equal scores give 50–50. A difference of one gives about 73.1%. Only the difference matters: scores 100 and 99 express the same pair probability as 1 and 0. The scale is not a meter of truth.
Try it · toy model
Judge candidate summaries and fit a small reward model. Change the rubric and watch the ranking change.
Label a pair, then fit the judge. For a sharper demonstration, open the two rubric presets. Faithfulness first favors the grounded summary. Confidence first rewards the invented conclusion. The same fitting procedure succeeds at learning either set of labels. This tiny model fits four free scores; a real reward model must also generalize to unseen language.
Optimization
Reward with restraint
Now let the assistant generate its own responses. The reward model scores them; reinforcement learning changes the policy to favor higher-scoring responses. In the InstructGPT recipe, PPO performs those updates after SFT. This is a common form of reinforcement learning from human feedback, RLHF.
There is a catch. A model actively searching for high reward can discover mistakes that a passive validation set did not expose. One restraint is to penalize departure from a frozen reference, commonly the SFT model:
The KL term says how much the distribution changed. It cannot say whether a claim is true. Our next lab makes the difference measurable: its flawed judge rewards the confident false response most highly. At β=1, optimization raises that response's probability from 20% to about 53%.
Try it · toy model
Optimize a three-answer policy and discover why more reward can mean less truth.
Lower the restraint. Reward rises while correctness falls. Then penalize the unsupported claim: changing the judge changes what optimization achieves. This is an exact three-action calculation, not a PPO simulation or an empirical LLM result.
Deep diveWhy PPO has another old modelShould know
PPO compares a token's new probability with its probability under the policy that collected the rollout. Its clipped objective limits the incentive for large updates to that sampled policy. This old rollout policy is distinct from the reference policy used for the KL penalty. The rollout policy refreshes; the reference can remain fixed. Clipping is not a guarantee of a maximum policy distance or safe behavior.
A shorter route
The DPO shortcut
The reward-model-plus-RL route is not the only way to use preferences. Direct Preference Optimization, DPO, fits the language model directly to chosen/rejected pairs, using a fixed reference to measure how each response's probability changes.
The arithmetic fits on a napkin. The reference log probabilities are −4 and −3. The policy's are −3.5 and −3.2. The chosen answer gained 0.5 log units; the rejected answer lost 0.2. The relative difference is 0.7. Multiply by β=0.5, apply a sigmoid and take negative log: the pair loss is 0.533, down from 0.693 at the reference.
Try it · toy model
Follow chosen and rejected sequence probabilities through a reference-relative preference loss.
Try Both probabilities fall. DPO can improve the comparison while making both answers absolutely less likely, because the rejected answer loses more relative to its reference. It does not independently promise to raise every chosen answer's probability.
- 01
Show an answer
Demonstrations train an SFT policy. Keep a frozen reference copy.
- 02
Compare answers
Judges rank candidates. A reward model learns to predict those comparisons.
- 03
Sample and improve
The policy generates responses. PPO uses reward and a penalty for drifting from the reference.
These are common teaching pipelines, not mandatory stages for every assistant. DPO-based systems can still use a reward model elsewhere, for example to select demonstrations.
Vanilla DPO avoids a separately trained reward model and online sampling during preference fitting. A whole training pipeline can still use judges elsewhere. Llama 3's report describes reward-based candidate selection alongside SFT and DPO. Do not confuse one objective with the entire development process.
The source of judgment
Whose feedback?
Human comparisons take time. Models can help write demonstrations, critique drafts or compare responses. The important question is which signal they supply and how it is checked.
Constitutional AI uses explicit principles to guide critique and revision, and AI-generated preferences for a feedback stage. Humans still choose the principles and judging setup. AI feedback shifts annotation work; it does not eliminate human choices or automatically create reliable ground truth.
Another operation is to generate several candidates, select good ones and train on the selected responses. This rejection sampling supplies an SFT dataset; it is different from optimizing a policy on sampled rewards. It also explains why inference efficiency matters while building a model: the training pipeline may generate and judge many answers before a user ever sees one.
“Pretraining gives knowledge; post-training only changes style” is a useful first approximation that becomes misleading when treated as a law Interpretation. Reported recipes such as Tulu 3 target multiple skills and combine SFT, DPO and verifiable rewards. Chapter 14 takes up reasoning-oriented reinforcement learning; here, measure the task changes instead of assuming they must be cosmetic.
Behavioral boundaries
Safety is more than refusal
A system that always says no would pass a test consisting only of requests it should decline. It would also fail as an assistant. Safety tuning needs examples of appropriate responses and appropriate refusals, including benign requests that resemble problematic ones.
Helpfulness, truthfulness and harmlessness are related but separate aims. A polite hallucination, an accurate but inappropriate response and an unnecessary refusal fail in different ways. Training examples and preferences can shape these behaviors; they do not replace the wider system checks covered in Chapter 16.
The exam
Test the assistant
After each stage, ask what actually improved. Keep a development set for choosing the recipe and an unseen test set for the eventual claim. Use prompt-level splits, blind model identity in comparisons, randomize answer order and report uncertainty rather than just a win percentage.
| Check | A concrete failure |
|---|---|
| Instruction following | The summary has five sentences instead of two. |
| Faithfulness | It invents a launch date. |
| Uncertainty | A short trial becomes proof of permanent demand. |
| Appropriate refusal | It refuses a harmless transport summary. |
| Retained ability | A familiar factual or coding task regresses after tuning. |
Reward is useful training feedback, but it is not the exam. The reward-overoptimization paper studies divergence from a synthetic gold reward; work on sycophancy investigates feedback that favors agreement with a user's belief over a more accurate response. Neither finding justifies a universal claim that preference training always fails. They tell us what to test.
For an engineer, this chapter is a reminder that schemas, labels and evaluation splits are part of model behavior. For a researcher, it exposes a coupled learning problem: the policy changes, the judge is imperfect, and the metric can become a target. More optimization is only helpful when the signal continues to represent what you want.
What comes next
The assistant still has to run
The output is another checkpoint. It still predicts tokens, but demonstrations and feedback have shaped which continuations it favors. Deploying it creates a different set of problems: growing context memory, serial token generation, many simultaneous requests and the cost of adapting large weights. Chapter 11 goes inside the engineering that makes those systems practical.
Concepts in this chapter
Mark each one as you go. Must-know concepts are the core path.
- Chat TemplatesA chat template serializes roles and messages into the token sequence a particular model expects.UnderstandMust know
- Direct Preference OptimizationDPO trains a language model directly on preferred and rejected responses using their log-probability changes relative to a reference model.ImplementMust know
- Instruction and Demonstration DataDemonstration data pairs tasks or conversations with target responses, teaching the behavior that supervised fine-tuning should reproduce.Know wellMust know
- Evaluating Post-TrainingEvaluate a post-trained assistant on held-out task success, factuality, uncertainty, safety and retained capability, separately from its training reward.Know wellMust know
- Learning from ComparisonsPreference data records which of two responses a judge favors for the same prompt under a stated rubric.Know wellMust know
- Reward ModelsA reward model learns a scalar score that predicts which response a judge will prefer for a prompt.ImplementMust know
- Reinforcement Learning from Human FeedbackRLHF improves a policy using a reward learned from human feedback, often with a penalty for departing from a reference model.Know wellMust know
- Safety Tuning and RefusalSafety tuning uses examples and feedback to shape when an assistant should answer, decline or offer a safer alternative.UnderstandMust know
- Supervised Fine-TuningSupervised fine-tuning continues a pretrained model's training on demonstrations of the responses it should produce.ImplementMust know
- Constitutional AI and AI FeedbackExplicit principles can guide model critiques and revisions and help generate preference feedback for later training.UnderstandShould know
- PPO for Language ModelsPPO uses a clipped policy objective to limit the incentive for large changes relative to the policy that collected a rollout.Know wellShould know
- Generate, Rank, Then TrainGenerate several responses, select candidates with a judge, and use the selected responses as supervised demonstrations.UnderstandShould know
What do I actually need to remember?
- Base and assistant models can share the same next-token generation mechanism.
- SFT updates weights using demonstrations; examples in a prompt do not.
- A loss mask selects training targets without removing the prompt from context.
- Preferences depend on a prompt, judge and rubric.
- Reward-model scores are proxies, not truth.
- RLHF can optimize sampled responses with a penalty for drifting from a reference.
- PPO’s rollout policy and the KL reference serve different purposes.
- DPO fits chosen/rejected pairs relative to a reference; the whole pipeline may still use judges.
- AI feedback still depends on human choices about principles, models and data.
- Evaluate task success, factuality, safety and retained capability separately.
You do not need to memorize everything else. This list is the revision sheet.
Key papers
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu et al. · 2022 · NeurIPS 2022
InstructGPT: the supervised fine-tuning + reward model + RL recipe that turned GPT-3 into an instruction-following assistant, and the template for ChatGPT.
- Problem
- Pretrained language models continue text; they don't reliably follow instructions or behave helpfully.
- What was new
- Fine-tune on human demonstrations, train a reward model on human rankings, then optimize the model against it with PPO.
How to read it: Figure 2 is the three-step RLHF pipeline you'll meet in Chapter 10.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafael Rafailov, Archit Sharma et al. · 2023
Provides a direct preference-training route with fewer moving parts.
- Problem
- A separate reward-model and online RL loop is operationally complex.
- What was new
- Derive a reference-relative policy objective fitted directly to preference pairs.
How to read it: Start with equations 3 and 7; inspect the assumptions behind the policy/reward connection.
Deep reinforcement learning from human preferences
Paul Christiano, Jan Leike et al. · 2017 · NeurIPS 2017
Showed that agents can be trained from human comparisons between behaviours rather than a hand-written reward — the foundation of RLHF.
- Problem
- For many tasks, writing a reward function that captures what we want is hard or impossible.
- What was new
- Learn a reward model from pairwise human preferences, then optimize a policy against it with reinforcement learning.
Proximal Policy Optimization Algorithms
John Schulman, Filip Wolski et al. · 2017
PPO: a simple, robust actor–critic policy-gradient method. It became the default RL algorithm in many labs and was the optimiser in InstructGPT-style RLHF.
- Problem
- Policy-gradient updates that are too large can wreck a policy in a single step; the principled fixes were complicated.
- What was new
- A clipped objective that removes the incentive to move the policy too far from the one that collected the data, so the same batch can be reused for several updates.
Finetuned Language Models Are Zero-Shot Learners
Jason Wei, Maarten Bosma et al. · 2021 · ICLR 2022
Instruction tuning: fine-tune on many tasks phrased as instructions and the model follows instructions for new tasks too. The 137B FLAN beat zero-shot GPT-3 on 20 of 25 tasks.
- Problem
- Base models were good at few-shot prompting but weak at simply following an instruction with no examples.
- What was new
- Fine-tune a pretrained model on over 60 datasets rewritten as natural-language instructions, then test on unseen task types.
Learning to summarize from human feedback
Nisan Stiennon, Long Ouyang et al. · 2020
Connects preference learning to generated language before instruction-following assistants.
- Problem
- Automatic summary metrics miss qualities that people care about.
- What was new
- Learn a reward from human summary comparisons and optimize a summarization policy.
How to read it: Read the comparison setup and human evaluation before the aggregate scores.
Multitask Prompted Training Enables Zero-Shot Task Generalization
Victor Sanh, Albert Webson et al. · 2021
Shows why task diversity and prompt design belong in the instruction-data story.
- Problem
- Supervised models often specialize in the tasks used for training.
- What was new
- Train on many tasks expressed through prompts and evaluate generalization to held-out tasks.
How to read it: Compare task splits and prompt formulations with FLAN.
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Yuntao Bai, Andy Jones et al. · 2022
Makes feedback data and the helpfulness/harmlessness tension concrete.
- Problem
- Helpfulness and harmlessness are not one easily specified objective.
- What was new
- Train assistants with preference feedback and examine behavior along these separate dimensions.
How to read it: Inspect the data and evaluation categories; do not reduce safety to refusal rate.
Scaling Laws for Reward Model Overoptimization
Leo Gao, John Schulman, Jacob Hilton · 2022
Explains why maximizing a learned judge can diverge from the intended outcome.
- Problem
- A policy can exploit errors in a proxy reward model.
- What was new
- Study optimization against a proxy while measuring a synthetic gold reward.
How to read it: Read the synthetic-gold-reward setup before interpreting the scaling curves.
Constitutional AI: Harmlessness from AI Feedback
Yuntao Bai, Saurav Kadavath et al. · 2022
A documented route for scaling some feedback while retaining human-selected principles.
- Problem
- Direct human feedback is costly and principles need consistent application.
- What was new
- Use principle-guided critique/revision and AI-generated preferences in separate training stages.
How to read it: Distinguish the supervised revision stage from the feedback/RL stage.
Self-Instruct: Aligning Language Models with Self-Generated Instructions
Yizhong Wang, Yeganeh Kordi et al. · 2022
Makes synthetic instruction generation an explicit data pipeline to evaluate.
- Problem
- Writing enough diverse instruction examples by hand is expensive.
- What was new
- Generate and filter synthetic instructions and instances from a language model.
How to read it: Trace seed examples, generation, filtering and held-out evaluation separately.
LIMA: Less Is More for Alignment
Chunting Zhou, Pengfei Liu et al. · 2023
A bounded case study of data quality, not a universal sample-size rule.
- Problem
- How far can carefully curated supervised examples shape a strong base model?
- What was new
- Study a 65B model fine-tuned on 1,000 curated examples without reinforcement learning.
How to read it: Separate the reported results from broader interpretations about where capability comes from.
UltraFeedback: Boosting Language Models with Scaled AI Feedback
Ganqu Cui, Lifan Yuan et al. · 2023
A concrete source of synthetic feedback used in open assistant research.
- Problem
- Feedback annotation needs scale and more than one quality dimension.
- What was new
- Collect model responses with AI-generated multidimensional feedback.
How to read it: Distinguish original annotations from a binarized derivative dataset.
Towards Understanding Sycophancy in Language Models
Mrinank Sharma, Meg Tong et al. · 2023
Separates being preferred from being accurate.
- Problem
- A response can please a user by agreeing with a mistaken belief.
- What was new
- Investigate sycophancy and how preference judgments can reward it.
How to read it: Inspect task construction and judge behavior before generalizing to all feedback-trained models.
Zephyr: Direct Distillation of LM Alignment
Lewis Tunstall, Edward Beeching et al. · 2023
A practical application of DPO and synthetic data.
- Problem
- Open assistant training needs an effective, reproducible preference route.
- What was new
- Combine distilled instruction data and direct preference optimization.
How to read it: Follow the data stages; do not attribute the entire outcome to one optimizer.
The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey et al. · 2024
The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.
- Problem
- Frontier labs had largely stopped publishing how their models were built.
- What was new
- A 405B model trained on 15.6T tokens with 3.8 × 10²⁵ FLOPs, with detailed data cleaning, a data mix chosen by scaling experiments, 4D parallelism at 38–43% utilisation, and 466 job interruptions in a 54-day period.
How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.
Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Nathan Lambert, Jacob Morrison et al. · 2024
Provides an open case study connecting data, objectives and evaluation.
- Problem
- Strong post-training recipes are often difficult to inspect or reproduce.
- What was new
- Release a pipeline combining SFT, DPO, verifiable rewards and explicit development/unseen evaluation.
How to read it: Read the evaluation split and training stages; leave reasoning RL detail for Chapter 14.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 15: Alignment - SFT/RLHF
A technical companion to the chapter’s demonstration, preference and reinforcement-learning pipeline.
Covers: Supervised fine-tuning and alignment through human feedback; connects data, reward and policy optimization.
Andrej Karpathy
Deep Dive into LLMs like ChatGPT
A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.
Covers: The full training stack behind a chat model, from internet text and tokens to a pretrained base model and the post-training that turns it into an assistant.
What came next?
Chapter 11
The original Transformer is too slow and memory-hungry to serve at modern scale.