Concept · Chapter 10: From Base Model to Assistant
Supervised Fine-Tuning
Supervised fine-tuning continues a pretrained model's training on demonstrations of the responses it should produce.
The problem
Predicting web text does not specify which role to play, which instructions to follow, or what a useful answer looks like.
The solution
Train on prompt-response demonstrations with the usual next-token loss, selecting the intended target tokens with a loss mask.
The consequence
The model becomes more likely to produce demonstrated behavior, but its coverage and errors still reflect the data.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Pretrain, Then Fine-Tune
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Supervised Fine-Tuning
Show what success looks like
Give a model a task and a good answer, then train it to predict that answer. A summary demonstrates what to keep; a code answer demonstrates a function and its explanation; a conversation demonstrates when to respond. This is ordinary supervised learning applied to a pretrained language model.
Instruction tuning is a common form of SFT: tasks are expressed as instructions. FLAN and T0 explored whether training on many such tasks transfers to instructions from held-out task families. That split matters: remembering examples is easier than following new kinds of request.
Same loss, different targets
Imagine the sequence 2 + 2 ? Four . END. In a response-only setup, the first four tokens provide context and the last three receive loss. Teacher forcing supplies the dataset prefix: when predicting the period, the model is given Four, even if it would have sampled a different word.
For an authored example, probabilities of 0.50, 0.25 and 0.80 for those three targets give losses of 0.693, 1.386 and 0.223 nats. Their mean is 0.768 nats. Changing a masked prompt probability leaves that mean unchanged.
Here is 1 for a supervised target and 0 otherwise. A mask with no targets has no mean loss. This formula shows one sequence's mean; a training system must also choose how to aggregate across batches and lengths.
Try it · toy model
Select training targets, change their probabilities and see exactly what a loss mask removes.
The mask removes a direct loss term, not the prompt's role in the computation. The response predictions still depend on prompt representations, so response loss can send gradients through them.
What changed, and what did not
SFT changes weights. In-context learning changes the supplied context without updating them. Both still use the causal next-token mechanism.
Post-training can change task performance as well as presentation. The useful question is empirical: which held-out abilities improved, which regressed, and under what evaluation? A narrow demonstration set can teach a format without robustly teaching the underlying task.
Try the distinction
In the lab, use Assistant only, move the prompt slider, then switch to All tokens. Explain why the first move leaves the loss unchanged and the second changes what the model is asked to learn. Finally clear the mask: no selected targets means no learning signal for this example.
Why should I care?
As a researcher
The same architecture can behave very differently after a change in data and objective; isolate those effects rather than attributing them all to scale.
As an engineer
Incorrect masks, role formatting and target examples can quietly train the wrong behavior even when the loss falls.
Modern systems that depend on it
- instruction-following assistants
- domain adaptation
- preference training
Historical context
Before
A broad pretrained model can continue many kinds of document, including conversations and quizzes, without consistently taking the assistant's role.
After
The same next-token architecture is trained to give the desired response to a task or conversation.
Used today
SFT appears in the reported InstructGPT, Llama 3 and Tulu 3 post-training recipes, using human and model-generated demonstrations.
What to remember
- SFT updates weights; examples in a prompt do not.
- Teacher forcing supplies earlier target tokens during training.
- A masked prompt still conditions the response, even though it contributes no direct loss term.
- A falling training loss does not establish held-out factuality or safety.
Key papers
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu et al. · 2022 · NeurIPS 2022
InstructGPT: the supervised fine-tuning + reward model + RL recipe that turned GPT-3 into an instruction-following assistant, and the template for ChatGPT.
How to read it: Figure 2 is the three-step RLHF pipeline you'll meet in Chapter 10.
Finetuned Language Models Are Zero-Shot Learners
Jason Wei, Maarten Bosma et al. · 2021 · ICLR 2022
Instruction tuning: fine-tune on many tasks phrased as instructions and the model follows instructions for new tasks too. The 137B FLAN beat zero-shot GPT-3 on 20 of 25 tasks.
LIMA: Less Is More for Alignment
Chunting Zhou, Pengfei Liu et al. · 2023
A bounded case study of data quality, not a universal sample-size rule.
How to read it: Separate the reported results from broader interpretations about where capability comes from.
Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Nathan Lambert, Jacob Morrison et al. · 2024
Provides an open case study connecting data, objectives and evaluation.
How to read it: Read the evaluation split and training stages; leave reasoning RL detail for Chapter 14.
Watch
Andrej Karpathy
Deep Dive into LLMs like ChatGPT
A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 15: Alignment - SFT/RLHF
A technical companion to the chapter’s demonstration, preference and reinforcement-learning pipeline.