Skip to content
Road to Intelligence

Concept · Chapter 10: From Base Model to Assistant

Supervised Fine-Tuning

Must knowImplement15 minDifficulty

Supervised fine-tuning continues a pretrained model's training on demonstrations of the responses it should produce.

The problem

Predicting web text does not specify which role to play, which instructions to follow, or what a useful answer looks like.

The solution

Train on prompt-response demonstrations with the usual next-token loss, selecting the intended target tokens with a loss mask.

The consequence

The model becomes more likely to produce demonstrated behavior, but its coverage and errors still reflect the data.

Show what success looks like

Give a model a task and a good answer, then train it to predict that answer. A summary demonstrates what to keep; a code answer demonstrates a function and its explanation; a conversation demonstrates when to respond. This is ordinary supervised learning applied to a pretrained language model.

Instruction tuning is a common form of SFT: tasks are expressed as instructions. FLAN and T0 explored whether training on many such tasks transfers to instructions from held-out task families. That split matters: remembering examples is easier than following new kinds of request.

Same loss, different targets

Imagine the sequence 2 + 2 ? Four . END. In a response-only setup, the first four tokens provide context and the last three receive loss. Teacher forcing supplies the dataset prefix: when predicting the period, the model is given Four, even if it would have sampled a different word.

For an authored example, probabilities of 0.50, 0.25 and 0.80 for those three targets give losses of 0.693, 1.386 and 0.223 nats. Their mean is 0.768 nats. Changing a masked prompt probability leaves that mean unchanged.

LSFT=−∑tmtlog⁡pθ(zt∣z<t)∑tmt\mathcal L_{\mathrm{SFT}}=-\frac{\sum_t m_t\log p_\theta(z_t\mid z_{<t})}{\sum_t m_t}

Here mtm_t is 1 for a supervised target and 0 otherwise. A mask with no targets has no mean loss. This formula shows one sequence's mean; a training system must also choose how to aggregate across batches and lengths.

Try it · toy model

What Does the Loss See?

Select training targets, change their probabilities and see exactly what a loss mask removes.

Implement10 min

The mask removes a direct loss term, not the prompt's role in the computation. The response predictions still depend on prompt representations, so response loss can send gradients through them.

What changed, and what did not

SFT changes weights. In-context learning changes the supplied context without updating them. Both still use the causal next-token mechanism.

Post-training can change task performance as well as presentation. The useful question is empirical: which held-out abilities improved, which regressed, and under what evaluation? A narrow demonstration set can teach a format without robustly teaching the underlying task.

Try the distinction

In the lab, use Assistant only, move the prompt slider, then switch to All tokens. Explain why the first move leaves the loss unchanged and the second changes what the model is asked to learn. Finally clear the mask: no selected targets means no learning signal for this example.

Why should I care?

As a researcher

The same architecture can behave very differently after a change in data and objective; isolate those effects rather than attributing them all to scale.

As an engineer

Incorrect masks, role formatting and target examples can quietly train the wrong behavior even when the loss falls.

Modern systems that depend on it

  • instruction-following assistants
  • domain adaptation
  • preference training

Historical context

Before

A broad pretrained model can continue many kinds of document, including conversations and quizzes, without consistently taking the assistant's role.

After

The same next-token architecture is trained to give the desired response to a task or conversation.

Used today

SFT appears in the reported InstructGPT, Llama 3 and Tulu 3 post-training recipes, using human and model-generated demonstrations.

What to remember

  • SFT updates weights; examples in a prompt do not.
  • Teacher forcing supplies earlier target tokens during training.
  • A masked prompt still conditions the response, even though it contributes no direct loss term.
  • A falling training loss does not establish held-out factuality or safety.

Key papers

Essential

Training language models to follow instructions with human feedback

Long Ouyang, Jeff Wu et al. · 2022 · NeurIPS 2022

InstructGPT: the supervised fine-tuning + reward model + RL recipe that turned GPT-3 into an instruction-following assistant, and the template for ChatGPT.

How to read it: Figure 2 is the three-step RLHF pipeline you'll meet in Chapter 10.

~1 h readarXiv:2203.02155✓ verified 2026-09-26
Important

Finetuned Language Models Are Zero-Shot Learners

Jason Wei, Maarten Bosma et al. · 2021 · ICLR 2022

Instruction tuning: fine-tune on many tasks phrased as instructions and the model follows instructions for new tasks too. The 137B FLAN beat zero-shot GPT-3 on 20 of 25 tasks.

~35 min readarXiv:2109.01652✓ verified 2026-09-26
Optional

LIMA: Less Is More for Alignment

Chunting Zhou, Pengfei Liu et al. · 2023

A bounded case study of data quality, not a universal sample-size rule.

How to read it: Separate the reported results from broader interpretations about where capability comes from.

~25 min readarXiv:2305.11206✓ verified 2026-10-04
Important

Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Nathan Lambert, Jacob Morrison et al. · 2024

Provides an open case study connecting data, objectives and evaluation.

How to read it: Read the evaluation split and training stages; leave reasoning RL detail for Chapter 14.

~40 min readarXiv:2411.15124✓ verified 2026-10-04

Watch

3 h 31 min

Andrej Karpathy

Deep Dive into LLMs like ChatGPT

A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.

Should know