Concept · Chapter 11: Inside Modern LLMs
Knowledge Distillation
Knowledge distillation trains a small 'student' model to match a large 'teacher' model's output probabilities, which carry more information than the correct answers alone, so the student gets closer to the teacher than training on labels would allow.
The problem
Big models predict better but cost too much to run everywhere; a small model trained from scratch on the same labels is noticeably worse.
The solution
Train the student on the teacher's full output distribution (soft targets), often softened with a temperature, usually alongside the ordinary loss on the true labels.
The consequence
Distillation is a standard way to make cheaper models. The student is limited by the teacher, and by the data on which the teacher's outputs are collected.
You should understand first
- Probability and Distributions
- Entropy
- Softmax
- Loss Functions
- Cross-Entropy Loss
- KL Divergence
- Knowledge Distillation
More than the right answer
A trained classifier asked about a picture of a 2 might put 0.9 on "2", 0.09 on "7" and almost nothing on "cat". The label says only "2". The teacher's distribution also says that this 2 looks a little like a 7, which is information about the task the label doesn't carry.
Hinton, Vinyals and Dean proposed raising the temperature of the teacher's softmax until it produces a suitably soft set of targets, and training the small model with the same high temperature to match them Established. The training loss is the KL divergence (equivalently, cross-entropy) between teacher and student distributions, usually combined with the ordinary loss on the true labels.
Temperature, in numbers
Tiny example. Teacher logits .
- At : . The third class is nearly invisible.
- At : . The ranking is the same, but the student now sees how the wrong answers compare.
This is the same temperature as in decoding, used for a different purpose.
For language models
DistilBERT applied distillation during pretraining and produced a model 40% smaller and 60% faster than BERT that retained 97% of its language-understanding performance Established. For generative LLMs the idea extends to training a student on the teacher's next-token distributions, or more simply on text the teacher generates.
What to remember
- Soft targets: the teacher's whole probability distribution, not just its top answer.
- Temperature T > 1 flattens the distribution so small probabilities carry signal.
- Loss: KL divergence (or cross-entropy) between teacher and student distributions, at the same T.
- DistilBERT: 40% smaller, 60% faster, 97% of BERT's language-understanding score.
Key papers
Distilling the Knowledge in a Neural Network
Geoffrey Hinton, Oriol Vinyals, Jeff Dean · 2015
Defined knowledge distillation: train a small student on a large teacher's softened output probabilities, which carry more information than hard labels.
How to read it: Section 2 explains temperature and soft targets in a page.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut et al. · 2019
A widely used demonstration that distillation during pretraining produces a much cheaper general-purpose language model.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 10: Inference
The closest single lecture to this chapter: the arithmetic of serving a language model and the main ways to make it cheaper.