Skip to content
Road to Intelligence

Concept · Chapter 11: Inside Modern LLMs

Knowledge Distillation

Should knowUnderstand9 minDifficulty

Knowledge distillation trains a small 'student' model to match a large 'teacher' model's output probabilities, which carry more information than the correct answers alone, so the student gets closer to the teacher than training on labels would allow.

The problem

Big models predict better but cost too much to run everywhere; a small model trained from scratch on the same labels is noticeably worse.

The solution

Train the student on the teacher's full output distribution (soft targets), often softened with a temperature, usually alongside the ordinary loss on the true labels.

The consequence

Distillation is a standard way to make cheaper models. The student is limited by the teacher, and by the data on which the teacher's outputs are collected.

You should understand first

  1. Probability and Distributions
  2. Entropy
  3. Softmax
  4. Loss Functions
  5. Cross-Entropy Loss
  6. KL Divergence
  7. Knowledge Distillation

More than the right answer

A trained classifier asked about a picture of a 2 might put 0.9 on "2", 0.09 on "7" and almost nothing on "cat". The label says only "2". The teacher's distribution also says that this 2 looks a little like a 7, which is information about the task the label doesn't carry.

Hinton, Vinyals and Dean proposed raising the temperature of the teacher's softmax until it produces a suitably soft set of targets, and training the small model with the same high temperature to match them Established. The training loss is the KL divergence (equivalently, cross-entropy) between teacher and student distributions, usually combined with the ordinary loss on the true labels.

Temperature, in numbers

qi=exp⁡(zi/T)∑jexp⁡(zj/T)q_i = \frac{\exp(z_i / T)}{\sum_j \exp(z_j / T)}

Tiny example. Teacher logits z=[4,2,0]z = [4, 2, 0].

  • At T=1T = 1: [0.867,0.117,0.016][0.867, 0.117, 0.016]. The third class is nearly invisible.
  • At T=4T = 4: [0.507,0.307,0.186][0.507, 0.307, 0.186]. The ranking is the same, but the student now sees how the wrong answers compare.

This is the same temperature as in decoding, used for a different purpose.

For language models

DistilBERT applied distillation during pretraining and produced a model 40% smaller and 60% faster than BERT that retained 97% of its language-understanding performance Established. For generative LLMs the idea extends to training a student on the teacher's next-token distributions, or more simply on text the teacher generates.

What to remember

  • Soft targets: the teacher's whole probability distribution, not just its top answer.
  • Temperature T > 1 flattens the distribution so small probabilities carry signal.
  • Loss: KL divergence (or cross-entropy) between teacher and student distributions, at the same T.
  • DistilBERT: 40% smaller, 60% faster, 97% of BERT's language-understanding score.

Key papers

Essential

Distilling the Knowledge in a Neural Network

Geoffrey Hinton, Oriol Vinyals, Jeff Dean · 2015

Defined knowledge distillation: train a small student on a large teacher's softened output probabilities, which carry more information than hard labels.

How to read it: Section 2 explains temperature and soft targets in a page.

~25 min readarXiv:1503.02531✓ verified 2026-10-05

Watch