Concept · Chapter 14: Reasoning Models
Synthetic Reasoning Data and Distillation
Reasoning distillation fine-tunes a model on solutions that a stronger or more expensive system generated and that passed a check, moving inference effort upstream into training data.
The problem
A strong reasoning system is too slow or expensive to serve everywhere, and hand-written step-by-step solutions are scarce.
The solution
Run the expensive system on many questions, keep the solutions that verify, and use them as supervised training data for a cheaper model, or for the same model, repeatedly.
The consequence
Smaller models gain much of the teacher's behaviour on similar tasks, and the filter, the deduplication and the evaluation split decide what exactly they learn.
You should understand first
- Probability and Distributions
- Entropy
- Softmax
- Loss Functions
- Cross-Entropy Loss
- KL Divergence
- Knowledge Distillation
- Text as Data
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Pretrain, Then Fine-Tune
- Supervised Fine-Tuning
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Autoregressive Next-Token Prediction
- Decoding: Greedy, Temperature, Top-k, Top-p
- Learning from Comparisons
- Features, Labels and Tasks
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Logistic Regression
- Reward Models
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- Chain of Thought
- Expected Value and Variance
- Sampling and Uncertainty
- Self-Consistency
- Verifiers and Best-of-N
- Synthetic Reasoning Data and Distillation
Spend once, teach many times
A teacher system can be slow: a large model, sampling 64 attempts, searching, running tests. Run it once over a large question set, keep what checks out, and you have training data. A student fine-tuned on that data may then answer similar questions with one cheap attempt. The expensive work has moved from every request to one dataset build.
This is knowledge distillation in a broad sense. Classic distillation trains a student to match a teacher's output probabilities. Established Reasoning distillation usually trains only on the teacher's sampled text, with ordinary supervised fine-tuning; the student never sees the teacher's probabilities.
A model can be its own teacher
STaR (the Self-Taught Reasoner) generated rationales for many questions using a few examples, retried failures with the correct answer given as a hint, fine-tuned on the rationales that ended in correct answers, and repeated; it improved over a model fine-tuned to predict answers directly on its datasets. EstablishedThe retry with a hint, which the paper calls rationalization, matters: without it the model only learns from problems it could already solve.
The DeepSeek-R1 report described fine-tuning smaller open models on reasoning data generated with R1, and reported strong reasoning-benchmark results for these distilled models. EstablishedThe pipeline, as a data engineer would build it
- Generate. Record the teacher checkpoint, prompt, sampling settings, tools and seed for every example.
- Verify. Check answers, run tests. Store the verdict, not just the survivors.
- Filter. Decide what passes. Final-answer filters keep lucky derivations (see the reward lab); step checks or a process scorer are stricter but costlier.
- Deduplicate. Near-duplicate problems inflate apparent learning.
- Split. Keep evaluation questions, and paraphrases of them, out of generation prompts and training data. This is contamination control.
- Fine-tune the student.
- Evaluate the student against its own pre-training baseline, at the same inference budget, on held-out task families.
Tiny example
The teacher answers 10,000 problems with 8 samples each. 62% of problems have at least one sample that passes the answer check; keeping one passing solution per problem gives 6,200 examples. If a step checker would reject 10% of those as lucky, the strict dataset has 5,580. Is the cleaner, smaller set better? Only an evaluation on held-out problems can say. (Illustrative numbers.)
Limits
- The student learns what is in the data. Problems the teacher never solved contribute nothing.
- Imitating long, confident text is not proof of the underlying skill; a student can copy the style of self-checking without checking.
- A distilled student can still need long outputs, so it is not automatically cheap to serve.
- Teacher accuracy, student accuracy, dataset cost and serving cost are four different numbers. Report all four.
Mini experiment
Take ten problems you can check. For each, write two candidate solutions: one valid and one that reaches the right answer by a flaw. Apply an answer-only filter and a step-by-step filter. How many flawed solutions would a student see under each, and what would you expect it to learn from them?
What to remember
- Pipeline: generate → verify → filter → deduplicate → split → fine-tune → evaluate.
- The teacher can be a model plus search, tools and many samples; record which system produced each example.
- STaR: keep your own rationales that reach correct answers, fine-tune, repeat.
- Filtering on final answers keeps lucky derivations too.
- Compare the student before and after at the same inference budget, on held-out tasks.
Key papers
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI et al. · 2025
An openly released reasoning model, with a detailed account of training long chains of reasoning mainly through reinforcement learning on verifiable problems.
How to read it: Read the R1-Zero and R1 sections separately: the first is RL straight from a base model, the second a multi-stage recipe with cold-start data, and the distillation results are a third story.
Distilling the Knowledge in a Neural Network
Geoffrey Hinton, Oriol Vinyals, Jeff Dean · 2015
Defined knowledge distillation: train a small student on a large teacher's softened output probabilities, which carry more information than hard labels.
How to read it: Section 2 explains temperature and soft targets in a page.
STaR: Bootstrapping Reasoning With Reasoning
Eric Zelikman, Yuhuai Wu et al. · 2022 · NeurIPS 2022
An early, clear version of the loop behind synthetic reasoning data: a model generates rationales, keeps the ones that reach correct answers, and trains on them.
How to read it: Notice the filter is the final answer only, so a rationale that reaches the right answer by a flawed route is kept; compare Chapter 14's lucky answer.