Concept · Chapter 10: From Base Model to Assistant
Instruction and Demonstration Data
Demonstration data pairs tasks or conversations with target responses, teaching the behavior that supervised fine-tuning should reproduce.
The problem
A broad web corpus contains many behaviors, only some of which are appropriate for an assistant.
The solution
Curate diverse, correct prompt-response demonstrations and separate training prompts from evaluation task families.
The consequence
Data coverage, formatting and judgment become part of the model's behavior rather than merely preprocessing details.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Pretrain, Then Fine-Tune
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Supervised Fine-Tuning
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Expected Value and Variance
- Sampling and Uncertainty
- Generalization, Overfitting and Underfitting
- Data Leakage
- Instruction and Demonstration Data
A dataset is a behavioral specification
“Summarize this paragraph” needs a paragraph and a target summary. “Return valid JSON” needs examples that really obey the schema. Multi-turn dialogue needs a consistent role sequence and answers that use earlier context. The label is an entire response, not one category.
A useful data record keeps the source, prompt, response, task family, creation method and quality checks together. Those fields make it possible to audit coverage, remove duplicates and keep evaluation prompts out of the training mixture. Model-generated variants of one seed prompt should not casually cross the split.
Quantity is not the only dial
FLAN and T0 test instruction generalization by holding out tasks. Self-Instruct generates and filters new instructions using a model. These approaches ask different questions: whether task diversity transfers, and whether synthetic data can expand coverage.
LIMA reports tuning a 65B model on 1,000 curated examples without reinforcement learning Established. Its result is a case study, not a universal sample-size rule. The base checkpoint, curation, task distribution and evaluation all matter.
A small audit before training
Take ten records. Check whether the response follows every explicit constraint, preserves uncertainty and avoids unsupported facts. Check that the role formatting matches the model. Then look across records: are nearly all answers long, apologetic or overly confident? Repeated stylistic choices can become unintended lessons.
What to remember
- A demonstration is an example to imitate, not a ranking of every alternative.
- Split prompts and task families before using variants in training.
- Synthetic responses need quality checks too.
Key papers
Finetuned Language Models Are Zero-Shot Learners
Jason Wei, Maarten Bosma et al. · 2021 · ICLR 2022
Instruction tuning: fine-tune on many tasks phrased as instructions and the model follows instructions for new tasks too. The 137B FLAN beat zero-shot GPT-3 on 20 of 25 tasks.
Multitask Prompted Training Enables Zero-Shot Task Generalization
Victor Sanh, Albert Webson et al. · 2021
Shows why task diversity and prompt design belong in the instruction-data story.
How to read it: Compare task splits and prompt formulations with FLAN.
Self-Instruct: Aligning Language Models with Self-Generated Instructions
Yizhong Wang, Yeganeh Kordi et al. · 2022
Makes synthetic instruction generation an explicit data pipeline to evaluate.
How to read it: Trace seed examples, generation, filtering and held-out evaluation separately.
LIMA: Less Is More for Alignment
Chunting Zhou, Pengfei Liu et al. · 2023
A bounded case study of data quality, not a universal sample-size rule.
How to read it: Separate the reported results from broader interpretations about where capability comes from.
Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Nathan Lambert, Jacob Morrison et al. · 2024
Provides an open case study connecting data, objectives and evaluation.
How to read it: Read the evaluation split and training stages; leave reasoning RL detail for Chapter 14.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 15: Alignment - SFT/RLHF
A technical companion to the chapter’s demonstration, preference and reinforcement-learning pipeline.