Skip to content
Road to Intelligence

Concept · Chapter 10: From Base Model to Assistant

Instruction and Demonstration Data

Must knowKnow well9 minDifficulty

Demonstration data pairs tasks or conversations with target responses, teaching the behavior that supervised fine-tuning should reproduce.

The problem

A broad web corpus contains many behaviors, only some of which are appropriate for an assistant.

The solution

Curate diverse, correct prompt-response demonstrations and separate training prompts from evaluation task families.

The consequence

Data coverage, formatting and judgment become part of the model's behavior rather than merely preprocessing details.

A dataset is a behavioral specification

“Summarize this paragraph” needs a paragraph and a target summary. “Return valid JSON” needs examples that really obey the schema. Multi-turn dialogue needs a consistent role sequence and answers that use earlier context. The label is an entire response, not one category.

A useful data record keeps the source, prompt, response, task family, creation method and quality checks together. Those fields make it possible to audit coverage, remove duplicates and keep evaluation prompts out of the training mixture. Model-generated variants of one seed prompt should not casually cross the split.

Quantity is not the only dial

FLAN and T0 test instruction generalization by holding out tasks. Self-Instruct generates and filters new instructions using a model. These approaches ask different questions: whether task diversity transfers, and whether synthetic data can expand coverage.

LIMA reports tuning a 65B model on 1,000 curated examples without reinforcement learning Established. Its result is a case study, not a universal sample-size rule. The base checkpoint, curation, task distribution and evaluation all matter.

A small audit before training

Take ten records. Check whether the response follows every explicit constraint, preserves uncertainty and avoids unsupported facts. Check that the role formatting matches the model. Then look across records: are nearly all answers long, apologetic or overly confident? Repeated stylistic choices can become unintended lessons.

What to remember

  • A demonstration is an example to imitate, not a ranking of every alternative.
  • Split prompts and task families before using variants in training.
  • Synthetic responses need quality checks too.

Key papers

Important

Finetuned Language Models Are Zero-Shot Learners

Jason Wei, Maarten Bosma et al. · 2021 · ICLR 2022

Instruction tuning: fine-tune on many tasks phrased as instructions and the model follows instructions for new tasks too. The 137B FLAN beat zero-shot GPT-3 on 20 of 25 tasks.

~35 min readarXiv:2109.01652✓ verified 2026-09-26
Important

Multitask Prompted Training Enables Zero-Shot Task Generalization

Victor Sanh, Albert Webson et al. · 2021

Shows why task diversity and prompt design belong in the instruction-data story.

How to read it: Compare task splits and prompt formulations with FLAN.

~30 min readarXiv:2110.08207✓ verified 2026-10-04
Optional

LIMA: Less Is More for Alignment

Chunting Zhou, Pengfei Liu et al. · 2023

A bounded case study of data quality, not a universal sample-size rule.

How to read it: Separate the reported results from broader interpretations about where capability comes from.

~25 min readarXiv:2305.11206✓ verified 2026-10-04
Important

Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Nathan Lambert, Jacob Morrison et al. · 2024

Provides an open case study connecting data, objectives and evaluation.

How to read it: Read the evaluation split and training stages; leave reasoning RL detail for Chapter 14.

~40 min readarXiv:2411.15124✓ verified 2026-10-04

Watch