Concept · Chapter 10: From Base Model to Assistant
Generate, Rank, Then Train
Generate several responses, select candidates with a judge, and use the selected responses as supervised demonstrations.
The problem
A model may produce a good response only sometimes, and manually writing every demonstration is expensive.
The solution
Sample multiple candidates, rank or filter them, and train on the retained examples.
The consequence
Selection can improve the demonstration pool, but it inherits judge errors and costs additional generation.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Pretrain, Then Fine-Tune
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Supervised Fine-Tuning
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Autoregressive Next-Token Prediction
- Decoding: Greedy, Temperature, Top-k, Top-p
- Learning from Comparisons
- Features, Labels and Tasks
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Logistic Regression
- Reward Models
- Generate, Rank, Then Train
Keep the better attempts
Ask a model for several summaries of one paragraph. A judge ranks them; a filter rejects malformed or unsupported answers; retained responses enter an SFT dataset. The weight update then imitates the selected text.
This chapter uses “rejection sampling” in the post-training data-selection sense. It is not a claim that best-of-N selection is the classical exact rejection-sampling algorithm for drawing from an arbitrary target distribution.
Selection is a separate stage
Best-of-N can also select a response at inference without training. If the selected response becomes a later training example, that is an additional step. A system can combine this selection with DPO on a separate set of preference pairs.
The Llama 3 report describes precisely such an integrated pipeline: reward scoring helps select demonstrations, while DPO performs preference optimization. An objective and an entire model-development process should not be treated as synonyms.
More candidates, more opportunities
More attempts can uncover a useful response, but also a response that exploits the judge. Independently check selected outputs, and compare the gain with the cost of generation and scoring. Chapter 11's efficient inference matters here even before the final assistant is deployed.
What to remember
- Candidate selection and policy optimization are different operations.
- A judge's favorite is not necessarily the most accurate response.
- Generating more candidates increases inference cost during training-data creation.
Key papers
The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey et al. · 2024
The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.
How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.
Scaling Laws for Reward Model Overoptimization
Leo Gao, John Schulman, Jacob Hilton · 2022
Explains why maximizing a learned judge can diverge from the intended outcome.
How to read it: Read the synthetic-gold-reward setup before interpreting the scaling curves.