Skip to content
Road to Intelligence

Concept · Chapter 10: From Base Model to Assistant

Generate, Rank, Then Train

Should knowUnderstand8 minDifficulty

Generate several responses, select candidates with a judge, and use the selected responses as supervised demonstrations.

The problem

A model may produce a good response only sometimes, and manually writing every demonstration is expensive.

The solution

Sample multiple candidates, rank or filter them, and train on the retained examples.

The consequence

Selection can improve the demonstration pool, but it inherits judge errors and costs additional generation.

Keep the better attempts

Ask a model for several summaries of one paragraph. A judge ranks them; a filter rejects malformed or unsupported answers; retained responses enter an SFT dataset. The weight update then imitates the selected text.

This chapter uses “rejection sampling” in the post-training data-selection sense. It is not a claim that best-of-N selection is the classical exact rejection-sampling algorithm for drawing from an arbitrary target distribution.

Selection is a separate stage

Best-of-N can also select a response at inference without training. If the selected response becomes a later training example, that is an additional step. A system can combine this selection with DPO on a separate set of preference pairs.

The Llama 3 report describes precisely such an integrated pipeline: reward scoring helps select demonstrations, while DPO performs preference optimization. An objective and an entire model-development process should not be treated as synonyms.

More candidates, more opportunities

More attempts can uncover a useful response, but also a response that exploits the judge. Independently check selected outputs, and compare the gain with the cost of generation and scoring. Chapter 11's efficient inference matters here even before the final assistant is deployed.

What to remember

  • Candidate selection and policy optimization are different operations.
  • A judge's favorite is not necessarily the most accurate response.
  • Generating more candidates increases inference cost during training-data creation.

Key papers

Essential

The Llama 3 Herd of Models

Aaron Grattafiori, Abhimanyu Dubey et al. · 2024

The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.

How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.

~2 h readarXiv:2407.21783✓ verified 2026-10-04
Important

Scaling Laws for Reward Model Overoptimization

Leo Gao, John Schulman, Jacob Hilton · 2022

Explains why maximizing a learned judge can diverge from the intended outcome.

How to read it: Read the synthetic-gold-reward setup before interpreting the scaling curves.

~30 min readarXiv:2210.10760✓ verified 2026-10-04