Concept · Chapter 9: How an LLM Is Actually Built
Data Mixtures, Repetition and Synthetic Data
A training set is a weighted mixture of sources (web, code, books, papers, maths), and the weights, the number of times each source is repeated, and any model-generated data are deliberate design choices.
The problem
Sources differ hugely in size and value: the web is vast but noisy, while books, code and papers are smaller and denser. Sampling in proportion to size would drown the best data.
The solution
Assign each source a sampling weight, upsample small high-value sources for a few passes, choose the weights by experiments on small models, and add generated data where real data is scarce.
The consequence
The mix shapes what the model is good at (more code and maths improve reasoning benchmarks), and high-quality data is now the scarce input, not compute.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- One-Hot Encoding
- Tokenization
- Building a Pretraining Dataset
- Data Mixtures, Repetition and Synthetic Data
Weights, not sizes
Suppose the web gives you 10 trillion tokens and Wikipedia gives you 5 billion. Sample in proportion to size and the model would barely see Wikipedia. So training sets are mixtures with chosen weights. LLaMA (2023) sampled 67% of its training tokens from Common Crawl, 15% from C4, 4.5% each from GitHub, Wikipedia and books, 2.5% from arXiv and 2% from Stack Exchange; at 1.4 trillion tokens that meant about 2.45 passes over Wikipedia and 2.23 over the books, but 1.1 over Common Crawl Established.
| Source | Share of tokens | Passes (epochs) |
|---|---|---|
| Common Crawl | 67.0% | 1.10 |
| C4 | 15.0% | 1.06 |
| GitHub | 4.5% | 0.64 |
| Wikipedia | 4.5% | 2.45 |
| Books | 4.5% | 2.23 |
| ArXiv | 2.5% | 1.06 |
| Stack Exchange | 2.0% | 1.03 |
Choosing the weights
Mixtures used to be set by hand. Now they are tuned like any other hyperparameter. Llama 3's team trained small models on candidate mixes, used scaling-law fits to predict how a large model would do on each, and settled on roughly 50% general knowledge, 25% maths and reasoning, 17% code and 8% multilingual tokens Established. They also used a classifier to down-sample categories over-represented on the web, such as arts and entertainment, and changed the mix during training, for example adding more recent web data later to move the knowledge cutoff forward Established.
At the very end of training, Llama 3 was "annealed" on a small upsampled set of high-quality code and maths, which improved the 8B model on maths benchmarks but made little difference to the 405B model Established.
When good data runs out: repeat it?
Scaling laws assume fresh tokens. Real high-quality text is finite. Muennighoff and colleagues found that, for a fixed compute budget, training on up to about 4 epochs of repeated data gave almost the same loss as unique data; beyond that, extra repetitions were worth less and less, approaching nothing after many more passes Established.
Synthetic data
A strong model can write training data for another. phi-1, a 1.3B-parameter coding model, was trained on 6B tokens of web code filtered for "textbook quality" plus 1B tokens of textbooks and exercises generated by GPT-3.5, and despite its small size reached 50.6% pass@1 on the HumanEval coding benchmark Established. How far synthetic data can substitute for human-written text in general pretraining, and whether training on model outputs slowly degrades diversity, are open questions Active research.
What to remember
- Sampling weights, not raw sizes, decide how much the model sees of each source.
- LLaMA (2023): 67% Common Crawl, 15% C4, 4.5% each GitHub, Wikipedia and books; Wikipedia and books seen more than twice.
- Llama 3: about 50% general knowledge, 25% maths and reasoning, 17% code, 8% multilingual.
- Up to about 4 epochs, repeated data was almost as good as fresh data (Muennighoff et al. 2023).
- Synthetic data can help a lot on narrow tasks; how far it can replace real data is debated.
Key papers
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril et al. · 2023
Showed that smaller models trained on more tokens, using only publicly available data, can rival much larger ones, and released weights to researchers, starting the open-weight wave.
The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Leo Gao, Stella Biderman et al. · 2020
An influential open pretraining dataset that mixed 22 sources, from academic papers to code, and documented them.
Scaling Data-Constrained Language Models
Niklas Muennighoff, Alexander M. Rush et al. · 2023
Asked what happens when high-quality text runs out: how much is repeated data worth?
Textbooks Are All You Need
Suriya Gunasekar, Yi Zhang et al. · 2023
A striking demonstration that carefully selected and synthetic data can let a small model compete with much larger ones on a narrow task.
How to read it: Read it as evidence about data quality, not as a general recipe: the evaluation is narrow (Python functions).
The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey et al. · 2024
The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.
How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 13: Data 1
A tour of what real pretraining datasets contain and how they were built, from a course that argues data matters most.