Concept · Chapter 9: How an LLM Is Actually Built
Building a Pretraining Dataset
A pretraining dataset is built by a pipeline that extracts text from web crawls and other sources, then removes most of it, through language identification, quality filters, safety filters and deduplication, before trillions of tokens remain.
The problem
The raw web is mostly boilerplate, spam, duplicates and machine-generated text in hundreds of languages, and a model learns whatever it is fed.
The solution
Run the crawl through an ETL pipeline: extract clean text from the HTML, identify the language, filter by heuristics and by quality classifiers, remove personal and unsafe content, and deduplicate at the URL, document and line level.
The consequence
Data work became one of the biggest levers on model quality, and one of the least disclosed. Open datasets such as FineWeb now document every step, with an ablation showing what each one is worth.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- One-Hot Encoding
- Tokenization
- Building a Pretraining Dataset
Intuition: an ETL job where the output is the model's world
A data engineer would recognise the shape immediately. Extract from a huge, messy source; transform through a chain of filters; load into a training set. The difference is the stakes: the model knows nothing except what survives the pipeline.
The usual source is Common Crawl, a public archive of the web released as a series of snapshots. A snapshot holds whatever was crawled: navigation menus, cookie banners, product listings, spam, machine translations and near-copies of other pages, alongside the text worth learning from.
The pipeline, step by step
The open FineWeb dataset and Meta's Llama 3 report describe similar steps. FineWeb built 15 trillion tokens from 96 Common Crawl snapshots Established.
Extract
Pull the main text out of the raw HTML. FineWeb found that extracting text from the raw archive files with a dedicated library gave better models than Common Crawl's own pre-extracted text, which kept too much boilerplate Established. Llama 3's team removed markdown markers after finding they hurt models trained mostly on web data Established.Identify the language
A fast classifier labels each document. FineWeb kept pages scored as English with probability at least 0.65; Llama 3 sorted documents into 176 languages Established.Filter with rules
Cheap heuristics remove obvious junk: too short, too repetitive, too many symbols, lines that look like logs or error messages, lists of "dirty words".Filter with models
A classifier scores quality. CCNet kept documents that a small language model trained on Wikipedia found plausible (low perplexity) Established. For FineWeb-Edu, Llama-3-70B-Instruct rated 460,000 pages for educational value; a small classifier trained on those ratings then scored all of FineWeb, keeping 1.3 trillion tokens Established.Remove unsafe and personal data
Block domains known for adult content, unsafe material or large amounts of personal information.Deduplicate
Remove exact and near-duplicate pages and repeated lines. This gets its own page: deduplication.
A real funnel. FineWeb's basic filters (a URL blocklist, English identification and repetition rules) left about 36 trillion tokens from 96 snapshots. Deduplicating each snapshot left about 20 trillion, and further filters brought the final dataset to 15 trillion. The educational-quality filter then kept 1.3 trillion of those for FineWeb-Edu Established. Each stage is a trade: fewer tokens, better ones.
Every filter is a choice about the world
Filters encode someone's idea of "quality". An audit of the C4 dataset found that its blocklist filter disproportionately removed text from and about minority individuals Established. A classifier trained to prefer Wikipedia-like text will under-represent informal writing, dialects and many languages. These are trade-offs, not bugs to be fixed once.
Why should I care?
As a researcher
Data choices are confounders in almost every comparison between models, and data curation is an active research area with large, cheap wins still being found.
As an engineer
It is a data-engineering problem at petabyte scale: extraction, filtering, fuzzy deduplication and lineage. The same techniques apply to building a retrieval corpus or a fine-tuning set.
Modern systems that depend on it
- data mixtures
- deduplication
- benchmark contamination
- what a model knows
Historical context
Before
Training sets were curated corpora (Wikipedia, books, news) of a few billion words, or lightly filtered web scrapes such as C4.
After
Pipelines over dozens of Common Crawl snapshots with model-based quality scoring, producing tens of trillions of candidate tokens, documented and ablated step by step.
Used today
Every pretrained model. Open examples: FineWeb (15T tokens) and FineWeb-Edu (1.3T); Llama 3's report describes Meta's pipeline in detail.
What to remember
- Source: mostly web crawls (Common Crawl), plus code, books, papers and other curated sources.
- Pipeline: extract text → language ID → heuristic filters → quality classifiers → safety/PII filters → deduplication.
- Most of the raw crawl is thrown away; what remains still needs mixing and weighting.
- Quality classifiers are often trained on labels from a stronger model (FineWeb-Edu used Llama 3's ratings).
- Each step is a judgement call with side effects: filters can remove whole dialects, topics or communities.
Key papers
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer et al. · 2019 · JMLR 2020
Framed every NLP task as text in, text out, using an encoder–decoder Transformer — and ran a huge, careful set of ablations that is still a model of empirical method.
How to read it: Long (67 pages). Read the introduction and Section 3.2's architecture comparison; treat the rest as a reference.
CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data
Guillaume Wenzek, Marie-Anne Lachaux et al. · 2019
A widely copied web-cleaning pipeline: deduplicate, identify the language, then keep documents that a language model trained on Wikipedia finds plausible.
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
Jesse Dodge, Maarten Sap et al. · 2021
One of the first audits of what a web-scale training corpus actually contains, including benchmark test data and the side effects of its filters.
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Guilherme Penedo, Hynek Kydlíček et al. · 2024 · NeurIPS 2024
The most thoroughly documented open recipe for turning Common Crawl into pretraining data, with an ablation for every step.
How to read it: Section 3 walks through the pipeline step by step; the global-deduplication surprise is in 3.4.
The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey et al. · 2024
The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.
How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.
Watch
Andrej Karpathy
Deep Dive into LLMs like ChatGPT
A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 13: Data 1
A tour of what real pretraining datasets contain and how they were built, from a course that argues data matters most.