Skip to content
Road to Intelligence

Concept · Chapter 9: How an LLM Is Actually Built

Building a Pretraining Dataset

Must knowKnow well16 minDifficulty

A pretraining dataset is built by a pipeline that extracts text from web crawls and other sources, then removes most of it, through language identification, quality filters, safety filters and deduplication, before trillions of tokens remain.

The problem

The raw web is mostly boilerplate, spam, duplicates and machine-generated text in hundreds of languages, and a model learns whatever it is fed.

The solution

Run the crawl through an ETL pipeline: extract clean text from the HTML, identify the language, filter by heuristics and by quality classifiers, remove personal and unsafe content, and deduplicate at the URL, document and line level.

The consequence

Data work became one of the biggest levers on model quality, and one of the least disclosed. Open datasets such as FineWeb now document every step, with an ablation showing what each one is worth.

Intuition: an ETL job where the output is the model's world

A data engineer would recognise the shape immediately. Extract from a huge, messy source; transform through a chain of filters; load into a training set. The difference is the stakes: the model knows nothing except what survives the pipeline.

The usual source is Common Crawl, a public archive of the web released as a series of snapshots. A snapshot holds whatever was crawled: navigation menus, cookie banners, product listings, spam, machine translations and near-copies of other pages, alongside the text worth learning from.

The pipeline, step by step

The open FineWeb dataset and Meta's Llama 3 report describe similar steps. FineWeb built 15 trillion tokens from 96 Common Crawl snapshots Established.

  1. Extract

    Pull the main text out of the raw HTML. FineWeb found that extracting text from the raw archive files with a dedicated library gave better models than Common Crawl's own pre-extracted text, which kept too much boilerplate Established. Llama 3's team removed markdown markers after finding they hurt models trained mostly on web data Established.
  2. Identify the language

    A fast classifier labels each document. FineWeb kept pages scored as English with probability at least 0.65; Llama 3 sorted documents into 176 languages Established.
  3. Filter with rules

    Cheap heuristics remove obvious junk: too short, too repetitive, too many symbols, lines that look like logs or error messages, lists of "dirty words".
  4. Filter with models

    A classifier scores quality. CCNet kept documents that a small language model trained on Wikipedia found plausible (low perplexity) Established. For FineWeb-Edu, Llama-3-70B-Instruct rated 460,000 pages for educational value; a small classifier trained on those ratings then scored all of FineWeb, keeping 1.3 trillion tokens Established.
  5. Remove unsafe and personal data

    Block domains known for adult content, unsafe material or large amounts of personal information.
  6. Deduplicate

    Remove exact and near-duplicate pages and repeated lines. This gets its own page: deduplication.

A real funnel. FineWeb's basic filters (a URL blocklist, English identification and repetition rules) left about 36 trillion tokens from 96 snapshots. Deduplicating each snapshot left about 20 trillion, and further filters brought the final dataset to 15 trillion. The educational-quality filter then kept 1.3 trillion of those for FineWeb-Edu Established. Each stage is a trade: fewer tokens, better ones.

Every filter is a choice about the world

Filters encode someone's idea of "quality". An audit of the C4 dataset found that its blocklist filter disproportionately removed text from and about minority individuals Established. A classifier trained to prefer Wikipedia-like text will under-represent informal writing, dialects and many languages. These are trade-offs, not bugs to be fixed once.

Why should I care?

As a researcher

Data choices are confounders in almost every comparison between models, and data curation is an active research area with large, cheap wins still being found.

As an engineer

It is a data-engineering problem at petabyte scale: extraction, filtering, fuzzy deduplication and lineage. The same techniques apply to building a retrieval corpus or a fine-tuning set.

Modern systems that depend on it

  • data mixtures
  • deduplication
  • benchmark contamination
  • what a model knows

Historical context

Before

Training sets were curated corpora (Wikipedia, books, news) of a few billion words, or lightly filtered web scrapes such as C4.

After

Pipelines over dozens of Common Crawl snapshots with model-based quality scoring, producing tens of trillions of candidate tokens, documented and ablated step by step.

Used today

Every pretrained model. Open examples: FineWeb (15T tokens) and FineWeb-Edu (1.3T); Llama 3's report describes Meta's pipeline in detail.

What to remember

  • Source: mostly web crawls (Common Crawl), plus code, books, papers and other curated sources.
  • Pipeline: extract text → language ID → heuristic filters → quality classifiers → safety/PII filters → deduplication.
  • Most of the raw crawl is thrown away; what remains still needs mixing and weighting.
  • Quality classifiers are often trained on labels from a stronger model (FineWeb-Edu used Llama 3's ratings).
  • Each step is a judgement call with side effects: filters can remove whole dialects, topics or communities.

Key papers

Important

Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Colin Raffel, Noam Shazeer et al. · 2019 · JMLR 2020

Framed every NLP task as text in, text out, using an encoder–decoder Transformer — and ran a huge, careful set of ablations that is still a model of empirical method.

How to read it: Long (67 pages). Read the introduction and Section 3.2's architecture comparison; treat the rest as a reference.

~2 h readarXiv:1910.10683✓ verified 2026-09-26
Important

The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Guilherme Penedo, Hynek Kydlíček et al. · 2024 · NeurIPS 2024

The most thoroughly documented open recipe for turning Common Crawl into pretraining data, with an ablation for every step.

How to read it: Section 3 walks through the pipeline step by step; the global-deduplication surprise is in 3.4.

~45 min readarXiv:2406.17557✓ verified 2026-10-04
Essential

The Llama 3 Herd of Models

Aaron Grattafiori, Abhimanyu Dubey et al. · 2024

The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.

How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.

~2 h readarXiv:2407.21783✓ verified 2026-10-04

Watch

3 h 31 min

Andrej Karpathy

Deep Dive into LLMs like ChatGPT

A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.

Should know