Concept · Chapter 9: How an LLM Is Actually Built
Designing the Tokenizer
Building a tokenizer means choosing an algorithm (BPE, WordPiece or unigram), training data, a vocabulary size, rules for numbers and unknown characters, and a set of special tokens, and every choice is frozen into the model for good.
The problem
The tokenizer is fixed before pretraining starts; a poor choice wastes compute on long token sequences or serves some languages and digits badly for the model's whole life.
The solution
Train the vocabulary on a sample of the actual training mix, size it to balance compression against embedding cost, split digits consistently, fall back to bytes for anything unknown, and reserve special tokens for structure.
The consequence
Better compression means more text per unit of compute and context; a broader vocabulary serves more languages; special tokens later carry chat formats and tool calls.
You should understand first
- Text as Data
- Vectors
- One-Hot Encoding
- Tokenization
- Designing the Tokenizer
The algorithm: three close cousins
Chapter 8 trained a BPE tokenizer by repeatedly merging the most frequent pair. Two relatives are common:
- WordPiece grows the vocabulary by adding the piece that most increases the likelihood of the training data. It was developed for Japanese and Korean voice search Established, and BERT used WordPiece with a 30,000-token vocabulary Established.
- Unigram starts from a large vocabulary and prunes the pieces that matter least.
The SentencePiece library trains either BPE or unigram vocabularies directly from raw sentences, treating the space as an ordinary symbol, so it works for languages written without spaces and the original text can always be recovered exactly Established.
The choices that matter
Vocabulary size. A bigger vocabulary packs more text into each token, so the same compute and context window cover more text. The cost is a larger embedding table and a larger output softmax. Llama 3 moved to a 128,000-token vocabulary (100,000 from OpenAI's tiktoken tokenizer plus 28,000 for non-English languages), which improved compression on English from 3.17 to 3.94 characters per token compared with Llama 2 Established.
Tiny example. At 3.17 characters per token, a 10,000-character document is about 3,150 tokens. At 3.94 it is about 2,540: 19% fewer tokens to process for the same text, in training and at inference.
Numbers and the unknown. LLaMA's tokenizer split every number into single digits and fell back to raw bytes for characters it did not know Established, so arithmetic sees consistent pieces and no input is ever out of vocabulary.
Special tokens. Reserved IDs mark structure rather than text. GPT-2's vocabulary includes <|endoftext|>, placed between documents in training so the model learns where one ends. Chat models add tokens for roles and turns, which Chapter 10 builds on.
Fixed for life
The tokenizer is chosen before pretraining and baked into the embedding table. Changing it later means retraining, or at least expensive surgery. That is why tokenizer quirks, such as poor compression for some scripts, persist across a model's lifetime.
What to remember
- The tokenizer is trained before the model and can't be changed without retraining.
- Bigger vocabulary: fewer tokens per text, but a bigger embedding table and output layer.
- Llama 3: 128K tokens; English compression rose from 3.17 to 3.94 characters per token versus Llama 2.
- LLaMA split numbers into single digits and fell back to bytes for unknown characters.
- Special tokens mark structure: GPT-2's <|endoftext|> separates documents.
Key papers
Neural Machine Translation of Rare Words with Subword Units
Rico Sennrich, Barry Haddow, Alexandra Birch · 2015 · ACL 2016
Brought byte-pair encoding (BPE) to neural NLP — the ancestor of the tokenizers in GPT-style models.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril et al. · 2023
Showed that smaller models trained on more tokens, using only publicly available data, can rival much larger ones, and released weights to researchers, starting the open-weight wave.
Japanese and Korean voice search
Mike Schuster, Kaisuke Nakajima · 2012 · ICASSP 2012
The origin of the WordPiece subword method, later used for BERT's 30,000-token vocabulary.
How to read it: Only the section on building the word inventory matters here; the rest is about speech recognition.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing
Taku Kudo, John Richardson · 2018
The tokenizer library behind many LLMs, including LLaMA: it trains BPE or unigram vocabularies directly from raw text in any language.
The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey et al. · 2024
The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.
How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.
Watch
Andrej Karpathy
Let's build the GPT Tokenizer
Many odd LLM behaviours trace back to tokenization; this shows you why by building a BPE tokenizer.