Skip to content
Road to Intelligence

Concept · Chapter 9: How an LLM Is Actually Built

Designing the Tokenizer

Should knowKnow well11 minDifficulty

Building a tokenizer means choosing an algorithm (BPE, WordPiece or unigram), training data, a vocabulary size, rules for numbers and unknown characters, and a set of special tokens, and every choice is frozen into the model for good.

The problem

The tokenizer is fixed before pretraining starts; a poor choice wastes compute on long token sequences or serves some languages and digits badly for the model's whole life.

The solution

Train the vocabulary on a sample of the actual training mix, size it to balance compression against embedding cost, split digits consistently, fall back to bytes for anything unknown, and reserve special tokens for structure.

The consequence

Better compression means more text per unit of compute and context; a broader vocabulary serves more languages; special tokens later carry chat formats and tool calls.

You should understand first

  1. Text as Data
  2. Vectors
  3. One-Hot Encoding
  4. Tokenization
  5. Designing the Tokenizer

The algorithm: three close cousins

Chapter 8 trained a BPE tokenizer by repeatedly merging the most frequent pair. Two relatives are common:

  • WordPiece grows the vocabulary by adding the piece that most increases the likelihood of the training data. It was developed for Japanese and Korean voice search Established, and BERT used WordPiece with a 30,000-token vocabulary Established.
  • Unigram starts from a large vocabulary and prunes the pieces that matter least.

The SentencePiece library trains either BPE or unigram vocabularies directly from raw sentences, treating the space as an ordinary symbol, so it works for languages written without spaces and the original text can always be recovered exactly Established.

The choices that matter

Vocabulary size. A bigger vocabulary packs more text into each token, so the same compute and context window cover more text. The cost is a larger embedding table and a larger output softmax. Llama 3 moved to a 128,000-token vocabulary (100,000 from OpenAI's tiktoken tokenizer plus 28,000 for non-English languages), which improved compression on English from 3.17 to 3.94 characters per token compared with Llama 2 Established.

Tiny example. At 3.17 characters per token, a 10,000-character document is about 3,150 tokens. At 3.94 it is about 2,540: 19% fewer tokens to process for the same text, in training and at inference.

Numbers and the unknown. LLaMA's tokenizer split every number into single digits and fell back to raw bytes for characters it did not know Established, so arithmetic sees consistent pieces and no input is ever out of vocabulary.

Special tokens. Reserved IDs mark structure rather than text. GPT-2's vocabulary includes <|endoftext|>, placed between documents in training so the model learns where one ends. Chat models add tokens for roles and turns, which Chapter 10 builds on.

Fixed for life

The tokenizer is chosen before pretraining and baked into the embedding table. Changing it later means retraining, or at least expensive surgery. That is why tokenizer quirks, such as poor compression for some scripts, persist across a model's lifetime.

What to remember

  • The tokenizer is trained before the model and can't be changed without retraining.
  • Bigger vocabulary: fewer tokens per text, but a bigger embedding table and output layer.
  • Llama 3: 128K tokens; English compression rose from 3.17 to 3.94 characters per token versus Llama 2.
  • LLaMA split numbers into single digits and fell back to bytes for unknown characters.
  • Special tokens mark structure: GPT-2's <|endoftext|> separates documents.

Key papers

Important

LLaMA: Open and Efficient Foundation Language Models

Hugo Touvron, Thibaut Lavril et al. · 2023

Showed that smaller models trained on more tokens, using only publicly available data, can rival much larger ones, and released weights to researchers, starting the open-weight wave.

~30 min readarXiv:2302.13971✓ verified 2026-09-26
Optional

Japanese and Korean voice search

Mike Schuster, Kaisuke Nakajima · 2012 · ICASSP 2012

The origin of the WordPiece subword method, later used for BERT's 30,000-token vocabulary.

How to read it: Only the section on building the word inventory matters here; the rest is about speech recognition.

~20 min readdoi:10.1109/ICASSP.2012.6289079✓ verified 2026-10-04
Essential

The Llama 3 Herd of Models

Aaron Grattafiori, Abhimanyu Dubey et al. · 2024

The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.

How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.

~2 h readarXiv:2407.21783✓ verified 2026-10-04

Watch