Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack
Chunking
Chunking splits documents into passages before embedding them, and the chunk size sets a trade-off: small chunks are precise but can cut an answer in half, large chunks keep context but blur their vectors and cost more prompt space.
The problem
A whole document makes a poor search unit: one vector can't represent many topics, embedding models only read a limited number of tokens, and the prompt can hold only a few documents.
The solution
Cut documents into passages of a chosen size, optionally overlapping and following the document's structure (headings, paragraphs), and attach a short context such as the document title to each.
The consequence
Retrieval quality depends on a preprocessing choice that has nothing to do with the model. There is no universally right size; it depends on the documents, the embedding model's input limit and the questions.
You should understand first
- Vectors
- Dot Product
- Embeddings
- Attention
- Probability and Distributions
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Text Embeddings
- Text as Data
- Keyword Search and BM25
- Semantic and Hybrid Search
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- Parametric vs Retrieved Knowledge
- Retrieval-Augmented Generation (RAG)
- One-Hot Encoding
- Tokenization
- Chunking
Why cut at all
Three reasons. A 20-page document covers many topics, and one embedding averages them into mush. Embedding models read only so many tokens. And you'll put a handful of passages into the prompt, not whole manuals. Dense Passage Retrieval and the original RAG system split Wikipedia into disjoint 100-word passages, about 21 million of them Established.
The trade-off
Picture the answer to a question as a 25-word span somewhere in a page.
- Small chunks (say 40 words) match precisely, but the span often crosses a boundary, so the best-matching chunk holds half the answer. They also lose context: a chunk that says "it uses 4 bits" doesn't say what "it" is.
- Large chunks (say 200 words) usually contain the whole span, but their vectors average several ideas, so they match less sharply, and each one retrieved costs more of the prompt. Very large chunks can exceed the embedding model's input limit and get silently truncated.
The lab below lets you watch both failures on this site's own pages, with real embeddings.
Common fixes
- Overlap: start each chunk a little before the previous one ends, so a span near a boundary appears whole in one of them. It adds index size and duplicates in results. The DPR authors report that splitting into overlapping passages was not advantageous compared with non-overlapping ones in their setting Established, so test it rather than assume.
- Structure-aware splitting: cut at headings, paragraphs or sentences instead of every N words; for code, at functions.
- Context headers: prefix each chunk with its document title (this chapter's index does) or a generated one-line description of where it sits. Anthropic reported in September 2024 that prepending chunk-specific context before embedding and BM25 indexing, which it calls contextual retrieval, reduced its top-20 retrieval failure rate by 49%, and by 67% with reranking added Established.
- Small to retrieve, large to read: match on small chunks but put the surrounding larger section into the prompt.
What to remember
- Too small: answers split across chunks, and chunks lose their context.
- Too large: diluted vectors, truncated embeddings, wasted prompt tokens.
- Overlap repeats text across boundaries; structure-aware splits avoid cutting sentences.
- Prefix each chunk with its document's title or a short summary of where it comes from.
- Check that chunks fit the embedding model's maximum input length.
Key papers
Dense Passage Retrieval for Open-Domain Question Answering
Vladimir Karpukhin, Barlas Oğuz et al. · 2020
Showed that learned embeddings alone can beat BM25 for finding answer passages, which made dense retrieval the default first stage of RAG.
How to read it: Section 3 (the dual encoder and in-batch negatives) is the core; Section 4.1 describes splitting Wikipedia into 21 million 100-word passages.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Patrick Lewis, Ethan Perez et al. · 2020
Named retrieval-augmented generation: a generator conditioned on passages fetched from a dense index, with knowledge that can be inspected and updated without retraining.
How to read it: Section 2 (methods: the DPR retriever, the BART generator and the two ways of combining passages) is the core; the index is 21 million 100-word chunks of a December 2018 Wikipedia dump.