Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack
Text Embeddings
A text embedding model turns a whole passage into one fixed-length vector, placed so that passages with similar meaning have a high cosine similarity, which lets a computer compare meaning with a dot product.
The problem
Word vectors give one point per word, but search needs to compare a question with whole passages, and comparing two texts by running a Transformer over every pair is far too slow.
The solution
Encode each text once with a Transformer encoder, pool its token vectors into a single vector and normalise it; train the encoder so that matching texts end up pointing the same way.
The consequence
Millions of passages can be embedded ahead of time and searched with dot products, which is the first stage of almost every RAG system. The price: one vector has to summarise everything in the passage, so details and exact terms get blurred.
You should understand first
From words to passages
Chapter 6 gave every word a vector, learned from the company it keeps (embeddings). Search needs more: a question like "how do I make inference cheaper?" has to be compared with whole paragraphs. A text embedding model takes a passage of any length (up to a limit) and returns a single vector, with the same promise as word vectors: similar meaning, similar direction.
How one vector is made
The usual recipe is a BERT-style encoder:
- Tokenize the text and run it through the encoder. Every token comes out as a vector that has attended to the whole passage.
- Pool: combine the token vectors into one, most often by averaging them (mean pooling).
- Normalise the result to length 1.
Reimers and Gurevych pointed out that BERT, which reads two sentences together, needs about 50 million passes (roughly 65 hours) to find the most similar pair among 10,000 sentences; their Sentence-BERT encodes each sentence once and does the same search in about 5 seconds Established. That difference, compare-by-reading versus compare-by-vector, is the whole reason embeddings power search.
The tiny numeric example
Take three 3-dimensional "embeddings", already normalised (real ones have hundreds of dimensions):
| Passage | Vector |
|---|---|
| A: "the KV cache stores keys and values" | (0.80, 0.60, 0.00) |
| B: "reuse earlier attention results while generating" | (0.60, 0.80, 0.00) |
| C: "bananas are rich in potassium" | (0.00, 0.00, 1.00) |
cos(A, B) = 0.80 × 0.60 + 0.60 × 0.80 + 0 = 0.96. cos(A, C) = 0. A and B share almost no words but point the same way; that is what the training is for.
Because each vector has length 1, the cosine is just the dot product:
Where the meaning comes from
A pretrained encoder's averaged vectors are not automatically good for search. Embedding models are fine-tuned on pairs that should match (a question and its answer, a title and its article) with contrastive learning. E5, trained contrastively on a large curated set of naturally occurring text pairs, was the first model reported to beat BM25 on the BEIR retrieval benchmark zero-shot, without labelled data Established.
Limits worth knowing
- A fixed input length. The model reads only so many tokens; the rest is ignored. The small model in this chapter's labs reads 256 word pieces, which is one reason documents are cut into chunks.
- Blurring. One vector summarises everything, so rare names, codes and numbers can get lost. Keyword search (BM25) catches those.
- No universal winner. The MTEB benchmark, covering 8 task types and 58 datasets, found that no single embedding method dominated across all tasks Established. Test on your own data.
Why should I care?
As a researcher
Embedding models are where representation learning meets retrieval: what a single vector can and can't capture is an open question with direct consequences for search quality.
As an engineer
Choosing the embedding model, its maximum input length and whether vectors are normalised decides what your search can find, before any language model is involved.
Modern systems that depend on it
- semantic search
- retrieval-augmented generation
- clustering and deduplication
- recommendations
Historical context
Before
Comparing two sentences with BERT meant feeding both through the network together; finding the most similar pair among 10,000 sentences took hours.
After
Each text is encoded once into a vector; similarity is a dot product, so comparing a query with millions of stored passages takes milliseconds.
Used today
Every vector database and RAG pipeline starts with an embedding model; Chapter 12's labs use all-MiniLM-L6-v2, a small open one with 384 dimensions.
What to remember
- One passage → one vector, typically a few hundred to a few thousand numbers.
- Built from a Transformer encoder: token vectors are pooled (often averaged) and normalised.
- Normalised vectors make cosine similarity and dot product the same number.
- Encode once, compare many times: that's what makes search over millions of passages possible.
- A vector blurs details, and models only read a limited number of tokens (256 word pieces for MiniLM).
Key papers
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Nils Reimers, Iryna Gurevych · 2019
Made BERT produce one vector per sentence that can be compared with cosine similarity, which is what turned Transformers into practical search engines.
How to read it: Section 3 (the model) explains the siamese network and pooling in a page.
MTEB: Massive Text Embedding Benchmark
Niklas Muennighoff, Nouamane Tazi et al. · 2022
The benchmark (and public leaderboard) people use to choose an embedding model.
Text Embeddings by Weakly-Supervised Contrastive Pre-training
Liang Wang, Nan Yang et al. · 2022
E5 showed the modern recipe for general-purpose embedding models: contrastive pretraining on huge numbers of naturally occurring text pairs, then fine-tuning.
Watch
3Blue1Brown
Dot products and duality | Chapter 9, Essence of linear algebra
Explains why the dot product measures alignment, connecting the arithmetic to the geometry.