Skip to content
Road to Intelligence

Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack

Text Embeddings

Must knowKnow well16 minDifficulty

A text embedding model turns a whole passage into one fixed-length vector, placed so that passages with similar meaning have a high cosine similarity, which lets a computer compare meaning with a dot product.

The problem

Word vectors give one point per word, but search needs to compare a question with whole passages, and comparing two texts by running a Transformer over every pair is far too slow.

The solution

Encode each text once with a Transformer encoder, pool its token vectors into a single vector and normalise it; train the encoder so that matching texts end up pointing the same way.

The consequence

Millions of passages can be embedded ahead of time and searched with dot products, which is the first stage of almost every RAG system. The price: one vector has to summarise everything in the passage, so details and exact terms get blurred.

From words to passages

Chapter 6 gave every word a vector, learned from the company it keeps (embeddings). Search needs more: a question like "how do I make inference cheaper?" has to be compared with whole paragraphs. A text embedding model takes a passage of any length (up to a limit) and returns a single vector, with the same promise as word vectors: similar meaning, similar direction.

How one vector is made

The usual recipe is a BERT-style encoder:

  1. Tokenize the text and run it through the encoder. Every token comes out as a vector that has attended to the whole passage.
  2. Pool: combine the token vectors into one, most often by averaging them (mean pooling).
  3. Normalise the result to length 1.

Reimers and Gurevych pointed out that BERT, which reads two sentences together, needs about 50 million passes (roughly 65 hours) to find the most similar pair among 10,000 sentences; their Sentence-BERT encodes each sentence once and does the same search in about 5 seconds Established. That difference, compare-by-reading versus compare-by-vector, is the whole reason embeddings power search.

The tiny numeric example

Take three 3-dimensional "embeddings", already normalised (real ones have hundreds of dimensions):

PassageVector
A: "the KV cache stores keys and values"(0.80, 0.60, 0.00)
B: "reuse earlier attention results while generating"(0.60, 0.80, 0.00)
C: "bananas are rich in potassium"(0.00, 0.00, 1.00)

cos(A, B) = 0.80 × 0.60 + 0.60 × 0.80 + 0 = 0.96. cos(A, C) = 0. A and B share almost no words but point the same way; that is what the training is for.

Because each vector has length 1, the cosine is just the dot product:

cos⁡(q,d)=q⋅d∥q∥ ∥d∥=q⋅dwhen ∥q∥=∥d∥=1.\cos(\mathbf{q}, \mathbf{d}) = \frac{\mathbf{q}\cdot\mathbf{d}}{\lVert\mathbf{q}\rVert\,\lVert\mathbf{d}\rVert} = \mathbf{q}\cdot\mathbf{d} \quad \text{when } \lVert\mathbf{q}\rVert = \lVert\mathbf{d}\rVert = 1.

Where the meaning comes from

A pretrained encoder's averaged vectors are not automatically good for search. Embedding models are fine-tuned on pairs that should match (a question and its answer, a title and its article) with contrastive learning. E5, trained contrastively on a large curated set of naturally occurring text pairs, was the first model reported to beat BM25 on the BEIR retrieval benchmark zero-shot, without labelled data Established.

Limits worth knowing

  • A fixed input length. The model reads only so many tokens; the rest is ignored. The small model in this chapter's labs reads 256 word pieces, which is one reason documents are cut into chunks.
  • Blurring. One vector summarises everything, so rare names, codes and numbers can get lost. Keyword search (BM25) catches those.
  • No universal winner. The MTEB benchmark, covering 8 task types and 58 datasets, found that no single embedding method dominated across all tasks Established. Test on your own data.

Why should I care?

As a researcher

Embedding models are where representation learning meets retrieval: what a single vector can and can't capture is an open question with direct consequences for search quality.

As an engineer

Choosing the embedding model, its maximum input length and whether vectors are normalised decides what your search can find, before any language model is involved.

Modern systems that depend on it

  • semantic search
  • retrieval-augmented generation
  • clustering and deduplication
  • recommendations

Historical context

Before

Comparing two sentences with BERT meant feeding both through the network together; finding the most similar pair among 10,000 sentences took hours.

After

Each text is encoded once into a vector; similarity is a dot product, so comparing a query with millions of stored passages takes milliseconds.

Used today

Every vector database and RAG pipeline starts with an embedding model; Chapter 12's labs use all-MiniLM-L6-v2, a small open one with 384 dimensions.

What to remember

  • One passage → one vector, typically a few hundred to a few thousand numbers.
  • Built from a Transformer encoder: token vectors are pooled (often averaged) and normalised.
  • Normalised vectors make cosine similarity and dot product the same number.
  • Encode once, compare many times: that's what makes search over millions of passages possible.
  • A vector blurs details, and models only read a limited number of tokens (256 word pieces for MiniLM).

Key papers

Essential

Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Nils Reimers, Iryna Gurevych · 2019

Made BERT produce one vector per sentence that can be compared with cosine similarity, which is what turned Transformers into practical search engines.

How to read it: Section 3 (the model) explains the siamese network and pooling in a page.

~30 min readarXiv:1908.10084✓ verified 2026-10-05
Important

MTEB: Massive Text Embedding Benchmark

Niklas Muennighoff, Nouamane Tazi et al. · 2022

The benchmark (and public leaderboard) people use to choose an embedding model.

~30 min readarXiv:2210.07316✓ verified 2026-10-05
Important

Text Embeddings by Weakly-Supervised Contrastive Pre-training

Liang Wang, Nan Yang et al. · 2022

E5 showed the modern recipe for general-purpose embedding models: contrastive pretraining on huge numbers of naturally occurring text pairs, then fine-tuning.

~35 min readarXiv:2212.03533✓ verified 2026-10-05

Watch