Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack
Semantic and Hybrid Search
Semantic (dense) search embeds the query and returns the stored passages whose vectors are most similar to it, finding matches that share meaning but not words; hybrid search runs keyword and semantic search together and merges their rankings.
The problem
Keyword search can't connect a question to an answer phrased in different words, and embeddings can miss exact names and codes.
The solution
Embed every passage once; at query time embed the question and take the top k by cosine similarity. For robustness, also run BM25 and fuse the two ranked lists, for example with reciprocal rank fusion.
The consequence
Search by meaning became practical and is the default first stage of RAG; hybrid search became common because neither method wins on every query.
You should understand first
- Vectors
- Dot Product
- Embeddings
- Attention
- Probability and Distributions
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Text Embeddings
- Text as Data
- Keyword Search and BM25
- Semantic and Hybrid Search
Search by meaning
Embed every passage once with a text embedding model and store the vectors. When a question arrives:
- Embed the question with the same model.
- Score every stored vector by cosine similarity (a dot product, since both are normalised).
- Return the top .
Karpukhin and colleagues' Dense Passage Retrieval, two BERT encoders trained on question–passage pairs, outperformed a strong Lucene BM25 system by 9–19% absolute in top-20 passage retrieval accuracy on a range of open-domain QA datasets Established. That result made dense retrieval the standard first stage for question answering.
Scoring every vector is exact but slow at millions of passages; real systems use approximate nearest-neighbour indexes.
Each method has blind spots
Dense search finds "how do models reuse earlier work while generating?" in a page about the KV cache even if the page never says "reuse". It can fail on a query that is just an identifier like a library name or an error code: a short string's embedding carries little meaning, and the model may never have seen the term. BM25 is the mirror image. On the corpus in this chapter's lab, each wins clear cases the other misses.
Hybrid search: merge the rankings
BM25 scores and cosine similarities live on different scales, so adding them directly needs careful calibration. Reciprocal rank fusion ignores the scores and uses only ranks:
A passage ranked 1st by one method and 3rd by the other scores 1/61 + 1/63 ≈ 0.0323; one ranked 1st by one method and absent from the other's list scores 1/61 ≈ 0.0164. Agreement wins. Cormack, Clarke and Buettcher found k = 60 near-optimal in pilot experiments, though the choice was not critical, and RRF outperformed more sophisticated rank-fusion and learning methods in their tests Established.
Hybrid search is not always best on any single query: a passage one method ranks first and the other ranks 30th can be overtaken by passages both rank moderately well. What it gives you is robustness on average, and averages over many queries are how retrieval should be judged (evaluation).
Rewriting the query
Questions and documents are written differently. HyDE has a language model write a hypothetical answer document, which may contain false details, embeds it, and retrieves the real documents nearest to it Established. Query rewriting, expansion and splitting a question into sub-questions are common variations on the same idea.
What to remember
- Dense search: embed the query, rank passages by cosine similarity, take the top k.
- It finds paraphrases that keyword search misses, and can miss exact identifiers keyword search finds.
- Reciprocal rank fusion: add 1 / (k + rank) from each list; k = 60 is the usual choice.
- Judge a retriever on many labelled queries, not one impressive example.
Key papers
Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods
Gordon V. Cormack, Charles L. A. Clarke, Stefan Buettcher · 2009 · SIGIR
A two-page paper whose one-line formula is how most hybrid search systems merge keyword and vector results.
Dense Passage Retrieval for Open-Domain Question Answering
Vladimir Karpukhin, Barlas Oğuz et al. · 2020
Showed that learned embeddings alone can beat BM25 for finding answer passages, which made dense retrieval the default first stage of RAG.
How to read it: Section 3 (the dual encoder and in-batch negatives) is the core; Section 4.1 describes splitting Wikipedia into 21 million 100-word passages.
Precise Zero-Shot Dense Retrieval without Relevance Labels
Luyu Gao, Xueguang Ma et al. · 2022
HyDE: let a language model write a hypothetical answer and search with its embedding, since answers look more like documents than questions do.
Watch
Stanford Online
Stanford CS25: V3 I Retrieval Augmented Language Models
Douwe Kiela, one of the authors of the original RAG paper, surveys retrieval-augmented language models and ends with the open questions.