Skip to content
Road to Intelligence

Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack

Semantic and Hybrid Search

Must knowKnow well14 minDifficulty

Semantic (dense) search embeds the query and returns the stored passages whose vectors are most similar to it, finding matches that share meaning but not words; hybrid search runs keyword and semantic search together and merges their rankings.

The problem

Keyword search can't connect a question to an answer phrased in different words, and embeddings can miss exact names and codes.

The solution

Embed every passage once; at query time embed the question and take the top k by cosine similarity. For robustness, also run BM25 and fuse the two ranked lists, for example with reciprocal rank fusion.

The consequence

Search by meaning became practical and is the default first stage of RAG; hybrid search became common because neither method wins on every query.

Search by meaning

Embed every passage once with a text embedding model and store the vectors. When a question arrives:

  1. Embed the question with the same model.
  2. Score every stored vector by cosine similarity (a dot product, since both are normalised).
  3. Return the top kk.

Karpukhin and colleagues' Dense Passage Retrieval, two BERT encoders trained on question–passage pairs, outperformed a strong Lucene BM25 system by 9–19% absolute in top-20 passage retrieval accuracy on a range of open-domain QA datasets Established. That result made dense retrieval the standard first stage for question answering.

Scoring every vector is exact but slow at millions of passages; real systems use approximate nearest-neighbour indexes.

Each method has blind spots

Dense search finds "how do models reuse earlier work while generating?" in a page about the KV cache even if the page never says "reuse". It can fail on a query that is just an identifier like a library name or an error code: a short string's embedding carries little meaning, and the model may never have seen the term. BM25 is the mirror image. On the corpus in this chapter's lab, each wins clear cases the other misses.

Hybrid search: merge the rankings

BM25 scores and cosine similarities live on different scales, so adding them directly needs careful calibration. Reciprocal rank fusion ignores the scores and uses only ranks:

RRF(d)=∑rankings r1k+r(d).\text{RRF}(d) = \sum_{\text{rankings } r} \frac{1}{k + r(d)}.

A passage ranked 1st by one method and 3rd by the other scores 1/61 + 1/63 ≈ 0.0323; one ranked 1st by one method and absent from the other's list scores 1/61 ≈ 0.0164. Agreement wins. Cormack, Clarke and Buettcher found k = 60 near-optimal in pilot experiments, though the choice was not critical, and RRF outperformed more sophisticated rank-fusion and learning methods in their tests Established.

Hybrid search is not always best on any single query: a passage one method ranks first and the other ranks 30th can be overtaken by passages both rank moderately well. What it gives you is robustness on average, and averages over many queries are how retrieval should be judged (evaluation).

Rewriting the query

Questions and documents are written differently. HyDE has a language model write a hypothetical answer document, which may contain false details, embeds it, and retrieves the real documents nearest to it Established. Query rewriting, expansion and splitting a question into sub-questions are common variations on the same idea.

What to remember

  • Dense search: embed the query, rank passages by cosine similarity, take the top k.
  • It finds paraphrases that keyword search misses, and can miss exact identifiers keyword search finds.
  • Reciprocal rank fusion: add 1 / (k + rank) from each list; k = 60 is the usual choice.
  • Judge a retriever on many labelled queries, not one impressive example.

Key papers

Essential

Dense Passage Retrieval for Open-Domain Question Answering

Vladimir Karpukhin, Barlas Oğuz et al. · 2020

Showed that learned embeddings alone can beat BM25 for finding answer passages, which made dense retrieval the default first stage of RAG.

How to read it: Section 3 (the dual encoder and in-batch negatives) is the core; Section 4.1 describes splitting Wikipedia into 21 million 100-word passages.

~40 min readarXiv:2004.04906✓ verified 2026-10-05
Optional

Precise Zero-Shot Dense Retrieval without Relevance Labels

Luyu Gao, Xueguang Ma et al. · 2022

HyDE: let a language model write a hypothetical answer and search with its embedding, since answers look more like documents than questions do.

~25 min readarXiv:2212.10496✓ verified 2026-10-05

Watch