Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack
Reranking
Reranking takes the top candidates from a fast retriever and rescores them with a slower, more accurate model, typically a cross-encoder that reads the query and each passage together.
The problem
Embedding search compresses each passage into one vector before it ever sees the query, so its ordering of near-misses is coarse; reading every passage with the query is accurate but far too slow for a whole collection.
The solution
Two stages: retrieve the top 50–100 cheaply (BM25, dense or hybrid), then run a cross-encoder over each query–passage pair and reorder by its score.
The consequence
Large quality gains for modest cost, because the expensive model only sees a handful of candidates. The first stage still has to find the right passage: a reranker can reorder, never recover what wasn't retrieved.
You should understand first
- Vectors
- Dot Product
- Embeddings
- Attention
- Probability and Distributions
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Text Embeddings
- Text as Data
- Keyword Search and BM25
- Semantic and Hybrid Search
- Reranking
Reading together versus apart
A text embedding model is a bi-encoder: the passage is summarised into a vector without knowing what will be asked. A cross-encoder puts the query and the passage into one input, [CLS] query [SEP] passage, so every query token can attend to every passage token, and outputs a single relevance score. It can notice that the passage mentions the right model but the wrong year, which a single vector tends to blur.
The cost is that nothing can be precomputed: scoring 10 million passages means 10 million forward passes per query. So it's used as a second stage.
The funnel
- Retrieve 50–100 candidates with BM25, dense search or both. Milliseconds.
- Rerank those candidates with the cross-encoder. One forward pass each, batched.
- Keep the top few for the prompt.
Nogueira and Cho re-ranked the top 1,000 BM25 passages with BERT and reached the top of the MS MARCO passage leaderboard, beating the previous state of the art by 27% (relative) in MRR@10 Established. On BEIR's 18 datasets, a BM25-plus-cross-encoder reranker performed best overall and beat BM25 on 16 of them, but at high computational cost Established.
In between: late interaction
ColBERT encodes the query and the document separately but keeps a vector per token, and scores them with a cheap interaction step; document vectors can be computed offline, and it was competitive with BERT re-rankers while two orders of magnitude faster Established. The price is storage: one vector per token instead of one per passage.
What to remember
- Bi-encoder: encode separately, compare vectors. Fast, coarse.
- Cross-encoder: one pass over query + passage together. Accurate, one pass per candidate.
- Late interaction (ColBERT): one vector per token, cheap max-similarity matching; in between.
- Retrieve many cheaply, rerank a few expensively.
- A reranker can't fix a passage the first stage never returned.
Key papers
Passage Re-ranking with BERT
Rodrigo Nogueira, Kyunghyun Cho · 2019
Showed that a BERT model reading the query and passage together reorders a keyword search's results far better, making the retrieve-then-rerank pipeline standard.
How to read it: Short: the method is one page, re-ranking the top 1,000 BM25 passages.
ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
Omar Khattab, Matei Zaharia · 2020
Late interaction: keep one vector per token and match them cheaply, getting close to a cross-encoder's quality at a fraction of the cost.
BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
Nandan Thakur, Nils Reimers et al. · 2021
Showed that retrievers which shine on their training domain often fall behind plain BM25 elsewhere, which is why hybrid search and re-ranking became standard.