Skip to content
Road to Intelligence

Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack

Reranking

Should knowKnow well11 minDifficulty

Reranking takes the top candidates from a fast retriever and rescores them with a slower, more accurate model, typically a cross-encoder that reads the query and each passage together.

The problem

Embedding search compresses each passage into one vector before it ever sees the query, so its ordering of near-misses is coarse; reading every passage with the query is accurate but far too slow for a whole collection.

The solution

Two stages: retrieve the top 50–100 cheaply (BM25, dense or hybrid), then run a cross-encoder over each query–passage pair and reorder by its score.

The consequence

Large quality gains for modest cost, because the expensive model only sees a handful of candidates. The first stage still has to find the right passage: a reranker can reorder, never recover what wasn't retrieved.

Reading together versus apart

A text embedding model is a bi-encoder: the passage is summarised into a vector without knowing what will be asked. A cross-encoder puts the query and the passage into one input, [CLS] query [SEP] passage, so every query token can attend to every passage token, and outputs a single relevance score. It can notice that the passage mentions the right model but the wrong year, which a single vector tends to blur.

The cost is that nothing can be precomputed: scoring 10 million passages means 10 million forward passes per query. So it's used as a second stage.

The funnel

  1. Retrieve 50–100 candidates with BM25, dense search or both. Milliseconds.
  2. Rerank those candidates with the cross-encoder. One forward pass each, batched.
  3. Keep the top few for the prompt.

Nogueira and Cho re-ranked the top 1,000 BM25 passages with BERT and reached the top of the MS MARCO passage leaderboard, beating the previous state of the art by 27% (relative) in MRR@10 Established. On BEIR's 18 datasets, a BM25-plus-cross-encoder reranker performed best overall and beat BM25 on 16 of them, but at high computational cost Established.

In between: late interaction

ColBERT encodes the query and the document separately but keeps a vector per token, and scores them with a cheap interaction step; document vectors can be computed offline, and it was competitive with BERT re-rankers while two orders of magnitude faster Established. The price is storage: one vector per token instead of one per passage.

What to remember

  • Bi-encoder: encode separately, compare vectors. Fast, coarse.
  • Cross-encoder: one pass over query + passage together. Accurate, one pass per candidate.
  • Late interaction (ColBERT): one vector per token, cheap max-similarity matching; in between.
  • Retrieve many cheaply, rerank a few expensively.
  • A reranker can't fix a passage the first stage never returned.

Key papers

Important

Passage Re-ranking with BERT

Rodrigo Nogueira, Kyunghyun Cho · 2019

Showed that a BERT model reading the query and passage together reorders a keyword search's results far better, making the retrieve-then-rerank pipeline standard.

How to read it: Short: the method is one page, re-ranking the top 1,000 BM25 passages.

~15 min readarXiv:1901.04085✓ verified 2026-10-05