Skip to content
Road to Intelligence

Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack

Evaluating Retrieval and RAG

Must knowKnow well13 minDifficulty

A RAG system is evaluated in two layers: did retrieval find the passages that contain the answer (recall@k, MRR, nDCG on labelled queries), and is the generated answer correct and supported by the passages it cites (faithfulness, citation quality)?

The problem

A wrong answer could come from retrieval, from the model, or from both; one end-to-end score can't tell you which part to fix, and a few good demos prove nothing.

The solution

Build a set of real questions with the passages that answer them; score retrieval with ranking metrics and answers with correctness and support checks, by people, by rules, or by a model judge whose own reliability you check.

The consequence

Retrieval and generation can be improved independently and regressions caught. Model-based judges make evaluation cheap but add their own errors and biases.

Layer 1: did retrieval find it?

You need a set of questions and, for each, the passages that answer it. Then for each question rank the collection and ask:

  • Recall@k: what fraction of the relevant passages are in the top kk? For RAG, "is at least one answer-bearing passage among the kk we'll put in the prompt" is the number that matters.
  • Reciprocal rank: 1 / (rank of the first relevant passage); averaged over queries it's MRR. Rank 1 scores 1, rank 2 scores 0.5, rank 10 scores 0.1.
  • nDCG@k: rewards relevant passages more the higher they appear, with a logarithmic discount; it handles several relevant passages and graded relevance.

These are precision and recall from Chapter 3, applied to rankings. Tiny example: three queries whose first relevant passage is at ranks 1, 3 and "not in the top 10" give MRR = (1 + 0.33 + 0) / 3 ≈ 0.44, and recall@5 = 2/3 if each has one relevant passage.

Public benchmarks help choose a model: BEIR covers 18 retrieval datasets across diverse tasks and domains Established and MTEB covers embedding tasks more broadly. Your own questions matter more: a retriever that's best on Wikipedia trivia can be mediocre on your support tickets.

Layer 2: is the answer right and supported?

  • Correctness: does it answer the question, matching a reference answer where one exists?
  • Faithfulness (groundedness): is every claim supported by the retrieved passages? An answer can be true but unfaithful (the model used its own memory) or faithful but wrong (the source was wrong). For RAG you usually want faithful, so the citations mean something.
  • Citation quality: does each cited passage support the sentence it's attached to, and is every claim cited? The ALCE benchmark measures this automatically; on its ELI5 portion even the best models of the time lacked complete citation support 50% of the time Established.
  • Abstention: when the sources don't contain the answer, does the system say so?

Frameworks such as Ragas score faithfulness, answer relevance and context relevance without human reference answers by prompting a language model Established. That makes evaluation cheap enough to run on every change, but a model judge has its own errors and biases, so its scores should be checked against human judgments on a sample Interpretation. Chapter 16 returns to evaluating models with models.

Error analysis first

When an answer is wrong, look at the retrieved passages before anything else. If the answer passage wasn't retrieved, no prompt change will fix it; tune chunking, retrieval or reranking. If it was retrieved and the model still got it wrong, look at position, distractors and instructions (context engineering).

What to remember

  • Evaluate retrieval and generation separately.
  • Recall@k: did the answer passage make the top k? MRR: how high was the first relevant one?
  • Faithful (grounded): every claim is supported by the cited passages, whether or not it's true in the world.
  • Use many labelled queries from real users; averages, not anecdotes.
  • Model judges are models: spot-check them against people.

Key papers

Important

MTEB: Massive Text Embedding Benchmark

Niklas Muennighoff, Nouamane Tazi et al. · 2022

The benchmark (and public leaderboard) people use to choose an embedding model.

~30 min readarXiv:2210.07316✓ verified 2026-10-05