Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack
Evaluating Retrieval and RAG
A RAG system is evaluated in two layers: did retrieval find the passages that contain the answer (recall@k, MRR, nDCG on labelled queries), and is the generated answer correct and supported by the passages it cites (faithfulness, citation quality)?
The problem
A wrong answer could come from retrieval, from the model, or from both; one end-to-end score can't tell you which part to fix, and a few good demos prove nothing.
The solution
Build a set of real questions with the passages that answer them; score retrieval with ranking metrics and answers with correctness and support checks, by people, by rules, or by a model judge whose own reliability you check.
The consequence
Retrieval and generation can be improved independently and regressions caught. Model-based judges make evaluation cheap but add their own errors and biases.
You should understand first
- Vectors
- Dot Product
- Embeddings
- Attention
- Probability and Distributions
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Text Embeddings
- Text as Data
- Keyword Search and BM25
- Semantic and Hybrid Search
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- Parametric vs Retrieved Knowledge
- Retrieval-Augmented Generation (RAG)
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Logistic Regression
- Evaluation Metrics for Classifiers
- Evaluating Retrieval and RAG
Layer 1: did retrieval find it?
You need a set of questions and, for each, the passages that answer it. Then for each question rank the collection and ask:
- Recall@k: what fraction of the relevant passages are in the top ? For RAG, "is at least one answer-bearing passage among the we'll put in the prompt" is the number that matters.
- Reciprocal rank: 1 / (rank of the first relevant passage); averaged over queries it's MRR. Rank 1 scores 1, rank 2 scores 0.5, rank 10 scores 0.1.
- nDCG@k: rewards relevant passages more the higher they appear, with a logarithmic discount; it handles several relevant passages and graded relevance.
These are precision and recall from Chapter 3, applied to rankings. Tiny example: three queries whose first relevant passage is at ranks 1, 3 and "not in the top 10" give MRR = (1 + 0.33 + 0) / 3 ≈ 0.44, and recall@5 = 2/3 if each has one relevant passage.
Public benchmarks help choose a model: BEIR covers 18 retrieval datasets across diverse tasks and domains Established and MTEB covers embedding tasks more broadly. Your own questions matter more: a retriever that's best on Wikipedia trivia can be mediocre on your support tickets.
Layer 2: is the answer right and supported?
- Correctness: does it answer the question, matching a reference answer where one exists?
- Faithfulness (groundedness): is every claim supported by the retrieved passages? An answer can be true but unfaithful (the model used its own memory) or faithful but wrong (the source was wrong). For RAG you usually want faithful, so the citations mean something.
- Citation quality: does each cited passage support the sentence it's attached to, and is every claim cited? The ALCE benchmark measures this automatically; on its ELI5 portion even the best models of the time lacked complete citation support 50% of the time Established.
- Abstention: when the sources don't contain the answer, does the system say so?
Frameworks such as Ragas score faithfulness, answer relevance and context relevance without human reference answers by prompting a language model Established. That makes evaluation cheap enough to run on every change, but a model judge has its own errors and biases, so its scores should be checked against human judgments on a sample Interpretation. Chapter 16 returns to evaluating models with models.
Error analysis first
When an answer is wrong, look at the retrieved passages before anything else. If the answer passage wasn't retrieved, no prompt change will fix it; tune chunking, retrieval or reranking. If it was retrieved and the model still got it wrong, look at position, distractors and instructions (context engineering).
What to remember
- Evaluate retrieval and generation separately.
- Recall@k: did the answer passage make the top k? MRR: how high was the first relevant one?
- Faithful (grounded): every claim is supported by the cited passages, whether or not it's true in the world.
- Use many labelled queries from real users; averages, not anecdotes.
- Model judges are models: spot-check them against people.
Key papers
BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
Nandan Thakur, Nils Reimers et al. · 2021
Showed that retrievers which shine on their training domain often fall behind plain BM25 elsewhere, which is why hybrid search and re-ranking became standard.
MTEB: Massive Text Embedding Benchmark
Niklas Muennighoff, Nouamane Tazi et al. · 2022
The benchmark (and public leaderboard) people use to choose an embedding model.
Enabling Large Language Models to Generate Text with Citations
Tianyu Gao, Howard Yen et al. · 2023
A benchmark with automatic metrics for whether a model's citations actually support what it says.
Ragas: Automated Evaluation of Retrieval Augmented Generation
Shahul Es, Jithin James et al. · 2023
A widely used framework for scoring RAG pipelines without human reference answers, using a language model as the judge.