Skip to content
Road to Intelligence

Part IV · Systems

Chapter 12

Embeddings, RAG & the LLM Application Stack

Giving models knowledge they weren't trained on.

2 h 30 min core path15 concepts4 interactivesCore path · 3 optional concepts foldedDeep · all 15 concepts shown in full

In one sentenceRetrieval-augmented generation finds relevant documents by embedding similarity and places them in the model's context, grounding answers in external knowledge.

The problem

The frozen library

Chapter 11 ended with a model that is fast and cheap to serve. It still has two old problems. What it knows stopped at the end of its training data, and it can't say where any answer came from. Ask about last month's release notes, your company's wiki, or a fact mentioned in only a handful of web pages, and it either doesn't know or guesses fluently.

Facts do get into the weights. Kandpal and colleagues found that a model's accuracy on a factual question tracks how many of its pretraining documents mention the question's entities; BLOOM-176B's accuracy rose from about 25% to above 55% as that count went from 10 to 10,000 Established. Popular facts are learned; the long tail mostly isn't, and scaling barely helps there, while retrieval-augmented models beat far larger unaided ones on less popular facts Established.

This chapter builds the standard answer: keep the knowledge in documents, find the relevant pieces when a question arrives, and hand them to the model to read. We'll build it on a collection you already know, this site's own concept pages, so you can judge the results yourself:

WhatAsk the RoadSection
Documents180 concept pages, about 64,000 wordsCutting documents
Embedding modelall-MiniLM-L6-v2: 384 numbers per passage, reads at most 256 word piecesMeaning as a vector
Keyword searchBM25 over the same passages, run in your browserWords still matter
Benchmark44 everyday questions and bare terms, each labelled with the page that answers itFind by meaning
Rerankerms-marco-MiniLM-L6-v2, a cross-encoderA second opinion

Every number the labs show is computed from real embeddings of these pages, made offline with open models. Nothing generates text: where a language model would read the result, the labs show you exactly what it would read.

Representation

Meaning as a vector

Chapter 6 gave every word a vector. Search needs one vector per passage, so that "how do models reuse earlier work while writing?" can be compared with a paragraph about the KV cache. A text embedding model does this: a BERT-style encoder reads the passage, its token vectors are averaged, and the result is normalised to length 1. Similar meaning, similar direction; similarity is a dot product.

Reimers and Gurevych pointed out that BERT, which compares two sentences by reading them together, needs about 65 hours to find the most similar pair among 10,000 sentences; their Sentence-BERT encodes each sentence once and does it in about 5 seconds Established. That gap, compare-by-reading versus compare-by-vector, is why embeddings run search.

Where does the meaning come from? Encoders are fine-tuned on pairs that should match, such as a question and its answer passage, with contrastive learning: score the true partner against every other passage in the batch, softmax, cross-entropy. It is next-token prediction's loss with the batch as the vocabulary.

OptionalContrastive Learning· folded on the core path. Open it, or switch to Deep to show it here.

The old answer

Words still matter

Long before embeddings, search engines matched words. An inverted index maps each term to the documents that contain it: a database index for words. A query looks up its few terms and scores only those documents. BM25 scores them: rare terms count for more (inverse document frequency), the tenth repetition of a word adds less than the first, and long documents are discounted.

It needs no training, it explains itself, and it is very good at the things a single vector blurs: names, error codes, version numbers, function names. When the BEIR benchmark compared ten retrieval systems on 18 datasets they hadn't been trained for, BM25 was a robust baseline, and many methods that beat it in-domain did poorly elsewhere Established.

Try it

Find by meaning

Semantic search embeds the question and returns the passages whose vectors point most nearly the same way. Dense Passage Retrieval beat a strong BM25 system by 9–19% absolute in top-20 passage accuracy on open-domain question answering Established, and dense retrieval became the default first stage of RAG.

Below, all three run over this site's pages, which are cut into 723 chunks of about 100 words; a page ranks by its best chunk. Start with "Fading signal": keyword search ranks the vanishing-gradients page 13th, because the question shares almost no words with it, while meaning search ranks it first. Then try "ZeRO-3": keyword search finds the sharded-training page immediately; meaning search puts it 31st, since a short identifier's embedding carries little meaning.

Try it

Find by Meaning

Search this site's own concept pages three ways (keywords with BM25, meaning with real embeddings, and both fused) and see where each wins, then score all three on 44 labelled queries.

Know well10 min

Neither method wins everywhere, so hybrid search runs both and merges the rankings with reciprocal rank fusion: each list adds 1/(60 + rank). Agreement between the two lists wins; scores never have to be compared.

Then run the benchmark. Over the 44 labelled queries, hybrid is best on average (mean reciprocal rank 0.83; the right page in the top five for 41 of 44), keyword search is next (0.81; 39), and meaning search alone is last (0.72; 39). That is one small, hand-labelled test on one corpus, not a general ranking. The lesson is the habit, not the numbers: compare methods on many labelled queries from your own users, never on one impressive example.

Scale

Searching millions of vectors

Comparing a question with every stored vector is exact and simple. At 10 million passages of 768 numbers it's 7.7 billion multiply-adds per question, plus about 31 GB of vectors to read. Approximate nearest-neighbour indexes look at a small part instead. An inverted-file (IVF) index clusters the vectors with k-means (Chapter 3 promised this) and scans only the clusters nearest the query. HNSW builds a layered graph of neighbours and searches it greedily from a sparse top layer down, with logarithmic complexity scaling Established. Both trade a little recall for a lot of speed.

Try it

Search Fewer Vectors

Cluster the real chunk embeddings into 32 lists and search only the nearest few: watch recall rise as you scan more of the index.

Understand6 min

On this site's chunks, probing 1 of 32 lists scans about 4% of the vectors and finds under half of the true ten nearest; probing 8 scans about a quarter and finds about nine in ten. The same trade-off, with millions of vectors, is what vector databases tune.

Precision

A second opinion

An embedding summarises a passage before the question exists. A cross-encoder reads the question and the passage together, so every question word can attend to every passage word, and outputs one relevance score. It is far more accurate and far too slow to run on everything, so it runs second: retrieve 50–100 candidates cheaply, rerank them, keep a few. Re-ranking BM25's top 1,000 passages with BERT beat the previous state of the art on MS MARCO by 27% (relative) in MRR@10 Established.

OptionalReranking· folded on the core path. Open it, or switch to Deep to show it here.

The pattern

Retrieve, then generate

Put the pieces together and you have retrieval-augmented generation: two pipelines that meet at the prompt.

A RAG system is two pipelines that meet at the prompt

Runs for every question, in about a second

DrQA (2017) answered questions over all of Wikipedia with TF-IDF retrieval and a neural reader Established. In 2020, Lewis and colleagues named RAG: a pretrained generator conditioned on passages from a dense index of Wikipedia (21 million 100-word chunks), which set the state of the art on three open-domain QA tasks Established. Those systems trained the retriever and generator together. Most of today's don't: simply prepending retrieved documents to an unchanged model's input, with an off-the-shelf retriever, gave gains equivalent to a 2–3× larger model Established. Update a document and the next answer uses it, with no retraining, and the answer can cite where it came from.

Try it

Cutting documents

Before anything is embedded, documents are cut into chunks, and that choice quietly decides what can be found. Small chunks match sharply but can slice an answer in half; large chunks keep answers whole but blur their vectors, cost more prompt space and can exceed what the embedding model reads. On these pages, more than half of the 200-word chunks are longer than the model's 256 word pieces, so their vectors ignore their endings.

Start with the default: Chinchilla's comparison with Gopher, small chunks, three of them. Only 78% of the answer sentence arrives, because a chunk boundary cuts it and the other half isn't retrieved even at k = 5. Switch on overlap and the whole sentence arrives at k = 3; add the reranker and it comes first. Medium chunks hold it whole at k = 1. Then compare the prompt sizes.

Try it

Build the Context

Pick a question, a chunk size, overlap, how many chunks and whether to rerank; see what is retrieved, whether it holds the answer, the exact prompt and what it costs.

Know well12 min

The reranker is usually right, not always: on large chunks it pushes the LoRA answer out of first place. And there is no universal chunk size. The DPR authors found overlapping passages no better than non-overlapping ones in their setting Established, while the LoRA question here needs overlap. Pick a size that fits your embedding model, keep sentences and sections intact, put the document's title on every chunk, and measure.

The prompt

Packing the context

What the model reads is a built artefact: instructions, retrieved passages with their sources, the conversation so far, the question. Every token costs prefill compute and KV-cache memory (Chapter 11): ten retrieved passages of about 150 tokens add about 200 MB of cache per request on Llama 3 8B. And models don't use what they read evenly.

Where the answer sits changes whether it’s used
50%60%70%80%no documents56.1%75.8157.2553.81055.41563.220position of the answer document (of 20)

GPT-3.5-Turbo, multi-document question answering with 20 retrieved documents (about 4,000 tokens). Data: Liu et al. (2023), Tables 1 and 6. A 2023 model; newer models may show smaller effects, so measure your own.

Liu and colleagues moved the answer-bearing document through 20 retrieved documents: GPT-3.5-Turbo was right 75.8% of the time with it first, 53.8% with it in the middle and 63.2% with it last, against 56.1% with no documents at all Established. In the middle, retrieval made it worse than not retrieving. Irrelevant sentences also lower accuracy Established, so retrieving twenty passages "to be safe" can hurt. Rerank, keep the few that matter, and put the best where they are used best.

Why not paste everything into a million-token window? Cost grows with length, and quality doesn't keep up: of 17 models claiming 32K-token contexts or more, only half kept satisfactory performance at 32K on the RULER benchmark Established. Long windows and retrieval are complements: retrieval chooses, a long window lets you include more of what was chosen Interpretation.

Measurement

Did it work?

A RAG system can fail in two places, so measure two layers. Retrieval: for labelled questions, is an answer-bearing passage in the top k (recall@k), and how high is the first one (MRR)? You ran exactly this in the search lab. Generation: is the answer correct, and is it faithful, every claim supported by the passages it cites? On the ALCE citation benchmark's ELI5 portion, even the best models of 2023 lacked complete citation support half the time Established.

Model judges make answer-level evaluation cheap enough to run on every change, but they are models too; check them against people on a sample. And when an answer is wrong, look at what was retrieved first. If the answer wasn't there, no prompt will fix it.

Try it

Talking to software

So far the model reads; an application also needs it to act. Three pieces turn a text generator into a component.

A system prompt sets the standing rules: the task, the constraints, the format, the tools. To the model it is just more tokens, placed first in the chat template (Chapter 10), which post-training taught it to weigh heavily. It is a strong default, not a lock.

Structured outputs make the reply parseable: JSON that matches a schema. Asking usually works and sometimes doesn't, and a pipeline that runs a million times hits "sometimes" constantly. Constrained decoding removes the "sometimes": at each step, zero the tokens that would break the format and renormalise the rest, the same mask-and-renormalise move as top-k and top-p in Chapter 8. Compiling the format into a finite-state machine and indexing the vocabulary against its states makes the allowed set cheap to look up, O(1) on average per step Established.

Tool calling puts structured output to work. The application describes its tools; the model, instead of prose, emits a call such as get_weather with JSON arguments; the application checks and runs it and returns the result as a message; the model continues. OpenAI announced function calling on 13 June 2023: the model chooses to output a JSON object containing arguments to call a described function Established. The model never runs anything itself.

In the lab, a toy model writes a weather-tool call. Sampled freely at temperature 1, about 23% of its calls break (the probabilities are made up, and real models break far less often); constrained, none do. Then ask without a date while the schema requires one: every constrained call is valid and every one contains a date the user never gave. Valid is not correct. Let the schema say null and most calls do; the fix was in the schema, not the decoder.

Try it · toy model

Make It Valid

A toy model writes a weather-tool call token by token. Sample freely and some calls break; mask invalid tokens and none do, but a required field can force the model to invent a value.

Know well8 min

The whole system

The application stack

A production assistant wraps the model in ordinary software:

  1. Check the input

    Is the request in scope? Route it, refuse it, or strip data that shouldn't be logged.
  2. Retrieve

    Rewrite the query if needed; run keyword and vector search; rerank; filter by what this user may see.
  3. Build the context

    System prompt, tool descriptions, the best few passages with sources, recent history, the question, within a token budget.
  4. Generate

    Free text, structured output or a tool call.
  5. Validate and act

    Check the schema and citations; check permissions before running any tool; confirm before anything irreversible.

Each step is code you can test. And each place the model reads outside text is an entry point: Greshake and colleagues showed that instructions planted in web pages likely to be retrieved could take over LLM-integrated applications, including Bing's GPT-4-powered chat Established. Retrieved passages and tool results are data, never instructions; the reliable defence is limiting what the model can do. Chapter 16 returns to attacks. The Model Context Protocol, open-sourced by Anthropic in November 2024, standardises how tools and data sources are offered to models Established.

Deep diveBeyond retrieve-then-readFrontier

Self-RAG trains the model itself to decide when to retrieve and to grade whether passages support its claims Active research. GraphRAG builds an entity graph and community summaries ahead of time so it can answer questions about a whole collection, which a few retrieved chunks can't Active research. HyDE has a model write a hypothetical answer and searches with its embedding Established. RETRO, which retrieves from a 2-trillion-token database, matched GPT-3 on the Pile with 25× fewer parameters Established, a reminder that retrieval can stand in for some parameters.

Why it matters

Why it matters

For a researcher, retrieval separates what a model knows from what it can look up. That raises questions without settled answers: how models combine parametric and retrieved knowledge, when they should retrieve at all, how to attribute an answer to its sources, and how much memorisation retrieval can replace.

For an engineer, this chapter is most of the job. The majority of LLM features that answer from your documents are search plus a prompt plus validation, and most of their failures are search failures. The skills are familiar from data engineering: an indexing pipeline, query plans with a cheap stage and an expensive one, evaluation sets, monitoring.

What depends on it: agents use search and tools in a loop; coding assistants retrieve the right files; enterprise assistants stand or fall on permissions-aware retrieval.

Next

What comes next

Everything here was one round trip: retrieve once, call one tool, answer. Chapter 13 lets the model loop: choose a tool, read the result, decide what to do next, and repeat until a task is done. That loop is an agent, and it inherits every problem in this chapter, multiplied by the number of steps.

Concepts in this chapter

Mark each one as you go. Must-know concepts are the core path.

What do I actually need to remember?

  • A model's weights are a frozen, unsourced snapshot; retrieval adds knowledge at question time and lets answers cite it.
  • A text embedding model turns a passage into one vector; similar meaning gives a high cosine similarity.
  • Embedding models are trained contrastively: pull matching pairs together, push the rest of the batch apart.
  • Keyword search (BM25) wins on names, codes and rare terms; hybrid search fuses both rankings with reciprocal rank fusion.
  • Approximate nearest-neighbour indexes (IVF, HNSW) search millions of vectors by scanning a small part and giving up a little recall.
  • Retrieve cheaply, then rerank the top candidates with a cross-encoder that reads the question and passage together.
  • RAG = retrieve passages, put them in the prompt, generate with citations; most RAG failures are retrieval failures.
  • Chunk size trades completeness against precision; every retrieved token costs prefill and KV cache, and position in the prompt matters.
  • Evaluate retrieval (recall@k, MRR) and answers (correctness, faithfulness, citations) separately, on many labelled questions.
  • Tool calling: the model emits structured arguments and the application runs the tool; constrained decoding guarantees valid syntax, not correct content.

You do not need to memorize everything else. This list is the revision sheet.

Key papers

Important

A Statistical Interpretation of Term Specificity and Its Application in Retrieval

Karen Spärck Jones · 1972 · Journal of Documentation

Introduced the idea behind inverse document frequency: a word that appears in few documents tells you more about them than a word that appears everywhere.

Problem
Matching documents by shared terms treats a common word like 'system' as just as informative as a rare one.
What was new
Weight each term by how specific it is, decreasing with the number of documents that contain it, which became the IDF in TF-IDF and BM25.
~25 min readdoi:10.1108/eb026526✓ verified 2026-10-05
Essential

The Probabilistic Relevance Framework: BM25 and Beyond

Stephen Robertson, Hugo Zaragoza · 2009 · Foundations and Trends in Information Retrieval

The authoritative account of BM25, the keyword-ranking function that is still the standard baseline, and often a component, of modern retrieval systems.

Problem
Ranking by raw term counts over-rewards long documents and repeated words.
What was new
Derives BM25 from a probabilistic model of relevance: IDF times a term frequency that saturates, normalised for document length, with two tunable parameters k1 and b.

How to read it: Section 3 derives BM25; its discussion of parameters notes that 1.2 < k1 < 2 and 0.5 < b < 0.8 are reasonable in many settings.

~1 h 30 min readdoi:10.1561/1500000019✓ verified 2026-10-05
Optional

Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods

Gordon V. Cormack, Charles L. A. Clarke, Stefan Buettcher · 2009 · SIGIR

A two-page paper whose one-line formula is how most hybrid search systems merge keyword and vector results.

Problem
Different retrieval systems produce scores on incomparable scales, so their results are hard to combine.
What was new
Ignore the scores and add 1 / (k + rank) across rankings; k = 60 was near-optimal in pilot experiments, though the choice was not critical.
~10 min readdoi:10.1145/1571941.1572114✓ verified 2026-10-05
Important

Product Quantization for Nearest Neighbor Search

Hervé Jégou, Matthijs Douze, Cordelia Schmid · 2011 · IEEE TPAMI

Showed how to compress high-dimensional vectors to a few bytes each and still estimate distances, the basis of billion-scale vector search.

Problem
Storing and scanning millions of full-precision vectors costs too much memory and time.
What was new
Split each vector into sub-vectors, replace each by the nearest of a small learned codebook, and compute distances from precomputed tables.
~1 h readdoi:10.1109/TPAMI.2010.57✓ verified 2026-10-05
Essential

Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs

Yu. A. Malkov, D. A. Yashunin · 2016

HNSW is the graph index behind most vector databases: fast, accurate approximate search over millions of embeddings.

Problem
Exact nearest-neighbour search compares the query with every stored vector, which is too slow at scale.
What was new
A layered proximity graph, like a skip list: each element joins a random number of layers, and a search descends greedily from the sparse top layer to the dense bottom one, with logarithmic complexity scaling.

How to read it: Section 4 (algorithm description) is the core; Figure 1 shows the layered search.

~45 min readarXiv:1603.09320✓ verified 2026-10-05
Important

Billion-scale similarity search with GPUs

Jeff Johnson, Matthijs Douze, Hervé Jégou · 2017

The paper behind FAISS, the library that made exact and compressed vector search fast on GPUs and that many RAG systems still use.

Problem
GPU similarity search was bottlenecked by selecting the top k results and by poor use of the memory hierarchy.
What was new
A k-selection design running at up to 55% of theoretical peak, giving nearest-neighbour search 8.5× faster than the previous GPU state of the art, for brute-force and product-quantized indexes.
~45 min readarXiv:1702.08734✓ verified 2026-10-05
Important

Reading Wikipedia to Answer Open-Domain Questions

Danqi Chen, Adam Fisch et al. · 2017

Set the retrieve-then-read template: a search component finds Wikipedia articles, and a neural reader extracts the answer from them.

Problem
Reading-comprehension models assumed the right paragraph was handed to them; real questions come with no paragraph.
What was new
DrQA combined bigram-hashing TF-IDF retrieval over all of Wikipedia with a recurrent reader trained to find answer spans.
~30 min readarXiv:1704.00051✓ verified 2026-10-05
Optional

Representation Learning with Contrastive Predictive Coding

Aaron van den Oord, Yazhe Li, Oriol Vinyals · 2018

Introduced the InfoNCE loss: pick the true partner out of a set of negatives with a softmax, the objective most embedding models are trained with.

Problem
Learning useful representations without labels needs a training signal that doesn't require predicting every detail of the input.
What was new
A probabilistic contrastive loss with negative sampling: score the true future sample against distractors and train with cross-entropy.
~40 min readarXiv:1807.03748✓ verified 2026-10-05
Important

Passage Re-ranking with BERT

Rodrigo Nogueira, Kyunghyun Cho · 2019

Showed that a BERT model reading the query and passage together reorders a keyword search's results far better, making the retrieve-then-rerank pipeline standard.

Problem
Fast first-stage retrieval like BM25 finds candidates but orders them poorly.
What was new
Feed query and passage as one input to BERT and score relevance; on MS MARCO it beat the previous state of the art by 27% (relative) in MRR@10.

How to read it: Short: the method is one page, re-ranking the top 1,000 BM25 passages.

~15 min readarXiv:1901.04085✓ verified 2026-10-05
Essential

Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Nils Reimers, Iryna Gurevych · 2019

Made BERT produce one vector per sentence that can be compared with cosine similarity, which is what turned Transformers into practical search engines.

Problem
BERT compares two sentences by reading them together; finding the most similar pair among 10,000 sentences takes about 50 million passes, roughly 65 hours.
What was new
Fine-tune BERT in a siamese setup with mean pooling so each sentence is encoded once; the same search takes about 5 seconds.

How to read it: Section 3 (the model) explains the siamese network and pooling in a page.

~30 min readarXiv:1908.10084✓ verified 2026-10-05
Important

Language Models as Knowledge Bases?

Fabio Petroni, Tim Rocktäschel et al. · 2019

Asked how much factual knowledge a pretrained model holds in its weights, framing the question that retrieval later answered differently.

Problem
Knowledge bases need hand-built schemas; pretrained language models might already store relational facts.
What was new
Probe models with fill-in-the-blank statements: without fine-tuning, BERT held relational knowledge competitive with some traditional methods, and some kinds of facts were learned far more readily than others.
~25 min readarXiv:1909.01066✓ verified 2026-10-05
Important

REALM: Retrieval-Augmented Language Model Pre-Training

Kelvin Guu, Kenton Lee et al. · 2020

Trained a retriever and a language model together, so knowledge could live in a searchable corpus instead of only in the weights.

Problem
Knowledge stored implicitly in parameters requires ever-larger networks to cover more facts, and is hard to inspect.
What was new
Add a latent retriever over Wikipedia to masked-language-model pretraining and backpropagate through the retrieval step; it beat previous open-domain QA methods by 4–16% absolute accuracy.
~45 min readarXiv:2002.08909✓ verified 2026-10-05
Essential

Dense Passage Retrieval for Open-Domain Question Answering

Vladimir Karpukhin, Barlas Oğuz et al. · 2020

Showed that learned embeddings alone can beat BM25 for finding answer passages, which made dense retrieval the default first stage of RAG.

Problem
Open-domain QA retrieved passages with TF-IDF or BM25, which miss passages that use different words for the same thing.
What was new
Two BERT encoders, one for questions and one for passages, trained with in-batch negatives; it beat a strong BM25 system by 9–19% absolute in top-20 passage accuracy.

How to read it: Section 3 (the dual encoder and in-batch negatives) is the core; Section 4.1 describes splitting Wikipedia into 21 million 100-word passages.

~40 min readarXiv:2004.04906✓ verified 2026-10-05
Important

ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT

Omar Khattab, Matei Zaharia · 2020

Late interaction: keep one vector per token and match them cheaply, getting close to a cross-encoder's quality at a fraction of the cost.

Problem
BERT re-rankers must run the full network on every query–passage pair, orders of magnitude more expensive than earlier methods.
What was new
Encode query and passage separately into token vectors and score with a sum of maximum similarities; passages can be encoded offline. It was two orders of magnitude faster with four orders fewer FLOPs per query than BERT re-rankers, at competitive quality.
~40 min readarXiv:2004.12832✓ verified 2026-10-05
Essential

Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Patrick Lewis, Ethan Perez et al. · 2020

Named retrieval-augmented generation: a generator conditioned on passages fetched from a dense index, with knowledge that can be inspected and updated without retraining.

Problem
Knowledge in a model's weights is hard to update, can't point to its sources, and lags task-specific systems on knowledge-intensive tasks.
What was new
Combine a pretrained seq2seq generator (BART) with a dense vector index of Wikipedia searched by a DPR retriever, trained end to end; it set the state of the art on three open-domain QA tasks.

How to read it: Section 2 (methods: the DPR retriever, the BART generator and the two ways of combining passages) is the core; the index is 21 million 100-word chunks of a December 2018 Wikipedia dump.

~45 min readarXiv:2005.11401✓ verified 2026-10-05
Important

BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models

Nandan Thakur, Nils Reimers et al. · 2021

Showed that retrievers which shine on their training domain often fall behind plain BM25 elsewhere, which is why hybrid search and re-ranking became standard.

Problem
Neural retrievers were evaluated in narrow, in-domain settings, hiding how they generalise.
What was new
18 datasets across diverse tasks and domains, 10 systems compared zero-shot: BM25 was a robust baseline; re-ranking and late interaction did best on average, at high computational cost.
~40 min readarXiv:2104.08663✓ verified 2026-10-05
Important

Improving language models by retrieving from trillions of tokens

Sebastian Borgeaud, Arthur Mensch et al. · 2021

A language model that looks things up in a 2-trillion-token database matched much larger models, evidence that retrieval can substitute for some parameters.

Problem
Scaling parameters is an expensive way to store more knowledge.
What was new
RETRO retrieves neighbouring chunks with a frozen BERT retriever and attends to them through chunked cross-attention; it matched GPT-3 and Jurassic-1 on the Pile with 25× fewer parameters.
~1 h readarXiv:2112.04426✓ verified 2026-10-05
Important

MTEB: Massive Text Embedding Benchmark

Niklas Muennighoff, Nouamane Tazi et al. · 2022

The benchmark (and public leaderboard) people use to choose an embedding model.

Problem
Embedding models were evaluated on a few datasets from a single task, so it was unclear how well they transfer.
What was new
8 tasks, 58 datasets, 112 languages and 33 models: no single embedding method dominated across all tasks.
~30 min readarXiv:2210.07316✓ verified 2026-10-05
Important

Large Language Models Struggle to Learn Long-Tail Knowledge

Nikhil Kandpal, Haikang Deng et al. · 2022

Measured that a model's chance of answering a factual question tracks how many pretraining documents discuss it, which is why rare facts are where retrieval helps most.

Problem
It was unclear which facts a language model can learn from its pretraining data, and why it fails on others.
What was new
Entity-link pretraining corpora and count relevant documents per question: strong correlational and causal links between that count and accuracy, across models up to 176B parameters.

How to read it: Section 3 (correlational and causal analysis): BLOOM-176B's accuracy rises from about 25% to above 55% as relevant documents go from 10 to 10,000.

~30 min readarXiv:2211.08411✓ verified 2026-10-05
Important

Text Embeddings by Weakly-Supervised Contrastive Pre-training

Liang Wang, Nan Yang et al. · 2022

E5 showed the modern recipe for general-purpose embedding models: contrastive pretraining on huge numbers of naturally occurring text pairs, then fine-tuning.

Problem
Good embedding models needed labelled relevance data, and zero-shot dense retrievers still lost to BM25.
What was new
Contrastive training on a curated web-scale dataset of text pairs; the first model to beat BM25 on the BEIR benchmark zero-shot without labelled data.
~35 min readarXiv:2212.03533✓ verified 2026-10-05
Optional

Precise Zero-Shot Dense Retrieval without Relevance Labels

Luyu Gao, Xueguang Ma et al. · 2022

HyDE: let a language model write a hypothetical answer and search with its embedding, since answers look more like documents than questions do.

Problem
Zero-shot dense retrieval is hard when no relevance labels exist to train on.
What was new
Generate a hypothetical document for the query (it may contain false details), embed it, and retrieve the real documents nearest to it.
~25 min readarXiv:2212.10496✓ verified 2026-10-05
Important

When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories

Alex Mallen, Akari Asai et al. · 2022

Showed that retrieval matters most for less popular facts and that scaling barely helps there, suggesting retrieval only when it's needed.

Problem
When should a system trust a model's own memory, and when should it look things up?
What was new
PopQA, 14k questions spanning popular to obscure entities: models struggle with less popular facts, retrieval-augmented models beat far larger unaided ones, and retrieving only when needed improves results at lower cost.
~35 min readarXiv:2212.10511✓ verified 2026-10-05
Important

In-Context Retrieval-Augmented Language Models

Ori Ram, Yoav Levine et al. · 2023

Showed that you don't need to change or retrain the model: just put retrieved documents in front of the input. That's how most RAG systems work today.

Problem
Earlier retrieval-augmented models changed the architecture, which complicated deployment.
What was new
Prepend retrieved documents to a frozen model's input; with off-the-shelf retrievers the gains were equivalent to a 2–3× larger model.

How to read it: The 'Our framework' section is two pages and describes the whole method.

~35 min readarXiv:2302.00083✓ verified 2026-10-05
Optional

Large Language Models Can Be Easily Distracted by Irrelevant Context

Freda Shi, Xinyun Chen et al. · 2023

A clean demonstration that adding irrelevant text to a prompt makes models worse, which is why retrieving more is not always better.

Problem
Benchmarks give models only relevant information; real contexts contain distractions.
What was new
GSM-IC, grade-school maths problems with an irrelevant sentence added: accuracy dropped sharply; self-consistency and an instruction to ignore irrelevant information helped.
~20 min readarXiv:2302.00093✓ verified 2026-10-05
Essential

Toolformer: Language Models Can Teach Themselves to Use Tools

Timo Schick, Jane Dwivedi-Yu et al. · 2023

Showed a model can learn when to call a calculator, search engine or calendar and how to use the result, the core idea behind tool calling.

Problem
Language models fail at things simple tools do well, like arithmetic and looking up facts.
What was new
Self-supervised: the model inserts candidate API calls into text, keeps the ones whose results make the following tokens easier to predict, and is fine-tuned on them. It improved zero-shot performance, often competitive with much larger models.
~40 min readarXiv:2302.04761✓ verified 2026-10-05
Important

Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection

Kai Greshake, Sahar Abdelnabi et al. · 2023

Showed that text a model retrieves (a web page, an email, a document) can carry instructions an attacker planted, so retrieval and tools are a security boundary.

Problem
Prompt injection was thought of as a user attacking their own chat; applications that read outside data blur data and instructions.
What was new
Indirect prompt injection: plant prompts in data likely to be retrieved. Demonstrated against real systems including Bing's GPT-4-powered chat, with a taxonomy of impacts such as data theft.
~40 min readarXiv:2302.12173✓ verified 2026-10-05
Important

Enabling Large Language Models to Generate Text with Citations

Tianyu Gao, Howard Yen et al. · 2023

A benchmark with automatic metrics for whether a model's citations actually support what it says.

Problem
Citation quality was judged by hand with commercial search engines, hard to reproduce or compare.
What was new
ALCE: questions plus retrieval corpora, scored for fluency, correctness and citation quality; on ELI5 even the best models lacked complete citation support 50% of the time.
~35 min readarXiv:2305.14627✓ verified 2026-10-05
Essential

Lost in the Middle: How Language Models Use Long Contexts

Nelson F. Liu, Kevin Lin et al. · 2023

Showed that where a passage sits in the prompt changes whether the model uses it: best at the start or end, worst in the middle.

Problem
Models accepted long contexts, but little was known about how well they used them.
What was new
Move the answer-bearing document through 10, 20 or 30 retrieved documents: accuracy is U-shaped, and in the worst case GPT-3.5-Turbo did worse than with no documents at all.

How to read it: Section 2 (multi-document QA) and its Figure 5 hold the main result; Appendix G tabulates the numbers.

~30 min readarXiv:2307.03172✓ verified 2026-10-05
Important

Efficient Guided Generation for Large Language Models

Brandon T. Willard, Rémi Louf · 2023

Made constrained decoding cheap: precompute which tokens are allowed in each state of a regular expression or grammar, so structured output can be guaranteed.

Problem
Masking invalid tokens naively means checking the whole vocabulary at every step.
What was new
Turn the pattern into a finite-state machine and index the vocabulary against its states once; at generation time the allowed tokens cost O(1) on average to look up.

How to read it: The section on iterative FSM processing and indexing, with Figure 1's floating-point-number example, is the core.

~25 min readarXiv:2307.09702✓ verified 2026-10-05
Optional

Ragas: Automated Evaluation of Retrieval Augmented Generation

Shahul Es, Jithin James et al. · 2023

A widely used framework for scoring RAG pipelines without human reference answers, using a language model as the judge.

Problem
Evaluating RAG needs judgments of retrieval relevance, faithfulness and answer quality, and human annotations are slow.
What was new
Reference-free metrics such as faithfulness, answer relevance and context relevance, each computed by prompting a language model.
~20 min readarXiv:2309.15217✓ verified 2026-10-05
Optional

Retrieval meets Long Context Large Language Models

Peng Xu, Wei Ping et al. · 2023

A direct comparison of the two ways to give a model more information: retrieval or a longer context window. They found retrieval helps even when the window is long.

Problem
Should you extend the context window or retrieve, and can the two be combined?
What was new
A 4K-context model with simple retrieval matched a 16K-context fine-tuned model on long-context tasks with much less computation, and retrieval improved models regardless of window size.
~30 min readarXiv:2310.03025✓ verified 2026-10-05
Optional

LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Huiqiang Jiang, Qianhui Wu et al. · 2023

Prompt compression: drop the tokens a small model finds predictable before sending a long prompt to a large one.

Problem
Prompts with retrieved documents and examples grow to tens of thousands of tokens, which is slow and expensive.
What was new
Coarse-to-fine compression with a budget controller and token-level iterative pruning; up to 20× compression with little performance loss on the tasks tested.
~30 min readarXiv:2310.05736✓ verified 2026-10-05
Optional

Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection

Akari Asai, Zeqiu Wu et al. · 2023

Trains the model itself to decide when to retrieve and to grade whether passages and its own claims are supported.

Problem
Always retrieving a fixed number of passages, relevant or not, can make answers worse.
What was new
Special reflection tokens let the model retrieve on demand and critique retrieved passages and its own output; it improved factuality and citation accuracy over the baselines tested.
~40 min readarXiv:2310.11511✓ verified 2026-10-05
Optional

RULER: What's the Real Context Size of Your Long-Context Language Models?

Cheng-Ping Hsieh, Simeng Sun et al. · 2024

Showed that a model's advertised context length overstates the length it can actually use well.

Problem
Needle-in-a-haystack tests only check simple retrieval from a long context.
What was new
Synthetic tasks with multiple needles, multi-hop tracing and aggregation: of 17 models claiming 32K tokens or more, only half kept satisfactory performance at 32K.
~30 min readarXiv:2404.06654✓ verified 2026-10-05
Optional

From Local to Global: A Graph RAG Approach to Query-Focused Summarization

Darren Edge, Ha Trinh et al. · 2024

Addresses a blind spot of chunk retrieval: questions about a whole collection, like 'what are the main themes?'.

Problem
RAG retrieves a few passages, so it fails on global questions that need the entire corpus.
What was new
Use an LLM to build an entity graph, summarise communities of related entities in advance, then answer from those summaries; it improved comprehensiveness and diversity on such questions over a conventional RAG baseline.
~40 min readarXiv:2404.16130✓ verified 2026-10-05

Watch

1 h 19 min

Stanford Online

Stanford CS25: V3 I Retrieval Augmented Language Models

Douwe Kiela, one of the authors of the original RAG paper, surveys retrieval-augmented language models and ends with the open questions.

Covers: Why language models need retrieval, the main families of retrieval-augmented models from the recent literature, and the questions that remain open.

Should know
1 h

Andrej Karpathy

[1hr Talk] Intro to Large Language Models

A clear one-hour overview of what LLMs are, how they are trained, and where they're going — good orientation for Part III.

Covers: Pretraining, fine-tuning, scaling, tool use, security issues.

Must know
3 h 31 min

Andrej Karpathy

Deep Dive into LLMs like ChatGPT

A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.

Covers: The full training stack behind a chat model, from internet text and tokens to a pretrained base model and the post-training that turns it into an assistant.

Should know

What came next?

Chapter 13

Agents →

Generating text isn't the same as getting something done.