Part IV · Systems
Chapter 12
Embeddings, RAG & the LLM Application Stack
Giving models knowledge they weren't trained on.
In one sentenceRetrieval-augmented generation finds relevant documents by embedding similarity and places them in the model's context, grounding answers in external knowledge.
The problem
The frozen library
Chapter 11 ended with a model that is fast and cheap to serve. It still has two old problems. What it knows stopped at the end of its training data, and it can't say where any answer came from. Ask about last month's release notes, your company's wiki, or a fact mentioned in only a handful of web pages, and it either doesn't know or guesses fluently.
Facts do get into the weights. Kandpal and colleagues found that a model's accuracy on a factual question tracks how many of its pretraining documents mention the question's entities; BLOOM-176B's accuracy rose from about 25% to above 55% as that count went from 10 to 10,000 Established. Popular facts are learned; the long tail mostly isn't, and scaling barely helps there, while retrieval-augmented models beat far larger unaided ones on less popular facts Established.
This chapter builds the standard answer: keep the knowledge in documents, find the relevant pieces when a question arrives, and hand them to the model to read. We'll build it on a collection you already know, this site's own concept pages, so you can judge the results yourself:
| What | Ask the Road | Section |
|---|---|---|
| Documents | 180 concept pages, about 64,000 words | Cutting documents |
| Embedding model | all-MiniLM-L6-v2: 384 numbers per passage, reads at most 256 word pieces | Meaning as a vector |
| Keyword search | BM25 over the same passages, run in your browser | Words still matter |
| Benchmark | 44 everyday questions and bare terms, each labelled with the page that answers it | Find by meaning |
| Reranker | ms-marco-MiniLM-L6-v2, a cross-encoder | A second opinion |
Every number the labs show is computed from real embeddings of these pages, made offline with open models. Nothing generates text: where a language model would read the result, the labs show you exactly what it would read.
Representation
Meaning as a vector
Chapter 6 gave every word a vector. Search needs one vector per passage, so that "how do models reuse earlier work while writing?" can be compared with a paragraph about the KV cache. A text embedding model does this: a BERT-style encoder reads the passage, its token vectors are averaged, and the result is normalised to length 1. Similar meaning, similar direction; similarity is a dot product.
Reimers and Gurevych pointed out that BERT, which compares two sentences by reading them together, needs about 65 hours to find the most similar pair among 10,000 sentences; their Sentence-BERT encodes each sentence once and does it in about 5 seconds Established. That gap, compare-by-reading versus compare-by-vector, is why embeddings run search.
Where does the meaning come from? Encoders are fine-tuned on pairs that should match, such as a question and its answer passage, with contrastive learning: score the true partner against every other passage in the batch, softmax, cross-entropy. It is next-token prediction's loss with the batch as the vocabulary.
The old answer
Words still matter
Long before embeddings, search engines matched words. An inverted index maps each term to the documents that contain it: a database index for words. A query looks up its few terms and scores only those documents. BM25 scores them: rare terms count for more (inverse document frequency), the tenth repetition of a word adds less than the first, and long documents are discounted.
It needs no training, it explains itself, and it is very good at the things a single vector blurs: names, error codes, version numbers, function names. When the BEIR benchmark compared ten retrieval systems on 18 datasets they hadn't been trained for, BM25 was a robust baseline, and many methods that beat it in-domain did poorly elsewhere Established.
Try it
Find by meaning
Semantic search embeds the question and returns the passages whose vectors point most nearly the same way. Dense Passage Retrieval beat a strong BM25 system by 9–19% absolute in top-20 passage accuracy on open-domain question answering Established, and dense retrieval became the default first stage of RAG.
Below, all three run over this site's pages, which are cut into 723 chunks of about 100 words; a page ranks by its best chunk. Start with "Fading signal": keyword search ranks the vanishing-gradients page 13th, because the question shares almost no words with it, while meaning search ranks it first. Then try "ZeRO-3": keyword search finds the sharded-training page immediately; meaning search puts it 31st, since a short identifier's embedding carries little meaning.
Try it
Search this site's own concept pages three ways (keywords with BM25, meaning with real embeddings, and both fused) and see where each wins, then score all three on 44 labelled queries.
Neither method wins everywhere, so hybrid search runs both and merges the rankings with reciprocal rank fusion: each list adds 1/(60 + rank). Agreement between the two lists wins; scores never have to be compared.
Then run the benchmark. Over the 44 labelled queries, hybrid is best on average (mean reciprocal rank 0.83; the right page in the top five for 41 of 44), keyword search is next (0.81; 39), and meaning search alone is last (0.72; 39). That is one small, hand-labelled test on one corpus, not a general ranking. The lesson is the habit, not the numbers: compare methods on many labelled queries from your own users, never on one impressive example.
Scale
Searching millions of vectors
Comparing a question with every stored vector is exact and simple. At 10 million passages of 768 numbers it's 7.7 billion multiply-adds per question, plus about 31 GB of vectors to read. Approximate nearest-neighbour indexes look at a small part instead. An inverted-file (IVF) index clusters the vectors with k-means (Chapter 3 promised this) and scans only the clusters nearest the query. HNSW builds a layered graph of neighbours and searches it greedily from a sparse top layer down, with logarithmic complexity scaling Established. Both trade a little recall for a lot of speed.
Try it
Cluster the real chunk embeddings into 32 lists and search only the nearest few: watch recall rise as you scan more of the index.
On this site's chunks, probing 1 of 32 lists scans about 4% of the vectors and finds under half of the true ten nearest; probing 8 scans about a quarter and finds about nine in ten. The same trade-off, with millions of vectors, is what vector databases tune.
Precision
A second opinion
An embedding summarises a passage before the question exists. A cross-encoder reads the question and the passage together, so every question word can attend to every passage word, and outputs one relevance score. It is far more accurate and far too slow to run on everything, so it runs second: retrieve 50–100 candidates cheaply, rerank them, keep a few. Re-ranking BM25's top 1,000 passages with BERT beat the previous state of the art on MS MARCO by 27% (relative) in MRR@10 Established.
The pattern
Retrieve, then generate
Put the pieces together and you have retrieval-augmented generation: two pipelines that meet at the prompt.
Runs for every question, in about a second
DrQA (2017) answered questions over all of Wikipedia with TF-IDF retrieval and a neural reader Established. In 2020, Lewis and colleagues named RAG: a pretrained generator conditioned on passages from a dense index of Wikipedia (21 million 100-word chunks), which set the state of the art on three open-domain QA tasks Established. Those systems trained the retriever and generator together. Most of today's don't: simply prepending retrieved documents to an unchanged model's input, with an off-the-shelf retriever, gave gains equivalent to a 2–3× larger model Established. Update a document and the next answer uses it, with no retraining, and the answer can cite where it came from.
Try it
Cutting documents
Before anything is embedded, documents are cut into chunks, and that choice quietly decides what can be found. Small chunks match sharply but can slice an answer in half; large chunks keep answers whole but blur their vectors, cost more prompt space and can exceed what the embedding model reads. On these pages, more than half of the 200-word chunks are longer than the model's 256 word pieces, so their vectors ignore their endings.
Start with the default: Chinchilla's comparison with Gopher, small chunks, three of them. Only 78% of the answer sentence arrives, because a chunk boundary cuts it and the other half isn't retrieved even at k = 5. Switch on overlap and the whole sentence arrives at k = 3; add the reranker and it comes first. Medium chunks hold it whole at k = 1. Then compare the prompt sizes.
Try it
Pick a question, a chunk size, overlap, how many chunks and whether to rerank; see what is retrieved, whether it holds the answer, the exact prompt and what it costs.
The reranker is usually right, not always: on large chunks it pushes the LoRA answer out of first place. And there is no universal chunk size. The DPR authors found overlapping passages no better than non-overlapping ones in their setting Established, while the LoRA question here needs overlap. Pick a size that fits your embedding model, keep sentences and sections intact, put the document's title on every chunk, and measure.
The prompt
Packing the context
What the model reads is a built artefact: instructions, retrieved passages with their sources, the conversation so far, the question. Every token costs prefill compute and KV-cache memory (Chapter 11): ten retrieved passages of about 150 tokens add about 200 MB of cache per request on Llama 3 8B. And models don't use what they read evenly.
GPT-3.5-Turbo, multi-document question answering with 20 retrieved documents (about 4,000 tokens). Data: Liu et al. (2023), Tables 1 and 6. A 2023 model; newer models may show smaller effects, so measure your own.
Liu and colleagues moved the answer-bearing document through 20 retrieved documents: GPT-3.5-Turbo was right 75.8% of the time with it first, 53.8% with it in the middle and 63.2% with it last, against 56.1% with no documents at all Established. In the middle, retrieval made it worse than not retrieving. Irrelevant sentences also lower accuracy Established, so retrieving twenty passages "to be safe" can hurt. Rerank, keep the few that matter, and put the best where they are used best.
Why not paste everything into a million-token window? Cost grows with length, and quality doesn't keep up: of 17 models claiming 32K-token contexts or more, only half kept satisfactory performance at 32K on the RULER benchmark Established. Long windows and retrieval are complements: retrieval chooses, a long window lets you include more of what was chosen Interpretation.
Measurement
Did it work?
A RAG system can fail in two places, so measure two layers. Retrieval: for labelled questions, is an answer-bearing passage in the top k (recall@k), and how high is the first one (MRR)? You ran exactly this in the search lab. Generation: is the answer correct, and is it faithful, every claim supported by the passages it cites? On the ALCE citation benchmark's ELI5 portion, even the best models of 2023 lacked complete citation support half the time Established.
Model judges make answer-level evaluation cheap enough to run on every change, but they are models too; check them against people on a sample. And when an answer is wrong, look at what was retrieved first. If the answer wasn't there, no prompt will fix it.
Try it
Talking to software
So far the model reads; an application also needs it to act. Three pieces turn a text generator into a component.
A system prompt sets the standing rules: the task, the constraints, the format, the tools. To the model it is just more tokens, placed first in the chat template (Chapter 10), which post-training taught it to weigh heavily. It is a strong default, not a lock.
Structured outputs make the reply parseable: JSON that matches a schema. Asking usually works and sometimes doesn't, and a pipeline that runs a million times hits "sometimes" constantly. Constrained decoding removes the "sometimes": at each step, zero the tokens that would break the format and renormalise the rest, the same mask-and-renormalise move as top-k and top-p in Chapter 8. Compiling the format into a finite-state machine and indexing the vocabulary against its states makes the allowed set cheap to look up, O(1) on average per step Established.
Tool calling puts structured output to work. The application describes its tools; the model, instead of prose, emits a call such as get_weather with JSON arguments; the application checks and runs it and returns the result as a message; the model continues. OpenAI announced function calling on 13 June 2023: the model chooses to output a JSON object containing arguments to call a described function Established. The model never runs anything itself.
In the lab, a toy model writes a weather-tool call. Sampled freely at temperature 1, about 23% of its calls break (the probabilities are made up, and real models break far less often); constrained, none do. Then ask without a date while the schema requires one: every constrained call is valid and every one contains a date the user never gave. Valid is not correct. Let the schema say null and most calls do; the fix was in the schema, not the decoder.
Try it · toy model
A toy model writes a weather-tool call token by token. Sample freely and some calls break; mask invalid tokens and none do, but a required field can force the model to invent a value.
The whole system
The application stack
A production assistant wraps the model in ordinary software:
Check the input
Is the request in scope? Route it, refuse it, or strip data that shouldn't be logged.Retrieve
Rewrite the query if needed; run keyword and vector search; rerank; filter by what this user may see.Build the context
System prompt, tool descriptions, the best few passages with sources, recent history, the question, within a token budget.Generate
Free text, structured output or a tool call.Validate and act
Check the schema and citations; check permissions before running any tool; confirm before anything irreversible.
Each step is code you can test. And each place the model reads outside text is an entry point: Greshake and colleagues showed that instructions planted in web pages likely to be retrieved could take over LLM-integrated applications, including Bing's GPT-4-powered chat Established. Retrieved passages and tool results are data, never instructions; the reliable defence is limiting what the model can do. Chapter 16 returns to attacks. The Model Context Protocol, open-sourced by Anthropic in November 2024, standardises how tools and data sources are offered to models Established.
Deep diveBeyond retrieve-then-readFrontier
Self-RAG trains the model itself to decide when to retrieve and to grade whether passages support its claims Active research. GraphRAG builds an entity graph and community summaries ahead of time so it can answer questions about a whole collection, which a few retrieved chunks can't Active research. HyDE has a model write a hypothetical answer and searches with its embedding Established. RETRO, which retrieves from a 2-trillion-token database, matched GPT-3 on the Pile with 25× fewer parameters Established, a reminder that retrieval can stand in for some parameters.
Why it matters
Why it matters
For a researcher, retrieval separates what a model knows from what it can look up. That raises questions without settled answers: how models combine parametric and retrieved knowledge, when they should retrieve at all, how to attribute an answer to its sources, and how much memorisation retrieval can replace.
For an engineer, this chapter is most of the job. The majority of LLM features that answer from your documents are search plus a prompt plus validation, and most of their failures are search failures. The skills are familiar from data engineering: an indexing pipeline, query plans with a cheap stage and an expensive one, evaluation sets, monitoring.
What depends on it: agents use search and tools in a loop; coding assistants retrieve the right files; enterprise assistants stand or fall on permissions-aware retrieval.
Next
What comes next
Everything here was one round trip: retrieve once, call one tool, answer. Chapter 13 lets the model loop: choose a tool, read the result, decide what to do next, and repeat until a task is done. That loop is an agent, and it inherits every problem in this chapter, multiplied by the number of steps.
From a frozen model to a connected one
Concepts in this chapter
Mark each one as you go. Must-know concepts are the core path.
- Approximate Nearest-Neighbour SearchApproximate nearest-neighbour indexes find the vectors most similar to a query without comparing it to every stored vector, by searching only promising clusters (IVF) or walking a proximity graph (HNSW), giving up a little recall for a large speed-up.Know wellMust know
- ChunkingChunking splits documents into passages before embedding them, and the chunk size sets a trade-off: small chunks are precise but can cut an answer in half, large chunks keep context but blur their vectors and cost more prompt space.Know wellMust know
- Context EngineeringContext engineering is deciding what goes into the model's limited context window, in what form and in what order (instructions, retrieved passages, conversation history, tool results), because the model can only use what it reads, and doesn't use everything it reads equally well.Know wellMust know
- Keyword Search and BM25Keyword search ranks documents by the query words they contain, weighting rare words more (inverse document frequency) and letting repeated words count less and less; BM25 is the standard formula, and it is still a strong baseline.Know wellMust know
- Parametric vs Retrieved KnowledgeWhat a language model knows from training is stored in its weights (parametric knowledge): frozen at the training cutoff, unreliable for rare facts and unable to name its source; retrieved knowledge is read from documents at question time instead.UnderstandMust know
- Evaluating Retrieval and RAGA RAG system is evaluated in two layers: did retrieval find the passages that contain the answer (recall@k, MRR, nDCG on labelled queries), and is the generated answer correct and supported by the passages it cites (faithfulness, citation quality)?Know wellMust know
- Retrieval-Augmented Generation (RAG)Retrieval-augmented generation answers a question by first searching a document collection for relevant passages and then giving those passages to a language model in its prompt, so the answer can use, and cite, knowledge that isn't in the model's weights.Know wellMust know
- Semantic and Hybrid SearchSemantic (dense) search embeds the query and returns the stored passages whose vectors are most similar to it, finding matches that share meaning but not words; hybrid search runs keyword and semantic search together and merges their rankings.Know wellMust know
- Structured Outputs and Constrained DecodingStructured outputs make a model produce data a program can parse, such as JSON matching a schema, and constrained decoding guarantees it by masking, at every step, the tokens that would break the format and renormalising over the rest.Know wellMust know
- System Prompts and InstructionsA system prompt is the developer's standing instruction at the start of the conversation (the role, rules, output format and available tools), and to the model it's just more tokens in the chat template that post-training taught it to give priority.Know wellMust know
- Text EmbeddingsA text embedding model turns a whole passage into one fixed-length vector, placed so that passages with similar meaning have a high cosine similarity, which lets a computer compare meaning with a dot product.Know wellMust know
- Tool CallingTool calling lets a model ask the application to run a function: the model emits a structured request (a tool name and JSON arguments), the application executes it and returns the result as a new message, and the model continues with that result in its context.Know wellMust know
- Contrastive LearningContrastive learning trains an encoder by pulling the vectors of matching pairs (a question and its answer) together and pushing non-matching ones apart, usually by asking the model to pick the true partner out of a batch with a softmax.UnderstandShould know
- GuardrailsGuardrails are the checks an application runs around a model (on inputs, on retrieved text and tool results, on outputs and on actions) because the model's own trained behaviour is a strong default, not a guarantee.UnderstandShould know
- RerankingReranking takes the top candidates from a fast retriever and rescores them with a slower, more accurate model, typically a cross-encoder that reads the query and each passage together.Know wellShould know
What do I actually need to remember?
- A model's weights are a frozen, unsourced snapshot; retrieval adds knowledge at question time and lets answers cite it.
- A text embedding model turns a passage into one vector; similar meaning gives a high cosine similarity.
- Embedding models are trained contrastively: pull matching pairs together, push the rest of the batch apart.
- Keyword search (BM25) wins on names, codes and rare terms; hybrid search fuses both rankings with reciprocal rank fusion.
- Approximate nearest-neighbour indexes (IVF, HNSW) search millions of vectors by scanning a small part and giving up a little recall.
- Retrieve cheaply, then rerank the top candidates with a cross-encoder that reads the question and passage together.
- RAG = retrieve passages, put them in the prompt, generate with citations; most RAG failures are retrieval failures.
- Chunk size trades completeness against precision; every retrieved token costs prefill and KV cache, and position in the prompt matters.
- Evaluate retrieval (recall@k, MRR) and answers (correctness, faithfulness, citations) separately, on many labelled questions.
- Tool calling: the model emits structured arguments and the application runs the tool; constrained decoding guarantees valid syntax, not correct content.
You do not need to memorize everything else. This list is the revision sheet.
Key papers
A Statistical Interpretation of Term Specificity and Its Application in Retrieval
Karen Spärck Jones · 1972 · Journal of Documentation
Introduced the idea behind inverse document frequency: a word that appears in few documents tells you more about them than a word that appears everywhere.
- Problem
- Matching documents by shared terms treats a common word like 'system' as just as informative as a rare one.
- What was new
- Weight each term by how specific it is, decreasing with the number of documents that contain it, which became the IDF in TF-IDF and BM25.
The Probabilistic Relevance Framework: BM25 and Beyond
Stephen Robertson, Hugo Zaragoza · 2009 · Foundations and Trends in Information Retrieval
The authoritative account of BM25, the keyword-ranking function that is still the standard baseline, and often a component, of modern retrieval systems.
- Problem
- Ranking by raw term counts over-rewards long documents and repeated words.
- What was new
- Derives BM25 from a probabilistic model of relevance: IDF times a term frequency that saturates, normalised for document length, with two tunable parameters k1 and b.
How to read it: Section 3 derives BM25; its discussion of parameters notes that 1.2 < k1 < 2 and 0.5 < b < 0.8 are reasonable in many settings.
Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods
Gordon V. Cormack, Charles L. A. Clarke, Stefan Buettcher · 2009 · SIGIR
A two-page paper whose one-line formula is how most hybrid search systems merge keyword and vector results.
- Problem
- Different retrieval systems produce scores on incomparable scales, so their results are hard to combine.
- What was new
- Ignore the scores and add 1 / (k + rank) across rankings; k = 60 was near-optimal in pilot experiments, though the choice was not critical.
Product Quantization for Nearest Neighbor Search
Hervé Jégou, Matthijs Douze, Cordelia Schmid · 2011 · IEEE TPAMI
Showed how to compress high-dimensional vectors to a few bytes each and still estimate distances, the basis of billion-scale vector search.
- Problem
- Storing and scanning millions of full-precision vectors costs too much memory and time.
- What was new
- Split each vector into sub-vectors, replace each by the nearest of a small learned codebook, and compute distances from precomputed tables.
Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs
Yu. A. Malkov, D. A. Yashunin · 2016
HNSW is the graph index behind most vector databases: fast, accurate approximate search over millions of embeddings.
- Problem
- Exact nearest-neighbour search compares the query with every stored vector, which is too slow at scale.
- What was new
- A layered proximity graph, like a skip list: each element joins a random number of layers, and a search descends greedily from the sparse top layer to the dense bottom one, with logarithmic complexity scaling.
How to read it: Section 4 (algorithm description) is the core; Figure 1 shows the layered search.
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, Hervé Jégou · 2017
The paper behind FAISS, the library that made exact and compressed vector search fast on GPUs and that many RAG systems still use.
- Problem
- GPU similarity search was bottlenecked by selecting the top k results and by poor use of the memory hierarchy.
- What was new
- A k-selection design running at up to 55% of theoretical peak, giving nearest-neighbour search 8.5× faster than the previous GPU state of the art, for brute-force and product-quantized indexes.
Reading Wikipedia to Answer Open-Domain Questions
Danqi Chen, Adam Fisch et al. · 2017
Set the retrieve-then-read template: a search component finds Wikipedia articles, and a neural reader extracts the answer from them.
- Problem
- Reading-comprehension models assumed the right paragraph was handed to them; real questions come with no paragraph.
- What was new
- DrQA combined bigram-hashing TF-IDF retrieval over all of Wikipedia with a recurrent reader trained to find answer spans.
Representation Learning with Contrastive Predictive Coding
Aaron van den Oord, Yazhe Li, Oriol Vinyals · 2018
Introduced the InfoNCE loss: pick the true partner out of a set of negatives with a softmax, the objective most embedding models are trained with.
- Problem
- Learning useful representations without labels needs a training signal that doesn't require predicting every detail of the input.
- What was new
- A probabilistic contrastive loss with negative sampling: score the true future sample against distractors and train with cross-entropy.
Passage Re-ranking with BERT
Rodrigo Nogueira, Kyunghyun Cho · 2019
Showed that a BERT model reading the query and passage together reorders a keyword search's results far better, making the retrieve-then-rerank pipeline standard.
- Problem
- Fast first-stage retrieval like BM25 finds candidates but orders them poorly.
- What was new
- Feed query and passage as one input to BERT and score relevance; on MS MARCO it beat the previous state of the art by 27% (relative) in MRR@10.
How to read it: Short: the method is one page, re-ranking the top 1,000 BM25 passages.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Nils Reimers, Iryna Gurevych · 2019
Made BERT produce one vector per sentence that can be compared with cosine similarity, which is what turned Transformers into practical search engines.
- Problem
- BERT compares two sentences by reading them together; finding the most similar pair among 10,000 sentences takes about 50 million passes, roughly 65 hours.
- What was new
- Fine-tune BERT in a siamese setup with mean pooling so each sentence is encoded once; the same search takes about 5 seconds.
How to read it: Section 3 (the model) explains the siamese network and pooling in a page.
Language Models as Knowledge Bases?
Fabio Petroni, Tim Rocktäschel et al. · 2019
Asked how much factual knowledge a pretrained model holds in its weights, framing the question that retrieval later answered differently.
- Problem
- Knowledge bases need hand-built schemas; pretrained language models might already store relational facts.
- What was new
- Probe models with fill-in-the-blank statements: without fine-tuning, BERT held relational knowledge competitive with some traditional methods, and some kinds of facts were learned far more readily than others.
REALM: Retrieval-Augmented Language Model Pre-Training
Kelvin Guu, Kenton Lee et al. · 2020
Trained a retriever and a language model together, so knowledge could live in a searchable corpus instead of only in the weights.
- Problem
- Knowledge stored implicitly in parameters requires ever-larger networks to cover more facts, and is hard to inspect.
- What was new
- Add a latent retriever over Wikipedia to masked-language-model pretraining and backpropagate through the retrieval step; it beat previous open-domain QA methods by 4–16% absolute accuracy.
Dense Passage Retrieval for Open-Domain Question Answering
Vladimir Karpukhin, Barlas Oğuz et al. · 2020
Showed that learned embeddings alone can beat BM25 for finding answer passages, which made dense retrieval the default first stage of RAG.
- Problem
- Open-domain QA retrieved passages with TF-IDF or BM25, which miss passages that use different words for the same thing.
- What was new
- Two BERT encoders, one for questions and one for passages, trained with in-batch negatives; it beat a strong BM25 system by 9–19% absolute in top-20 passage accuracy.
How to read it: Section 3 (the dual encoder and in-batch negatives) is the core; Section 4.1 describes splitting Wikipedia into 21 million 100-word passages.
ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
Omar Khattab, Matei Zaharia · 2020
Late interaction: keep one vector per token and match them cheaply, getting close to a cross-encoder's quality at a fraction of the cost.
- Problem
- BERT re-rankers must run the full network on every query–passage pair, orders of magnitude more expensive than earlier methods.
- What was new
- Encode query and passage separately into token vectors and score with a sum of maximum similarities; passages can be encoded offline. It was two orders of magnitude faster with four orders fewer FLOPs per query than BERT re-rankers, at competitive quality.
- Built on
- Passage Re-ranking with BERT
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Patrick Lewis, Ethan Perez et al. · 2020
Named retrieval-augmented generation: a generator conditioned on passages fetched from a dense index, with knowledge that can be inspected and updated without retraining.
- Problem
- Knowledge in a model's weights is hard to update, can't point to its sources, and lags task-specific systems on knowledge-intensive tasks.
- What was new
- Combine a pretrained seq2seq generator (BART) with a dense vector index of Wikipedia searched by a DPR retriever, trained end to end; it set the state of the art on three open-domain QA tasks.
How to read it: Section 2 (methods: the DPR retriever, the BART generator and the two ways of combining passages) is the core; the index is 21 million 100-word chunks of a December 2018 Wikipedia dump.
BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
Nandan Thakur, Nils Reimers et al. · 2021
Showed that retrievers which shine on their training domain often fall behind plain BM25 elsewhere, which is why hybrid search and re-ranking became standard.
- Problem
- Neural retrievers were evaluated in narrow, in-domain settings, hiding how they generalise.
- What was new
- 18 datasets across diverse tasks and domains, 10 systems compared zero-shot: BM25 was a robust baseline; re-ranking and late interaction did best on average, at high computational cost.
- Built on
- The Probabilistic Relevance Framework: BM25 and Beyond; Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks; ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
- Influenced
- MTEB: Massive Text Embedding Benchmark
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch et al. · 2021
A language model that looks things up in a 2-trillion-token database matched much larger models, evidence that retrieval can substitute for some parameters.
- Problem
- Scaling parameters is an expensive way to store more knowledge.
- What was new
- RETRO retrieves neighbouring chunks with a frozen BERT retriever and attends to them through chunked cross-attention; it matched GPT-3 and Jurassic-1 on the Pile with 25× fewer parameters.
MTEB: Massive Text Embedding Benchmark
Niklas Muennighoff, Nouamane Tazi et al. · 2022
The benchmark (and public leaderboard) people use to choose an embedding model.
- Problem
- Embedding models were evaluated on a few datasets from a single task, so it was unclear how well they transfer.
- What was new
- 8 tasks, 58 datasets, 112 languages and 33 models: no single embedding method dominated across all tasks.
Large Language Models Struggle to Learn Long-Tail Knowledge
Nikhil Kandpal, Haikang Deng et al. · 2022
Measured that a model's chance of answering a factual question tracks how many pretraining documents discuss it, which is why rare facts are where retrieval helps most.
- Problem
- It was unclear which facts a language model can learn from its pretraining data, and why it fails on others.
- What was new
- Entity-link pretraining corpora and count relevant documents per question: strong correlational and causal links between that count and accuracy, across models up to 176B parameters.
How to read it: Section 3 (correlational and causal analysis): BLOOM-176B's accuracy rises from about 25% to above 55% as relevant documents go from 10 to 10,000.
Text Embeddings by Weakly-Supervised Contrastive Pre-training
Liang Wang, Nan Yang et al. · 2022
E5 showed the modern recipe for general-purpose embedding models: contrastive pretraining on huge numbers of naturally occurring text pairs, then fine-tuning.
- Problem
- Good embedding models needed labelled relevance data, and zero-shot dense retrievers still lost to BM25.
- What was new
- Contrastive training on a curated web-scale dataset of text pairs; the first model to beat BM25 on the BEIR benchmark zero-shot without labelled data.
Precise Zero-Shot Dense Retrieval without Relevance Labels
Luyu Gao, Xueguang Ma et al. · 2022
HyDE: let a language model write a hypothetical answer and search with its embedding, since answers look more like documents than questions do.
- Problem
- Zero-shot dense retrieval is hard when no relevance labels exist to train on.
- What was new
- Generate a hypothetical document for the query (it may contain false details), embed it, and retrieve the real documents nearest to it.
When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories
Alex Mallen, Akari Asai et al. · 2022
Showed that retrieval matters most for less popular facts and that scaling barely helps there, suggesting retrieval only when it's needed.
- Problem
- When should a system trust a model's own memory, and when should it look things up?
- What was new
- PopQA, 14k questions spanning popular to obscure entities: models struggle with less popular facts, retrieval-augmented models beat far larger unaided ones, and retrieving only when needed improves results at lower cost.
In-Context Retrieval-Augmented Language Models
Ori Ram, Yoav Levine et al. · 2023
Showed that you don't need to change or retrain the model: just put retrieved documents in front of the input. That's how most RAG systems work today.
- Problem
- Earlier retrieval-augmented models changed the architecture, which complicated deployment.
- What was new
- Prepend retrieved documents to a frozen model's input; with off-the-shelf retrievers the gains were equivalent to a 2–3× larger model.
How to read it: The 'Our framework' section is two pages and describes the whole method.
Large Language Models Can Be Easily Distracted by Irrelevant Context
Freda Shi, Xinyun Chen et al. · 2023
A clean demonstration that adding irrelevant text to a prompt makes models worse, which is why retrieving more is not always better.
- Problem
- Benchmarks give models only relevant information; real contexts contain distractions.
- What was new
- GSM-IC, grade-school maths problems with an irrelevant sentence added: accuracy dropped sharply; self-consistency and an instruction to ignore irrelevant information helped.
Toolformer: Language Models Can Teach Themselves to Use Tools
Timo Schick, Jane Dwivedi-Yu et al. · 2023
Showed a model can learn when to call a calculator, search engine or calendar and how to use the result, the core idea behind tool calling.
- Problem
- Language models fail at things simple tools do well, like arithmetic and looking up facts.
- What was new
- Self-supervised: the model inserts candidate API calls into text, keeps the ones whose results make the following tokens easier to predict, and is fine-tuned on them. It improved zero-shot performance, often competitive with much larger models.
Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
Kai Greshake, Sahar Abdelnabi et al. · 2023
Showed that text a model retrieves (a web page, an email, a document) can carry instructions an attacker planted, so retrieval and tools are a security boundary.
- Problem
- Prompt injection was thought of as a user attacking their own chat; applications that read outside data blur data and instructions.
- What was new
- Indirect prompt injection: plant prompts in data likely to be retrieved. Demonstrated against real systems including Bing's GPT-4-powered chat, with a taxonomy of impacts such as data theft.
Enabling Large Language Models to Generate Text with Citations
Tianyu Gao, Howard Yen et al. · 2023
A benchmark with automatic metrics for whether a model's citations actually support what it says.
- Problem
- Citation quality was judged by hand with commercial search engines, hard to reproduce or compare.
- What was new
- ALCE: questions plus retrieval corpora, scored for fluency, correctness and citation quality; on ELI5 even the best models lacked complete citation support 50% of the time.
Lost in the Middle: How Language Models Use Long Contexts
Nelson F. Liu, Kevin Lin et al. · 2023
Showed that where a passage sits in the prompt changes whether the model uses it: best at the start or end, worst in the middle.
- Problem
- Models accepted long contexts, but little was known about how well they used them.
- What was new
- Move the answer-bearing document through 10, 20 or 30 retrieved documents: accuracy is U-shaped, and in the worst case GPT-3.5-Turbo did worse than with no documents at all.
How to read it: Section 2 (multi-document QA) and its Figure 5 hold the main result; Appendix G tabulates the numbers.
Efficient Guided Generation for Large Language Models
Brandon T. Willard, Rémi Louf · 2023
Made constrained decoding cheap: precompute which tokens are allowed in each state of a regular expression or grammar, so structured output can be guaranteed.
- Problem
- Masking invalid tokens naively means checking the whole vocabulary at every step.
- What was new
- Turn the pattern into a finite-state machine and index the vocabulary against its states once; at generation time the allowed tokens cost O(1) on average to look up.
How to read it: The section on iterative FSM processing and indexing, with Figure 1's floating-point-number example, is the core.
Ragas: Automated Evaluation of Retrieval Augmented Generation
Shahul Es, Jithin James et al. · 2023
A widely used framework for scoring RAG pipelines without human reference answers, using a language model as the judge.
- Problem
- Evaluating RAG needs judgments of retrieval relevance, faithfulness and answer quality, and human annotations are slow.
- What was new
- Reference-free metrics such as faithfulness, answer relevance and context relevance, each computed by prompting a language model.
Retrieval meets Long Context Large Language Models
Peng Xu, Wei Ping et al. · 2023
A direct comparison of the two ways to give a model more information: retrieval or a longer context window. They found retrieval helps even when the window is long.
- Problem
- Should you extend the context window or retrieve, and can the two be combined?
- What was new
- A 4K-context model with simple retrieval matched a 16K-context fine-tuned model on long-context tasks with much less computation, and retrieval improved models regardless of window size.
LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
Huiqiang Jiang, Qianhui Wu et al. · 2023
Prompt compression: drop the tokens a small model finds predictable before sending a long prompt to a large one.
- Problem
- Prompts with retrieved documents and examples grow to tens of thousands of tokens, which is slow and expensive.
- What was new
- Coarse-to-fine compression with a budget controller and token-level iterative pruning; up to 20× compression with little performance loss on the tasks tested.
Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
Akari Asai, Zeqiu Wu et al. · 2023
Trains the model itself to decide when to retrieve and to grade whether passages and its own claims are supported.
- Problem
- Always retrieving a fixed number of passages, relevant or not, can make answers worse.
- What was new
- Special reflection tokens let the model retrieve on demand and critique retrieved passages and its own output; it improved factuality and citation accuracy over the baselines tested.
RULER: What's the Real Context Size of Your Long-Context Language Models?
Cheng-Ping Hsieh, Simeng Sun et al. · 2024
Showed that a model's advertised context length overstates the length it can actually use well.
- Problem
- Needle-in-a-haystack tests only check simple retrieval from a long context.
- What was new
- Synthetic tasks with multiple needles, multi-hop tracing and aggregation: of 17 models claiming 32K tokens or more, only half kept satisfactory performance at 32K.
From Local to Global: A Graph RAG Approach to Query-Focused Summarization
Darren Edge, Ha Trinh et al. · 2024
Addresses a blind spot of chunk retrieval: questions about a whole collection, like 'what are the main themes?'.
- Problem
- RAG retrieves a few passages, so it fails on global questions that need the entire corpus.
- What was new
- Use an LLM to build an entity graph, summarise communities of related entities in advance, then answer from those summaries; it improved comprehensiveness and diversity on such questions over a conventional RAG baseline.
Watch
Stanford Online
Stanford CS25: V3 I Retrieval Augmented Language Models
Douwe Kiela, one of the authors of the original RAG paper, surveys retrieval-augmented language models and ends with the open questions.
Covers: Why language models need retrieval, the main families of retrieval-augmented models from the recent literature, and the questions that remain open.
3Blue1Brown
How might LLMs store facts | Deep Learning Chapter 7
The feed-forward half of a Transformer block gets less attention than attention; this fixes that.
Covers: MLP sublayers, directions in embedding space, superposition (intuition).
Andrej Karpathy
[1hr Talk] Intro to Large Language Models
A clear one-hour overview of what LLMs are, how they are trained, and where they're going — good orientation for Part III.
Covers: Pretraining, fine-tuning, scaling, tool use, security issues.
Andrej Karpathy
Deep Dive into LLMs like ChatGPT
A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.
Covers: The full training stack behind a chat model, from internet text and tokens to a pretrained base model and the post-training that turns it into an assistant.