Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack
Retrieval-Augmented Generation (RAG)
Retrieval-augmented generation answers a question by first searching a document collection for relevant passages and then giving those passages to a language model in its prompt, so the answer can use, and cite, knowledge that isn't in the model's weights.
The problem
A model's knowledge is frozen at its training cutoff, misses private and rare information, and can't point to where an answer came from.
The solution
Index a document collection offline (chunk, embed, store). At question time retrieve the most relevant chunks, put them in the prompt with the question and instructions to answer from them and cite them, and generate.
The consequence
Knowledge can be updated by editing documents instead of retraining, answers can carry citations, and small models can answer questions about large private collections. New failure points appear: the right passage may not be retrieved, and the model may ignore or misuse what was.
You should understand first
- Vectors
- Dot Product
- Embeddings
- Attention
- Probability and Distributions
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Text Embeddings
- Text as Data
- Keyword Search and BM25
- Semantic and Hybrid Search
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- Parametric vs Retrieved Knowledge
- Retrieval-Augmented Generation (RAG)
Two pipelines
A RAG system is two data pipelines that meet at the prompt.
Offline (indexing), an ETL job: load documents → split into chunks → embed each chunk → store vectors with their text and metadata in an index. Rerun when documents change.
Online (per question): embed the question → retrieve the top candidates (semantic, keyword or hybrid) → optionally rerank → assemble a prompt with instructions, numbered passages and the question (context engineering) → generate → check and show citations.
Where it came from
DrQA (2017) answered open-domain questions over Wikipedia by combining TF-IDF retrieval with a neural reader that extracted the answer span Established. In 2020, REALM added a learned retriever to language-model pretraining and backpropagated through the retrieval step Established, and Lewis and colleagues introduced RAG models combining a pretrained seq2seq generator with a dense vector index of Wikipedia searched by a neural retriever, setting the state of the art on three open-domain QA tasks Established. Their index was 21 million 100-word chunks of Wikipedia. RETRO (2021) retrieved from a 2-trillion-token database and matched GPT-3 and Jurassic-1 on the Pile with 25× fewer parameters Established.
Those systems trained the retriever and the model to work together. Most deployed systems today don't. Ram and colleagues showed that simply prepending retrieved documents to a frozen language model's input, with an off-the-shelf retriever, gave gains equivalent to a 2–3× larger model Established. The prepend-to-a-frozen-model pattern is the common practice in today's applications Interpretation, because it works with any model, including one behind an API.
A tiny worked prompt
Answer using only the sources below. Cite them like [1]. If they don't contain the answer, say so.
[1] The KV Cache: ... OPT-13B's cache as 2 (key and value) × 5,120 (hidden size) × 40 (layers) × 2 bytes = 800 KB per token ...
[2] Paged attention: ... only 20.4–38.2% of KV-cache memory held actual token states ...
Question: How much cache memory does OPT-13B need per token?
A good answer is "800 KB per token [1]." The model didn't need to remember OPT-13B's dimensions; it needed to read.
Where it fails
- Retrieval misses: the right passage isn't in the top k (wrong chunking, a paraphrase the retriever didn't catch, an identifier embeddings blurred).
- The model ignores or misreads the passage, especially among many distractors or in the middle of a long context.
- Stale or conflicting sources: the index answers with what it contains.
- Questions about the whole collection ("what are the main themes?") aren't answered by a few chunks; graph-based approaches such as GraphRAG pre-summarise the corpus for such questions Active research.
- Injected instructions: a retrieved page can contain text addressed to the model (guardrails).
Why should I care?
As a researcher
RAG separates what a model knows from what it can look up, which raises questions about how models combine the two, when retrieval should be used at all, and how to attribute answers.
As an engineer
Most production LLM features that answer from company documents, tickets or code are RAG pipelines, and most of their failures are retrieval failures.
Modern systems that depend on it
- enterprise search assistants
- coding assistants with repository context
- agents with search tools
- citation and attribution
Historical context
Before
Facts had to be learned into the weights during training; updating them meant fine-tuning, and answers had no sources.
After
Facts live in an index you control; the model reads the relevant pieces at question time and can cite them.
Used today
Search-backed chat assistants, documentation and support bots, and coding tools that pull in relevant files all follow the retrieve-then-generate pattern, usually with a frozen model and an off-the-shelf retriever.
What to remember
- Offline: load → chunk → embed → index. Online: embed query → retrieve → (rerank) → build prompt → generate → cite.
- The model is usually unchanged: retrieved passages are just prompt text (in-context RALM).
- Update knowledge by updating documents, not weights.
- Retrieval quality caps answer quality: a missed passage can't be cited.
- Retrieved text is untrusted input; treat it as data, not instructions.
Key papers
Reading Wikipedia to Answer Open-Domain Questions
Danqi Chen, Adam Fisch et al. · 2017
Set the retrieve-then-read template: a search component finds Wikipedia articles, and a neural reader extracts the answer from them.
REALM: Retrieval-Augmented Language Model Pre-Training
Kelvin Guu, Kenton Lee et al. · 2020
Trained a retriever and a language model together, so knowledge could live in a searchable corpus instead of only in the weights.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Patrick Lewis, Ethan Perez et al. · 2020
Named retrieval-augmented generation: a generator conditioned on passages fetched from a dense index, with knowledge that can be inspected and updated without retraining.
How to read it: Section 2 (methods: the DPR retriever, the BART generator and the two ways of combining passages) is the core; the index is 21 million 100-word chunks of a December 2018 Wikipedia dump.
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch et al. · 2021
A language model that looks things up in a 2-trillion-token database matched much larger models, evidence that retrieval can substitute for some parameters.
In-Context Retrieval-Augmented Language Models
Ori Ram, Yoav Levine et al. · 2023
Showed that you don't need to change or retrain the model: just put retrieved documents in front of the input. That's how most RAG systems work today.
How to read it: The 'Our framework' section is two pages and describes the whole method.
Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
Akari Asai, Zeqiu Wu et al. · 2023
Trains the model itself to decide when to retrieve and to grade whether passages and its own claims are supported.
From Local to Global: A Graph RAG Approach to Query-Focused Summarization
Darren Edge, Ha Trinh et al. · 2024
Addresses a blind spot of chunk retrieval: questions about a whole collection, like 'what are the main themes?'.
Watch
Stanford Online
Stanford CS25: V3 I Retrieval Augmented Language Models
Douwe Kiela, one of the authors of the original RAG paper, surveys retrieval-augmented language models and ends with the open questions.