Skip to content
Road to Intelligence

Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack

Retrieval-Augmented Generation (RAG)

Must knowKnow well18 minDifficulty

Retrieval-augmented generation answers a question by first searching a document collection for relevant passages and then giving those passages to a language model in its prompt, so the answer can use, and cite, knowledge that isn't in the model's weights.

The problem

A model's knowledge is frozen at its training cutoff, misses private and rare information, and can't point to where an answer came from.

The solution

Index a document collection offline (chunk, embed, store). At question time retrieve the most relevant chunks, put them in the prompt with the question and instructions to answer from them and cite them, and generate.

The consequence

Knowledge can be updated by editing documents instead of retraining, answers can carry citations, and small models can answer questions about large private collections. New failure points appear: the right passage may not be retrieved, and the model may ignore or misuse what was.

Two pipelines

A RAG system is two data pipelines that meet at the prompt.

Offline (indexing), an ETL job: load documents → split into chunks → embed each chunk → store vectors with their text and metadata in an index. Rerun when documents change.

Online (per question): embed the question → retrieve the top candidates (semantic, keyword or hybrid) → optionally rerank → assemble a prompt with instructions, numbered passages and the question (context engineering) → generate → check and show citations.

Where it came from

DrQA (2017) answered open-domain questions over Wikipedia by combining TF-IDF retrieval with a neural reader that extracted the answer span Established. In 2020, REALM added a learned retriever to language-model pretraining and backpropagated through the retrieval step Established, and Lewis and colleagues introduced RAG models combining a pretrained seq2seq generator with a dense vector index of Wikipedia searched by a neural retriever, setting the state of the art on three open-domain QA tasks Established. Their index was 21 million 100-word chunks of Wikipedia. RETRO (2021) retrieved from a 2-trillion-token database and matched GPT-3 and Jurassic-1 on the Pile with 25× fewer parameters Established.

Those systems trained the retriever and the model to work together. Most deployed systems today don't. Ram and colleagues showed that simply prepending retrieved documents to a frozen language model's input, with an off-the-shelf retriever, gave gains equivalent to a 2–3× larger model Established. The prepend-to-a-frozen-model pattern is the common practice in today's applications Interpretation, because it works with any model, including one behind an API.

A tiny worked prompt

Answer using only the sources below. Cite them like [1]. If they don't contain the answer, say so.

[1] The KV Cache: ... OPT-13B's cache as 2 (key and value) × 5,120 (hidden size) × 40 (layers) × 2 bytes = 800 KB per token ...
[2] Paged attention: ... only 20.4–38.2% of KV-cache memory held actual token states ...

Question: How much cache memory does OPT-13B need per token?

A good answer is "800 KB per token [1]." The model didn't need to remember OPT-13B's dimensions; it needed to read.

Where it fails

  • Retrieval misses: the right passage isn't in the top k (wrong chunking, a paraphrase the retriever didn't catch, an identifier embeddings blurred).
  • The model ignores or misreads the passage, especially among many distractors or in the middle of a long context.
  • Stale or conflicting sources: the index answers with what it contains.
  • Questions about the whole collection ("what are the main themes?") aren't answered by a few chunks; graph-based approaches such as GraphRAG pre-summarise the corpus for such questions Active research.
  • Injected instructions: a retrieved page can contain text addressed to the model (guardrails).

Why should I care?

As a researcher

RAG separates what a model knows from what it can look up, which raises questions about how models combine the two, when retrieval should be used at all, and how to attribute answers.

As an engineer

Most production LLM features that answer from company documents, tickets or code are RAG pipelines, and most of their failures are retrieval failures.

Modern systems that depend on it

  • enterprise search assistants
  • coding assistants with repository context
  • agents with search tools
  • citation and attribution

Historical context

Before

Facts had to be learned into the weights during training; updating them meant fine-tuning, and answers had no sources.

After

Facts live in an index you control; the model reads the relevant pieces at question time and can cite them.

Used today

Search-backed chat assistants, documentation and support bots, and coding tools that pull in relevant files all follow the retrieve-then-generate pattern, usually with a frozen model and an off-the-shelf retriever.

What to remember

  • Offline: load → chunk → embed → index. Online: embed query → retrieve → (rerank) → build prompt → generate → cite.
  • The model is usually unchanged: retrieved passages are just prompt text (in-context RALM).
  • Update knowledge by updating documents, not weights.
  • Retrieval quality caps answer quality: a missed passage can't be cited.
  • Retrieved text is untrusted input; treat it as data, not instructions.

Key papers

Important

Reading Wikipedia to Answer Open-Domain Questions

Danqi Chen, Adam Fisch et al. · 2017

Set the retrieve-then-read template: a search component finds Wikipedia articles, and a neural reader extracts the answer from them.

~30 min readarXiv:1704.00051✓ verified 2026-10-05
Important

REALM: Retrieval-Augmented Language Model Pre-Training

Kelvin Guu, Kenton Lee et al. · 2020

Trained a retriever and a language model together, so knowledge could live in a searchable corpus instead of only in the weights.

~45 min readarXiv:2002.08909✓ verified 2026-10-05
Essential

Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Patrick Lewis, Ethan Perez et al. · 2020

Named retrieval-augmented generation: a generator conditioned on passages fetched from a dense index, with knowledge that can be inspected and updated without retraining.

How to read it: Section 2 (methods: the DPR retriever, the BART generator and the two ways of combining passages) is the core; the index is 21 million 100-word chunks of a December 2018 Wikipedia dump.

~45 min readarXiv:2005.11401✓ verified 2026-10-05
Important

Improving language models by retrieving from trillions of tokens

Sebastian Borgeaud, Arthur Mensch et al. · 2021

A language model that looks things up in a 2-trillion-token database matched much larger models, evidence that retrieval can substitute for some parameters.

~1 h readarXiv:2112.04426✓ verified 2026-10-05
Important

In-Context Retrieval-Augmented Language Models

Ori Ram, Yoav Levine et al. · 2023

Showed that you don't need to change or retrain the model: just put retrieved documents in front of the input. That's how most RAG systems work today.

How to read it: The 'Our framework' section is two pages and describes the whole method.

~35 min readarXiv:2302.00083✓ verified 2026-10-05

Watch