Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack
Parametric vs Retrieved Knowledge
What a language model knows from training is stored in its weights (parametric knowledge): frozen at the training cutoff, unreliable for rare facts and unable to name its source; retrieved knowledge is read from documents at question time instead.
The problem
Users ask about recent events, private documents and obscure facts, and expect answers they can check; a model's weights contain none of the first two and only patchy coverage of the third.
The solution
Keep facts that change, are private, or are rare in a document collection the system can search, and give the relevant pieces to the model in its prompt.
The consequence
A system's knowledge can be updated without retraining and answers can be traced to sources; the model's own knowledge still matters for understanding the question and the passages.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- Feed-Forward Sublayer (MLP)
- Parametric vs Retrieved Knowledge
Knowledge in the weights
Pretraining compresses a huge amount of text into parameters, and some of it comes back as facts: ask a model for the capital of France and it answers without looking anything up. Petroni and colleagues found that BERT, without fine-tuning, held relational knowledge competitive with some traditional knowledge-base methods, and that some kinds of facts were learned much more readily than others Established. 3Blue1Brown's video on how LLMs might store facts walks through one proposed mechanism, the feed-forward layers acting like key–value lookups Interpretation.
This parametric knowledge has three limits:
- A cutoff. It stops at the training data's date; anything later is unknown or guessed.
- No private data. Your company's wiki was not in the pretraining set.
- A long tail. Kandpal and colleagues found that a model's accuracy on a factual question depends strongly on how many pretraining documents mention the question's entities; BLOOM-176B's accuracy rose from about 25% to above 55% as that count went from 10 to 10,000 Established. Mallen and colleagues found that scaling barely improved memorisation of less popular facts, while retrieval-augmented models beat far larger unaided ones on them Established.
And a fourth, practical one: weights can't tell you where a fact came from.
Three ways to add knowledge
| Fine-tune | Long context | Retrieve | |
|---|---|---|---|
| Update when facts change | retrain | edit the documents | edit the index |
| Cite sources | no | possible | yes |
| Cost per question | none extra | high (all tokens, every time) | moderate (a few passages) |
| Good for | behaviour, style, format | one long document | large, changing collections |
Fine-tuning changes behaviour well, but is an unreliable way to insert new facts. Putting everything in the context works for one long document but not for millions. Retrieval is the general answer, and the subject of this chapter.
What to remember
- Parametric: learned into the weights during training. Frozen, unsourced.
- Retrieved (non-parametric): read from documents at question time. Updatable, citable.
- Models answer popular facts well and rare ones poorly; retrieval helps most on the long tail.
- Fine-tuning is a poor way to add facts; it's better for behaviour and format.
Key papers
Language Models as Knowledge Bases?
Fabio Petroni, Tim Rocktäschel et al. · 2019
Asked how much factual knowledge a pretrained model holds in its weights, framing the question that retrieval later answered differently.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Patrick Lewis, Ethan Perez et al. · 2020
Named retrieval-augmented generation: a generator conditioned on passages fetched from a dense index, with knowledge that can be inspected and updated without retraining.
How to read it: Section 2 (methods: the DPR retriever, the BART generator and the two ways of combining passages) is the core; the index is 21 million 100-word chunks of a December 2018 Wikipedia dump.
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch et al. · 2021
A language model that looks things up in a 2-trillion-token database matched much larger models, evidence that retrieval can substitute for some parameters.
Large Language Models Struggle to Learn Long-Tail Knowledge
Nikhil Kandpal, Haikang Deng et al. · 2022
Measured that a model's chance of answering a factual question tracks how many pretraining documents discuss it, which is why rare facts are where retrieval helps most.
How to read it: Section 3 (correlational and causal analysis): BLOOM-176B's accuracy rises from about 25% to above 55% as relevant documents go from 10 to 10,000.
When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories
Alex Mallen, Akari Asai et al. · 2022
Showed that retrieval matters most for less popular facts and that scaling barely helps there, suggesting retrieval only when it's needed.
Watch
3Blue1Brown
How might LLMs store facts | Deep Learning Chapter 7
The feed-forward half of a Transformer block gets less attention than attention; this fixes that.
Stanford Online
Stanford CS25: V3 I Retrieval Augmented Language Models
Douwe Kiela, one of the authors of the original RAG paper, surveys retrieval-augmented language models and ends with the open questions.