Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack
Context Engineering
Context engineering is deciding what goes into the model's limited context window, in what form and in what order (instructions, retrieved passages, conversation history, tool results), because the model can only use what it reads, and doesn't use everything it reads equally well.
The problem
The context window is a fixed budget; every extra token costs prefill compute and cache memory, and models use information unevenly: relevant facts buried among distractors or in the middle of a long context are often missed.
The solution
Treat the prompt as a built artefact: put stable instructions first, include only the most relevant passages with IDs and sources, place the most important material where it is used best, summarise or drop old history, and measure the effect of each choice.
The consequence
Answer quality, cost and latency depend on prompt construction as much as on the model. A longer context window is not a substitute for selecting well.
You should understand first
- Vectors
- Dot Product
- Embeddings
- Attention
- Probability and Distributions
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Text Embeddings
- Text as Data
- Keyword Search and BM25
- Semantic and Hybrid Search
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- Parametric vs Retrieved Knowledge
- Retrieval-Augmented Generation (RAG)
- Matrix Multiplication
- Prefill, Decode and the Memory Wall
- The KV Cache
- Context Engineering
The window is a budget
Everything the model sees for a request competes for one window: system instructions, tool descriptions, the conversation so far, retrieved passages, the question. Each token is processed in prefill and stored in the KV cache for the rest of the reply. Retrieving 10 passages of about 150 tokens adds 1,500 tokens; for Llama 3 8B, whose cache takes 128 KiB per token in 16 bits, that's about 200 MB of cache for that one request (our arithmetic from Chapter 11's formula).
Position matters
Liu and colleagues moved the document containing the answer through a list of 20 retrieved documents: GPT-3.5-Turbo answered 75.8% correctly with it first, 53.8% with it in the middle (position 10) and 63.2% with it last, against 56.1% with no documents at all Established. In the middle, the model did worse than closed-book. The paper reports that performance is often highest when relevant information is at the beginning or end of the input, and that extended-context versions of models were not necessarily better at using their context Established.
What to do with that: rank passages well, keep the list short, and put the best ones where they are used best. Reversing the order so the best passage sits last, next to the question, is a common trick; test it on your model.
Less can be more
Shi and colleagues added a single irrelevant sentence to grade-school maths problems and found model accuracy dropped dramatically; telling the model to ignore irrelevant information helped Established. Retrieving 20 passages "to be safe" can hurt. Rerank, then keep the few that matter.
Long context or retrieval?
Context windows have grown into the hundreds of thousands of tokens, so why not paste everything? Cost and latency grow with length, and quality doesn't keep up. Xu and colleagues found a 4K-context model with simple retrieval comparable to a 16K-context model on long-context tasks with much less computation, and that retrieval helped regardless of window size Established. The RULER benchmark found that only half of 17 models claiming 32K tokens or more kept satisfactory performance at 32K Established. Long windows and retrieval are complements: retrieval picks what matters, a long window lets you include more of it when needed Interpretation.
History and compression
Conversations grow. Options: keep the last few turns verbatim, summarise older ones into a short running note, or treat past turns as another collection to retrieve from. Agents (Chapter 13) need the same for their tool results. Prompt compression methods such as LLMLingua drop tokens that a small model finds predictable and report up to 20× compression with little performance loss on the tasks they tested Established.
A template that works
- Stable instructions first (and they can be cached between requests).
- Retrieved passages, each with an ID, title and source, best placed deliberately.
- Recent conversation, summarised beyond a budget.
- The question last, with the reminder to answer only from the sources and cite them.
Why should I care?
As a researcher
How models use long contexts (position effects, distraction, effective versus advertised length) is an active research area with clear measurements and few settled explanations.
As an engineer
The prompt is the interface you control most directly; its size drives cost and latency, and its order and selection drive accuracy.
Modern systems that depend on it
- RAG answer quality
- long conversations
- agents' working memory
- prompt caching
Historical context
Before
Prompts were hand-written strings with a few examples; the main question was wording.
After
Prompts are assembled per request from instructions, retrieved passages, history and tool outputs under a token budget, and the assembly logic is versioned and tested like code.
Used today
Every chat product trims or summarises history, every RAG system chooses how many passages to include and how to label them, and agent frameworks manage what each step can see.
What to remember
- The window is a budget: tokens cost prefill compute and KV-cache memory (Chapter 11).
- Lost in the middle: relevant information at the start or end is used best.
- Irrelevant passages can make answers worse, not just longer.
- Advertised context length ≠ length used well.
- Long history: summarise, retrieve from it, or drop it.
Key papers
Large Language Models Can Be Easily Distracted by Irrelevant Context
Freda Shi, Xinyun Chen et al. · 2023
A clean demonstration that adding irrelevant text to a prompt makes models worse, which is why retrieving more is not always better.
Lost in the Middle: How Language Models Use Long Contexts
Nelson F. Liu, Kevin Lin et al. · 2023
Showed that where a passage sits in the prompt changes whether the model uses it: best at the start or end, worst in the middle.
How to read it: Section 2 (multi-document QA) and its Figure 5 hold the main result; Appendix G tabulates the numbers.
Retrieval meets Long Context Large Language Models
Peng Xu, Wei Ping et al. · 2023
A direct comparison of the two ways to give a model more information: retrieval or a longer context window. They found retrieval helps even when the window is long.
LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
Huiqiang Jiang, Qianhui Wu et al. · 2023
Prompt compression: drop the tokens a small model finds predictable before sending a long prompt to a large one.
RULER: What's the Real Context Size of Your Long-Context Language Models?
Cheng-Ping Hsieh, Simeng Sun et al. · 2024
Showed that a model's advertised context length overstates the length it can actually use well.