Skip to content
Road to Intelligence

Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack

Context Engineering

Must knowKnow well16 minDifficulty

Context engineering is deciding what goes into the model's limited context window, in what form and in what order (instructions, retrieved passages, conversation history, tool results), because the model can only use what it reads, and doesn't use everything it reads equally well.

The problem

The context window is a fixed budget; every extra token costs prefill compute and cache memory, and models use information unevenly: relevant facts buried among distractors or in the middle of a long context are often missed.

The solution

Treat the prompt as a built artefact: put stable instructions first, include only the most relevant passages with IDs and sources, place the most important material where it is used best, summarise or drop old history, and measure the effect of each choice.

The consequence

Answer quality, cost and latency depend on prompt construction as much as on the model. A longer context window is not a substitute for selecting well.

The window is a budget

Everything the model sees for a request competes for one window: system instructions, tool descriptions, the conversation so far, retrieved passages, the question. Each token is processed in prefill and stored in the KV cache for the rest of the reply. Retrieving 10 passages of about 150 tokens adds 1,500 tokens; for Llama 3 8B, whose cache takes 128 KiB per token in 16 bits, that's about 200 MB of cache for that one request (our arithmetic from Chapter 11's formula).

Position matters

Liu and colleagues moved the document containing the answer through a list of 20 retrieved documents: GPT-3.5-Turbo answered 75.8% correctly with it first, 53.8% with it in the middle (position 10) and 63.2% with it last, against 56.1% with no documents at all Established. In the middle, the model did worse than closed-book. The paper reports that performance is often highest when relevant information is at the beginning or end of the input, and that extended-context versions of models were not necessarily better at using their context Established.

What to do with that: rank passages well, keep the list short, and put the best ones where they are used best. Reversing the order so the best passage sits last, next to the question, is a common trick; test it on your model.

Less can be more

Shi and colleagues added a single irrelevant sentence to grade-school maths problems and found model accuracy dropped dramatically; telling the model to ignore irrelevant information helped Established. Retrieving 20 passages "to be safe" can hurt. Rerank, then keep the few that matter.

Long context or retrieval?

Context windows have grown into the hundreds of thousands of tokens, so why not paste everything? Cost and latency grow with length, and quality doesn't keep up. Xu and colleagues found a 4K-context model with simple retrieval comparable to a 16K-context model on long-context tasks with much less computation, and that retrieval helped regardless of window size Established. The RULER benchmark found that only half of 17 models claiming 32K tokens or more kept satisfactory performance at 32K Established. Long windows and retrieval are complements: retrieval picks what matters, a long window lets you include more of it when needed Interpretation.

History and compression

Conversations grow. Options: keep the last few turns verbatim, summarise older ones into a short running note, or treat past turns as another collection to retrieve from. Agents (Chapter 13) need the same for their tool results. Prompt compression methods such as LLMLingua drop tokens that a small model finds predictable and report up to 20× compression with little performance loss on the tasks they tested Established.

A template that works

  1. Stable instructions first (and they can be cached between requests).
  2. Retrieved passages, each with an ID, title and source, best placed deliberately.
  3. Recent conversation, summarised beyond a budget.
  4. The question last, with the reminder to answer only from the sources and cite them.

Why should I care?

As a researcher

How models use long contexts (position effects, distraction, effective versus advertised length) is an active research area with clear measurements and few settled explanations.

As an engineer

The prompt is the interface you control most directly; its size drives cost and latency, and its order and selection drive accuracy.

Modern systems that depend on it

  • RAG answer quality
  • long conversations
  • agents' working memory
  • prompt caching

Historical context

Before

Prompts were hand-written strings with a few examples; the main question was wording.

After

Prompts are assembled per request from instructions, retrieved passages, history and tool outputs under a token budget, and the assembly logic is versioned and tested like code.

Used today

Every chat product trims or summarises history, every RAG system chooses how many passages to include and how to label them, and agent frameworks manage what each step can see.

What to remember

  • The window is a budget: tokens cost prefill compute and KV-cache memory (Chapter 11).
  • Lost in the middle: relevant information at the start or end is used best.
  • Irrelevant passages can make answers worse, not just longer.
  • Advertised context length ≠ length used well.
  • Long history: summarise, retrieve from it, or drop it.

Key papers

Essential

Lost in the Middle: How Language Models Use Long Contexts

Nelson F. Liu, Kevin Lin et al. · 2023

Showed that where a passage sits in the prompt changes whether the model uses it: best at the start or end, worst in the middle.

How to read it: Section 2 (multi-document QA) and its Figure 5 hold the main result; Appendix G tabulates the numbers.

~30 min readarXiv:2307.03172✓ verified 2026-10-05
Optional

Retrieval meets Long Context Large Language Models

Peng Xu, Wei Ping et al. · 2023

A direct comparison of the two ways to give a model more information: retrieval or a longer context window. They found retrieval helps even when the window is long.

~30 min readarXiv:2310.03025✓ verified 2026-10-05