Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack
Guardrails
Guardrails are the checks an application runs around a model (on inputs, on retrieved text and tool results, on outputs and on actions) because the model's own trained behaviour is a strong default, not a guarantee.
The problem
A model can be asked, tricked or nudged into off-topic, unsafe or wrong outputs, and text it reads from documents or tools can contain instructions planted by someone else.
The solution
Wrap the model in ordinary software: validate and classify inputs, mark retrieved and tool content as data, check outputs against schemas, citations and safety classifiers, and require permission or confirmation before actions.
The consequence
Reliability and safety become properties of the whole system, testable like other code. No filter is complete, so the most important control is limiting what the model is able to do.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Decoding: Greedy, Temperature, Top-k, Top-p
- One-Hot Encoding
- Tokenization
- Chat Templates
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- System Prompts and Instructions
- Structured Outputs and Constrained Decoding
- Tool Calling
- Multi-Head Attention
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Pretrain, Then Fine-Tune
- Supervised Fine-Tuning
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Learning from Comparisons
- Safety Tuning and Refusal
- Guardrails
Why code around the model
Safety tuning and a good system prompt shape what a model usually does. Applications need some things always: never show another customer's data, never send money without confirmation, always return parseable output. "Always" belongs in code.
The checkpoints
- Input: is this request in scope? Does it contain personal data that shouldn't be logged? A small classifier or rules can route or refuse.
- Retrieved text and tool results: these are data. Delimit them clearly, record their source, and never let them change what the model is allowed to do.
- Output: does it match the schema (structured outputs)? Does every cited passage exist and support the claim? Does a safety classifier flag it?
- Actions: is this user allowed to do this? Is it reversible? Ask for confirmation before anything that isn't.
Indirect prompt injection
Greshake and colleagues showed that attackers can plant instructions in data likely to be retrieved, such as web pages, and that LLM-integrated applications, including Bing's GPT-4-powered chat, would follow them, enabling attacks such as data theft Established. In a RAG system, any indexed document is a possible attack surface; with tools, so is every tool result. There is no complete defence yet; delimiting untrusted content, training models to ignore instructions in data, and filtering help, but don't guarantee it Active research. What does work reliably is limiting damage: least-privilege tools, no secrets in the context that the model doesn't need, and human confirmation for consequential actions.
What to remember
- Trained refusals are a default; guardrails are code around the model.
- Check inputs, retrieved content, outputs and actions.
- Indirect prompt injection: instructions hidden in data the model reads.
- Least privilege: a tool the model can't call can't be misused.
- Chapter 16 covers attacks and defences in depth.
Key papers
Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
Kai Greshake, Sahar Abdelnabi et al. · 2023
Showed that text a model retrieves (a web page, an email, a document) can carry instructions an attacker planted, so retrieval and tools are a security boundary.