Skip to content
Road to Intelligence

Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack

Guardrails

Should knowUnderstand10 minDifficulty

Guardrails are the checks an application runs around a model (on inputs, on retrieved text and tool results, on outputs and on actions) because the model's own trained behaviour is a strong default, not a guarantee.

The problem

A model can be asked, tricked or nudged into off-topic, unsafe or wrong outputs, and text it reads from documents or tools can contain instructions planted by someone else.

The solution

Wrap the model in ordinary software: validate and classify inputs, mark retrieved and tool content as data, check outputs against schemas, citations and safety classifiers, and require permission or confirmation before actions.

The consequence

Reliability and safety become properties of the whole system, testable like other code. No filter is complete, so the most important control is limiting what the model is able to do.

Why code around the model

Safety tuning and a good system prompt shape what a model usually does. Applications need some things always: never show another customer's data, never send money without confirmation, always return parseable output. "Always" belongs in code.

The checkpoints

  • Input: is this request in scope? Does it contain personal data that shouldn't be logged? A small classifier or rules can route or refuse.
  • Retrieved text and tool results: these are data. Delimit them clearly, record their source, and never let them change what the model is allowed to do.
  • Output: does it match the schema (structured outputs)? Does every cited passage exist and support the claim? Does a safety classifier flag it?
  • Actions: is this user allowed to do this? Is it reversible? Ask for confirmation before anything that isn't.

Indirect prompt injection

Greshake and colleagues showed that attackers can plant instructions in data likely to be retrieved, such as web pages, and that LLM-integrated applications, including Bing's GPT-4-powered chat, would follow them, enabling attacks such as data theft Established. In a RAG system, any indexed document is a possible attack surface; with tools, so is every tool result. There is no complete defence yet; delimiting untrusted content, training models to ignore instructions in data, and filtering help, but don't guarantee it Active research. What does work reliably is limiting damage: least-privilege tools, no secrets in the context that the model doesn't need, and human confirmation for consequential actions.

What to remember

  • Trained refusals are a default; guardrails are code around the model.
  • Check inputs, retrieved content, outputs and actions.
  • Indirect prompt injection: instructions hidden in data the model reads.
  • Least privilege: a tool the model can't call can't be misused.
  • Chapter 16 covers attacks and defences in depth.

Key papers