Concept · Chapter 13: Agents
Agent Trust Boundaries
Agent trust boundaries prevent untrusted observations from granting permissions or directing unauthorized tool use.
The problem
An external page, issue or tool result can embed instructions that conflict with the real task.
The solution
Treat observations as data and enforce tool scope and authorization outside the model.
The consequence
A relevant document is not a trusted source of authority.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Expected Value and Variance
- Reinforcement Learning
- MDPs, Policies and Value
- Decoding: Greedy, Temperature, Top-k, Top-p
- One-Hot Encoding
- Tokenization
- Chat Templates
- System Prompts and Instructions
- Structured Outputs and Constrained Decoding
- Tool Calling
- LLM Agents
- Agent Runtime
- Tool Execution in Agents
- Agent Trust Boundaries
The issue is data
An issue describing a cart crash may also demand that the assistant read an environment file and expose it. The user asked for a bug fix; the issue author cannot grant access to secrets. Indirect prompt injection exploits confusion between external content and authorized instructions.
The model is not the only defense
A runtime can reject unknown tools, out-of-scope paths and unapproved operations. Sandboxing and least privilege reduce reachable harm. Delimiting untrusted text is useful but does not guarantee the model ignores its instructions.
Try it · toy model
A planted line in a bug report tells the agent to leak a secret. Choose the runtime's rules, approve or deny writes as the run pauses for you, and see what leaked, what got fixed and what changed that shouldn't have.
The lab uses a deliberately narrow filename allowlist. Production controls must account for identities, symlinks, races, network destinations and data returned to the user. A well-formed call is neither proof of authority nor proof of safety.
Why should I care?
As a researcher
Test whether attacker-controlled observations can redirect an agent across a trust boundary.
As an engineer
Enforce tool allowlists, least privilege and scoped authorization before execution.
Modern systems that depend on it
- Coding assistants
- Tool-using applications
Historical context
Before
An external page, issue or tool result can embed instructions that conflict with the real task.
After
Treat observations as data and enforce tool scope and authorization outside the model.
Used today
Agents reading repositories, documents or websites while holding tools that can change external state.
What to remember
- Agent trust boundaries prevent untrusted observations from granting permissions or directing unauthorized tool use.
- Treat observations as data and enforce tool scope and authorization outside the model.
- A relevant document is not a trusted source of authority.
Key papers
Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
Kai Greshake, Sahar Abdelnabi et al. · 2023
Showed that text a model retrieves (a web page, an email, a document) can carry instructions an attacker planted, so retrieval and tools are a security boundary.