Skip to content
Road to Intelligence

Concept · Chapter 13: Agents

Agent Trust Boundaries

Must knowKnow well12 minDifficulty

Agent trust boundaries prevent untrusted observations from granting permissions or directing unauthorized tool use.

The problem

An external page, issue or tool result can embed instructions that conflict with the real task.

The solution

Treat observations as data and enforce tool scope and authorization outside the model.

The consequence

A relevant document is not a trusted source of authority.

The issue is data

An issue describing a cart crash may also demand that the assistant read an environment file and expose it. The user asked for a bug fix; the issue author cannot grant access to secrets. Indirect prompt injection exploits confusion between external content and authorized instructions.

The model is not the only defense

A runtime can reject unknown tools, out-of-scope paths and unapproved operations. Sandboxing and least privilege reduce reachable harm. Delimiting untrusted text is useful but does not guarantee the model ignores its instructions.

Try it · toy model

Hold the Boundary

A planted line in a bug report tells the agent to leak a secret. Choose the runtime's rules, approve or deny writes as the run pauses for you, and see what leaked, what got fixed and what changed that shouldn't have.

Know well8 min

The lab uses a deliberately narrow filename allowlist. Production controls must account for identities, symlinks, races, network destinations and data returned to the user. A well-formed call is neither proof of authority nor proof of safety.

Why should I care?

As a researcher

Test whether attacker-controlled observations can redirect an agent across a trust boundary.

As an engineer

Enforce tool allowlists, least privilege and scoped authorization before execution.

Modern systems that depend on it

  • Coding assistants
  • Tool-using applications

Historical context

Before

An external page, issue or tool result can embed instructions that conflict with the real task.

After

Treat observations as data and enforce tool scope and authorization outside the model.

Used today

Agents reading repositories, documents or websites while holding tools that can change external state.

What to remember

  • Agent trust boundaries prevent untrusted observations from granting permissions or directing unauthorized tool use.
  • Treat observations as data and enforce tool scope and authorization outside the model.
  • A relevant document is not a trusted source of authority.

Key papers