Concept · Chapter 13: Agents
Tool Execution in Agents
A model proposes a tool call, but separate software validates, authorizes and executes it.
The problem
Well-formed output can still request an irrelevant or unauthorized action.
The solution
Check syntax, scope and permission before dispatch, then inspect the resulting evidence.
The consequence
A schema is an interface contract rather than a guarantee of correctness.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Expected Value and Variance
- Reinforcement Learning
- MDPs, Policies and Value
- Decoding: Greedy, Temperature, Top-k, Top-p
- One-Hot Encoding
- Tokenization
- Chat Templates
- System Prompts and Instructions
- Structured Outputs and Constrained Decoding
- Tool Calling
- LLM Agents
- Agent Runtime
- Tool Execution in Agents
Three checks
For a proposed read of src/total.ts, the schema can require a string path. The permission layer checks whether that file is in scope. The returned content determines whether the action actually helped.
Clear tool descriptions reduce ambiguity. Clear observations reduce uncertainty: return the affected version, relevant excerpt or explicit failure rather than a generic “OK.” Toolformer studies learning API use; SWE-agent studies an interface designed around an agent's needs.
Writes need stronger semantics
A tool that patches a file can require its expected current version to prevent applying an edit to stale content. A payment tool can use an idempotency key to avoid repeating a transaction after a timeout. These properties belong to the tool and runtime, even if the model knows how to describe them.
Try it · toy model
A planted line in a bug report tells the agent to leak a secret. Choose the runtime's rules, approve or deny writes as the run pauses for you, and see what leaked, what got fixed and what changed that shouldn't have.
What to remember
- A model proposes a tool call, but separate software validates, authorizes and executes it.
- Check syntax, scope and permission before dispatch, then inspect the resulting evidence.
- A schema is an interface contract rather than a guarantee of correctness.
Key papers
Toolformer: Language Models Can Teach Themselves to Use Tools
Timo Schick, Jane Dwivedi-Yu et al. · 2023
Showed a model can learn when to call a calculator, search engine or calendar and how to use the result, the core idea behind tool calling.
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
John Yang et al. · 2024
Treats the interface to a computer as part of an agent system worth designing and evaluating.
How to read it: Read the interface design and ablations; ask what changed besides the model.