Concept · Chapter 13: Agents
Agent Runtime
The agent runtime assembles context, routes authorized tool calls, records observations and enforces execution limits.
The problem
Generated action text has no effect until software interprets and executes it.
The solution
Keep execution, permissions, state and budget accounting in application code.
The consequence
A harness can enforce boundaries even when the model proposes a bad action.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Expected Value and Variance
- Reinforcement Learning
- MDPs, Policies and Value
- Decoding: Greedy, Temperature, Top-k, Top-p
- One-Hot Encoding
- Tokenization
- Chat Templates
- System Prompts and Instructions
- Structured Outputs and Constrained Decoding
- Tool Calling
- LLM Agents
- Agent Runtime
The program around the model
The runtime decides what tools exist, which arguments are valid and what permissions apply. It records results and passes relevant observations back to the model. A model can request a stop, but the runtime also needs independent cancellation and budget limits.
01 · Assemble context
Goal, relevant history, current observations and available tools.
02 · Choose an action
The model requests a tool, answers, or asks for help.
03 · Enforce the boundary
The runtime validates arguments, permissions and remaining budget.
04 · Observe and repeat
A result or error enters the next context. Completion needs evidence.
04 → 01 until verified completion, a stop request, a limit, or a handoff. The external world may change between observations.
Failure is a state
A tool response may be success, a known failure or an unknown outcome. If a write times out after the server applied it, repeating it could duplicate the effect. Preserve an operation ID, query the state and use idempotent operations where available.
Durable checkpoints should retain the goal, permission scope, completed operations and pending work. Reconstructing state from a fluent conversation summary alone can lose those distinctions.
Test the boundary: what happens if a model invents a tool name, provides malformed arguments, asks for an unavailable file, or never returns a final answer? Each needs defined runtime behavior.
Why should I care?
As a researcher
Compare harness designs while holding the underlying model and budget fixed.
As an engineer
Handle tool rejection, uncertain outcomes, durable checkpoints and cancellation explicitly.
Modern systems that depend on it
- Coding assistants
- Tool-using applications
Historical context
Before
Generated action text has no effect until software interprets and executes it.
After
Keep execution, permissions, state and budget accounting in application code.
Used today
Applications that route model tool requests to local programs or remote APIs.
What to remember
- The agent runtime assembles context, routes authorized tool calls, records observations and enforces execution limits.
- Keep execution, permissions, state and budget accounting in application code.
- A harness can enforce boundaries even when the model proposes a bad action.