Concept · Chapter 13: Agents
Agent Reliability and Budgets
Agent reliability measures completing a whole task correctly under limits, not merely choosing one plausible action.
The problem
Small action errors accumulate while retries consume time and compute.
The solution
Measure end-to-end success, add observable checks and bound recovery attempts.
The consequence
Feedback can improve completion rates but does not eliminate correlated errors.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Expected Value and Variance
- Reinforcement Learning
- MDPs, Policies and Value
- Decoding: Greedy, Temperature, Top-k, Top-p
- One-Hot Encoding
- Tokenization
- Chat Templates
- System Prompts and Instructions
- Structured Outputs and Constrained Decoding
- Tool Calling
- LLM Agents
- Agent Reliability and Budgets
A chain of required steps
Under independent identical success probability , completing all required steps has probability . At 95% per step, twenty required successes occur only about 35.8% of the time.
This is a toy calculation. In the general case, use the chain rule: multiply each step's success probability conditional on the previous steps succeeding. Different steps need not share a probability or be independent.
Allow one recovery
If a fraction of initial failures is detected, and each detected failure gets one independent retry with the same success probability, then
At , , this becomes 0.988; twenty such steps succeed about 78.5% of the time. Perfect detection still leaves the possibility of a failed retry.
Try it · toy model
Simulate a hundred multi-step tasks and watch small per-step failure rates compound; add retries, then make retries repeat the same mistake and see the textbook formula overstate success.
Cost changes the question
Expected attempts are if every step is visited. This excludes checking cost and is not the cost of a fail-fast pipeline. Report tool calls, token usage, latency and success within a specified budget. A retry driven by the same wrong belief may fail again, violating independence.
Why should I care?
As a researcher
Measure conditional failure and repeated-trial consistency instead of extrapolating from single-step accuracy.
As an engineer
Set budgets, instrument recovery and know when a retry could duplicate a side effect.
Modern systems that depend on it
- Coding assistants
- Tool-using applications
Historical context
Before
Small action errors accumulate while retries consume time and compute.
After
Measure end-to-end success, add observable checks and bound recovery attempts.
Used today
Evaluation and operational monitoring of multi-step tool-using systems.
What to remember
- Agent reliability measures completing a whole task correctly under limits, not merely choosing one plausible action.
- Measure end-to-end success, add observable checks and bound recovery attempts.
- Feedback can improve completion rates but does not eliminate correlated errors.
Key papers
$τ$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Shunyu Yao et al. · 2024
Evaluates goal-state correctness and consistency over repeated tool-and-user interactions.
How to read it: Distinguish pass^k (all trials succeed) from pass@k (at least one succeeds).