Concept · Chapter 13: Agents
Agent Evaluation
Agent evaluation checks final outcomes, constraint compliance and repeated-run reliability under a specified budget.
The problem
Fluent transcripts and selected successful runs can hide failed tasks.
The solution
Evaluate from a controlled initial state against explicit success conditions and retain the action trace.
The consequence
Scores depend on the task, harness, tools, model and available compute.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Expected Value and Variance
- Reinforcement Learning
- MDPs, Policies and Value
- Decoding: Greedy, Temperature, Top-k, Top-p
- One-Hot Encoding
- Tokenization
- Chat Templates
- System Prompts and Instructions
- Structured Outputs and Constrained Decoding
- Tool Calling
- LLM Agents
- Agent Reliability and Budgets
- Features, Labels and Tasks
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Logistic Regression
- Evaluation Metrics for Classifiers
- Agent Evaluation
The final state is the target
A coding task can be checked with tests and diff review. A scheduling task needs the correct saved record. Retain traces for diagnosis, but do not substitute apparent effort for a verified result.
τ-bench evaluates tool-and-user interaction against a target database state and examines consistency across repeated trials.
Two questions about repeated attempts
For independent runs of a fixed task with success probability , at least one success in attempts has probability . Success on all has probability . These correspond to the questions behind pass@k and pass^k; empirical benchmarks may use finite-sample estimators.
At , five attempts give a 99.968% chance of at least one success, but only 32.768% of all five succeeding. A showcase choosing the best run and a user depending on repeated operations experience different systems.
Report failures, permission violations, costs and latency as well as success. Compare architectures under similar budgets and check for benchmark contamination or access to hidden evaluation artifacts.
What to remember
- Agent evaluation checks final outcomes, constraint compliance and repeated-run reliability under a specified budget.
- Evaluate from a controlled initial state against explicit success conditions and retain the action trace.
- Scores depend on the task, harness, tools, model and available compute.
Key papers
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Carlos E. Jimenez et al. · 2023
Makes real repository issue resolution an executable evaluation problem.
How to read it: Check how tasks and tests are constructed before interpreting a score.
$τ$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Shunyu Yao et al. · 2024
Evaluates goal-state correctness and consistency over repeated tool-and-user interactions.
How to read it: Distinguish pass^k (all trials succeed) from pass@k (at least one succeeds).
WebArena: A Realistic Web Environment for Building Autonomous Agents
Shuyan Zhou et al. · 2023
Provides reproducible web tasks evaluated for functional completion.
How to read it: Inspect the success evaluator and available observations before comparing agents.
Watch
Latent Space
Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Hear a ReAct author discuss the transition from language-model reasoning to acting and the role of computer interfaces.