Skip to content
Road to Intelligence

Concept · Chapter 13: Agents

Agent Evaluation

Must knowUnderstand7 minDifficulty

Agent evaluation checks final outcomes, constraint compliance and repeated-run reliability under a specified budget.

The problem

Fluent transcripts and selected successful runs can hide failed tasks.

The solution

Evaluate from a controlled initial state against explicit success conditions and retain the action trace.

The consequence

Scores depend on the task, harness, tools, model and available compute.

The final state is the target

A coding task can be checked with tests and diff review. A scheduling task needs the correct saved record. Retain traces for diagnosis, but do not substitute apparent effort for a verified result.

τ-bench evaluates tool-and-user interaction against a target database state and examines consistency across repeated trials.

Two questions about repeated attempts

For independent runs of a fixed task with success probability pp, at least one success in kk attempts has probability 1−(1−p)k1-(1-p)^k. Success on all kk has probability pkp^k. These correspond to the questions behind pass@k and pass^k; empirical benchmarks may use finite-sample estimators.

At p=0.8p=0.8, five attempts give a 99.968% chance of at least one success, but only 32.768% of all five succeeding. A showcase choosing the best run and a user depending on repeated operations experience different systems.

Report failures, permission violations, costs and latency as well as success. Compare architectures under similar budgets and check for benchmark contamination or access to hidden evaluation artifacts.

What to remember

  • Agent evaluation checks final outcomes, constraint compliance and repeated-run reliability under a specified budget.
  • Evaluate from a controlled initial state against explicit success conditions and retain the action trace.
  • Scores depend on the task, harness, tools, model and available compute.

Key papers

Essential

SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Carlos E. Jimenez et al. · 2023

Makes real repository issue resolution an executable evaluation problem.

How to read it: Check how tasks and tests are constructed before interpreting a score.

~35 min readarXiv:2310.06770✓ verified 2026-10-05

Watch