Skip to content
Road to Intelligence

Concept · Chapter 13: Agents

Agent Reliability and Budgets

Must knowKnow well12 minDifficulty

Agent reliability measures completing a whole task correctly under limits, not merely choosing one plausible action.

The problem

Small action errors accumulate while retries consume time and compute.

The solution

Measure end-to-end success, add observable checks and bound recovery attempts.

The consequence

Feedback can improve completion rates but does not eliminate correlated errors.

A chain of required steps

Under independent identical success probability pp, completing all nn required steps has probability pnp^n. At 95% per step, twenty required successes occur only about 35.8% of the time.

This is a toy calculation. In the general case, use the chain rule: multiply each step's success probability conditional on the previous steps succeeding. Different steps need not share a probability or be independent.

Allow one recovery

If a fraction dd of initial failures is detected, and each detected failure gets one independent retry with the same success probability, then

peffective=p+(1−p)dp.p_{\text{effective}} = p + (1-p)dp.

At p=0.95p=0.95, d=0.8d=0.8, this becomes 0.988; twenty such steps succeed about 78.5% of the time. Perfect detection still leaves the possibility of a failed retry.

Try it · toy model

Run It 100 Times

Simulate a hundred multi-step tasks and watch small per-step failure rates compound; add retries, then make retries repeat the same mistake and see the textbook formula overstate success.

Know well8 min

Cost changes the question

Expected attempts are n[1+(1−p)d]n[1+(1-p)d] if every step is visited. This excludes checking cost and is not the cost of a fail-fast pipeline. Report tool calls, token usage, latency and success within a specified budget. A retry driven by the same wrong belief may fail again, violating independence.

Why should I care?

As a researcher

Measure conditional failure and repeated-trial consistency instead of extrapolating from single-step accuracy.

As an engineer

Set budgets, instrument recovery and know when a retry could duplicate a side effect.

Modern systems that depend on it

  • Coding assistants
  • Tool-using applications

Historical context

Before

Small action errors accumulate while retries consume time and compute.

After

Measure end-to-end success, add observable checks and bound recovery attempts.

Used today

Evaluation and operational monitoring of multi-step tool-using systems.

What to remember

  • Agent reliability measures completing a whole task correctly under limits, not merely choosing one plausible action.
  • Measure end-to-end success, add observable checks and bound recovery attempts.
  • Feedback can improve completion rates but does not eliminate correlated errors.

Key papers