Skip to content
Road to Intelligence

Concept · Chapter 13: Agents

Coding Agents

Should knowUnderstand7 minDifficulty

Coding agents inspect and modify software through tools, using executable checks to guide and verify changes.

The problem

A plausible patch may misunderstand the repository or break existing behavior.

The solution

Alternate repository inspection, targeted edits and checks tied to the current code.

The consequence

Executable feedback is valuable but only covers the tested behavior.

The repository is the environment

The task is more than completing a function. An agent needs to locate relevant code, understand its callers, reproduce an issue, make a change and check the result. SWE-bench evaluates real issue-derived repository tasks; SWE-agent explores a purpose-built interface for this work.

A green test is scoped evidence

A test result applies to a particular version in a particular environment. Editing invalidates it. A broken test environment should not be reported as a product regression. Passing a small suite does not establish correctness beyond its coverage.

Try it · toy model

Watch the Loop

Watch three agents attack the same bug (one checks its work, one trusts itself, one gets stuck) while the budget drains and the context grows; turn off the runtime's evidence check and see a broken fix ship. Or drive the loop yourself.

Know well8 min

For the toy bug, returning zero demonstrates why reproducing the reported case is insufficient: the ordinary cart still matters. In a real repository, also review the diff for unintended changes and preserve the human review boundary.

What to remember

  • Coding agents inspect and modify software through tools, using executable checks to guide and verify changes.
  • Alternate repository inspection, targeted edits and checks tied to the current code.
  • Executable feedback is valuable but only covers the tested behavior.

Key papers

Essential

SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Carlos E. Jimenez et al. · 2023

Makes real repository issue resolution an executable evaluation problem.

How to read it: Check how tasks and tests are constructed before interpreting a score.

~35 min readarXiv:2310.06770✓ verified 2026-10-05