Concept · Chapter 13: Agents
Coding Agents
Coding agents inspect and modify software through tools, using executable checks to guide and verify changes.
The problem
A plausible patch may misunderstand the repository or break existing behavior.
The solution
Alternate repository inspection, targeted edits and checks tied to the current code.
The consequence
Executable feedback is valuable but only covers the tested behavior.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Expected Value and Variance
- Reinforcement Learning
- MDPs, Policies and Value
- Decoding: Greedy, Temperature, Top-k, Top-p
- One-Hot Encoding
- Tokenization
- Chat Templates
- System Prompts and Instructions
- Structured Outputs and Constrained Decoding
- Tool Calling
- LLM Agents
- Agent Runtime
- Tool Execution in Agents
- Search
- Planning
- Agent Planning and ReAct
- Coding Agents
The repository is the environment
The task is more than completing a function. An agent needs to locate relevant code, understand its callers, reproduce an issue, make a change and check the result. SWE-bench evaluates real issue-derived repository tasks; SWE-agent explores a purpose-built interface for this work.
A green test is scoped evidence
A test result applies to a particular version in a particular environment. Editing invalidates it. A broken test environment should not be reported as a product regression. Passing a small suite does not establish correctness beyond its coverage.
Try it · toy model
Watch three agents attack the same bug (one checks its work, one trusts itself, one gets stuck) while the budget drains and the context grows; turn off the runtime's evidence check and see a broken fix ship. Or drive the loop yourself.
For the toy bug, returning zero demonstrates why reproducing the reported case is insufficient: the ordinary cart still matters. In a real repository, also review the diff for unintended changes and preserve the human review boundary.
What to remember
- Coding agents inspect and modify software through tools, using executable checks to guide and verify changes.
- Alternate repository inspection, targeted edits and checks tied to the current code.
- Executable feedback is valuable but only covers the tested behavior.
Key papers
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Carlos E. Jimenez et al. · 2023
Makes real repository issue resolution an executable evaluation problem.
How to read it: Check how tasks and tests are constructed before interpreting a score.
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
John Yang et al. · 2024
Treats the interface to a computer as part of an agent system worth designing and evaluating.
How to read it: Read the interface design and ablations; ask what changed besides the model.