Skip to content
Road to Intelligence

Concept · Chapter 13: Agents

LLM Agents

Must knowKnow well12 minDifficulty

An LLM agent repeatedly selects actions from observations to pursue a goal within a software-controlled environment.

The problem

A generated answer does not inspect files, change state or verify the outcome.

The solution

Put model-selected actions inside an observation-and-feedback loop.

The consequence

Useful autonomy depends on the surrounding tools, controls and evidence.

From an answer to an intervention

A model can explain a bug without touching a repository. An agent system lets it request a file, see the content, propose an edit and examine the test result before choosing the next step. External feedback changes what happens next.

This is a particular kind of agent: classical agents can be rule-based or trained with reinforcement learning. An LLM agent need not train during its run. Its context changes while its weights can remain fixed.

The model proposes. The runtime executes.
  1. 01 · Assemble context

    Goal, relevant history, current observations and available tools.

  2. 02 · Choose an action

    The model requests a tool, answers, or asks for help.

  3. 03 · Enforce the boundary

    The runtime validates arguments, permissions and remaining budget.

  4. 04 · Observe and repeat

    A result or error enters the next context. Completion needs evidence.

04 → 01 until verified completion, a stop request, a limit, or a handoff. The external world may change between observations.

Goal, action, observation

For the cart task, the goal is a correct total including the empty list. Reading a file and running tests are actions; their returned contents are observations. The final claim needs evidence about the changed implementation.

The environment is only partially visible. A failed command might reveal a missing dependency rather than a software defect. Good action selection distinguishes those cases.

Try the smallest complete system

Try it · toy model

Watch the Loop

Watch three agents attack the same bug (one checks its work, one trusts itself, one gets stuck) while the budget drains and the context grows; turn off the runtime's evidence check and see a broken fix ship. Or drive the loop yourself.

Know well8 min

A successful run illustrates the mechanism, not general coding ability: the lab is deterministic and has only two candidate patches. Real agents must generate or discover their own alternatives.

Why should I care?

As a researcher

Study action selection under partial observations and distinguish inference-time adaptation from learning weights.

As an engineer

Build a loop whose actions and outcomes can be inspected and reproduced.

Modern systems that depend on it

  • Coding assistants
  • Tool-using applications

Historical context

Before

A generated answer does not inspect files, change state or verify the outcome.

After

Put model-selected actions inside an observation-and-feedback loop.

Used today

Repository assistants, tool-driven research and bounded application automation.

What to remember

  • An LLM agent repeatedly selects actions from observations to pursue a goal within a software-controlled environment.
  • Put model-selected actions inside an observation-and-feedback loop.
  • Useful autonomy depends on the surrounding tools, controls and evidence.

Key papers

Essential

ReAct: Synergizing Reasoning and Acting in Language Models

Shunyu Yao et al. · 2022

Interleaves reasoning, actions and observations in language-model task solving.

How to read it: Compare the action-only and reasoning-only examples with the interleaved trajectory.

~35 min readarXiv:2210.03629✓ verified 2026-10-05

Watch