Concept · Chapter 13: Agents
Agent Planning and ReAct
Agent planning organizes actions toward a goal and revises them as new observations arrive.
The problem
A plan made before inspecting the environment can become wrong after the first action.
The solution
Use short plans with observable checkpoints and update them from feedback.
The consequence
The useful plan is the one that directs the next evidence-gathering action.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Expected Value and Variance
- Reinforcement Learning
- MDPs, Policies and Value
- Decoding: Greedy, Temperature, Top-k, Top-p
- One-Hot Encoding
- Tokenization
- Chat Templates
- System Prompts and Instructions
- Structured Outputs and Constrained Decoding
- Tool Calling
- LLM Agents
- Search
- Planning
- Agent Planning and ReAct
Open loop and closed loop
“Read, edit, test” is an initial plan. If reading reveals that the function is imported from another package, the plan must change. Closed-loop control uses the new observation; an open-loop sequence blindly continues.
ReAct interleaves generated reasoning, actions and observations. This does not require every modern agent to display private reasoning or produce a thought paragraph before each tool call.
Plan around uncertainty
For the cart bug, first establish whether the failure reproduces. After patching, establish whether normal inputs still work. Each check answers a question that affects the next action.
A checklist is not a proof. “Tests passed” needs a result tied to the current code, not a tick added because the agent intended to run them. Compare a planning system against a simpler loop at equal budget before attributing better results to its plan format.
What to remember
- Agent planning organizes actions toward a goal and revises them as new observations arrive.
- Use short plans with observable checkpoints and update them from feedback.
- The useful plan is the one that directs the next evidence-gathering action.
Key papers
ReAct: Synergizing Reasoning and Acting in Language Models
Shunyu Yao et al. · 2022
Interleaves reasoning, actions and observations in language-model task solving.
How to read it: Compare the action-only and reasoning-only examples with the interleaved trajectory.
Watch
Latent Space
Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Hear a ReAct author discuss the transition from language-model reasoning to acting and the role of computer interfaces.