Concept · Chapter 13: Agents
Computer-Use Agents
Computer-use agents act through interface controls and inspect fresh observations to determine the resulting state.
The problem
Many tasks are exposed through applications rather than a stable dedicated API.
The solution
Use an observe-act-observe loop over the page, accessibility tree or screen.
The consequence
Dynamic interfaces make stale observations and ambiguous outcomes costly.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Expected Value and Variance
- Reinforcement Learning
- MDPs, Policies and Value
- Decoding: Greedy, Temperature, Top-k, Top-p
- One-Hot Encoding
- Tokenization
- Chat Templates
- System Prompts and Instructions
- Structured Outputs and Constrained Decoding
- Tool Calling
- LLM Agents
- Agent Runtime
- Tool Execution in Agents
- Computer-Use Agents
Clicking is not completion
An agent can select Save and still fail to save: validation may block it, the network may time out, or the click may hit a different control after layout changes. Inspect the resulting state before declaring success.
WebArena provides reproducible web environments and evaluates functional task completion. It motivates checking what changed, not merely whether the action trace looks plausible.
Prefer observable actions
Use stable semantic controls when available. After navigation, obtain a fresh observation. A screenshot supplies visual evidence but not an omniscient view of application state; hidden records and pending requests can matter.
Side-effecting UI operations need the same permission boundaries as API calls. A purchase confirmation screen is an opportunity to review a specific operation, not proof that the purchase has already happened. A timeout may require reconciling state rather than clicking again.
What to remember
- Computer-use agents act through interface controls and inspect fresh observations to determine the resulting state.
- Use an observe-act-observe loop over the page, accessibility tree or screen.
- Dynamic interfaces make stale observations and ambiguous outcomes costly.
Key papers
WebArena: A Realistic Web Environment for Building Autonomous Agents
Shuyan Zhou et al. · 2023
Provides reproducible web tasks evaluated for functional completion.
How to read it: Inspect the success evaluator and available observations before comparing agents.