Agent system design: patterns, context & control
Last reviewed: 2026-07-16 — see the freshness policy.
Learning objectives
After this chapter you will be able to:
- choose between a deterministic workflow, a single agent, and a multi-agent system;
- map a workload to routing, parallel, sequential, critique, coordinator, or autonomous patterns;
- design context, tools, state, budgets, stop conditions, and human control as explicit subsystems;
- identify when multi-agent decomposition adds ceremony rather than capability;
- produce an agent design record that can be evaluated and operated.
The decision
What is the least autonomous control structure that can solve this task reliably within its risk, latency, and cost limits?
This chapter synthesizes two complementary primary guides: Anthropic's Building effective agents and Google Cloud's agentic architecture component and design-pattern guides. Product names differ; the durable architecture lessons largely agree:
- use simple composition before open-ended autonomy;
- make control flow and stopping behavior explicit;
- give agents well-designed tools and just enough context;
- evaluate the whole harness and environment, not only the model;
- accept multi-agent cost and complexity only for a measurable benefit.
Workflow and agent are different control choices
Anthropic distinguishes:
- workflows — models and tools follow control flow predefined in code;
- agents — the model dynamically chooses its process and tool use.
Both are agentic systems in the broad sense, but they have different failure surfaces. A deterministic sequence is easier to test, resume, and explain. Dynamic agents are useful when the correct path cannot be enumerated economically in advance.
Use ordinary code for arithmetic, validation, authorization, state transitions, and known business rules. Use model judgment where language ambiguity, open-ended planning, or semantic interpretation makes fixed rules brittle.
A unified pattern catalog
| Pattern | Control flow | Use when | Avoid when |
|---|---|---|---|
| Prompt chain / sequential | Fixed ordered stages | Each stage has a clear contract and gate | Later stages need to skip or reorder dynamically |
| Router | Classifier selects a known path | Inputs fall into stable categories | Categories overlap heavily or change constantly |
| Parallel fan-out / gather | Independent work runs concurrently | Subtasks are independent and latency matters | Synthesis is harder than the original task |
| Orchestrator-workers | Model decomposes and delegates | Number and shape of subtasks are not known ahead | A fixed decomposition works |
| Evaluator-optimizer | Draft and critique repeat | Criteria are explicit and iteration measurably helps | There is no reliable evaluator or stop rule |
| Review-and-critique | Separate validation step | A distinct checker catches meaningful failures | Reviewer shares the same blind spot and evidence |
| Human-in-the-loop | Workflow pauses for a person | Consequence, uncertainty, or accountability demands review | Review is ceremonial and reviewers lack context/time |
| Autonomous tool loop | Model plans, acts, observes, adapts | Environment is open-ended and steps cannot be scripted | Actions are high impact, irreversible, or poorly observable |
| Hierarchical / swarm | Multiple dynamic agents coordinate | Specialized context or parallel search shows measured gains | Adopted because “multi-agent” sounds scalable |
Google's pattern guide distinguishes deterministic sequential/parallel flows from dynamic coordinator, hierarchical, swarm, ReAct, iterative-refinement, and custom-logic patterns. Anthropic describes prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer, and autonomous agents. The naming is less important than answering: who chooses the next step—code, a model, or a human?
The agent architecture is more than the loop
| Component | Required decision |
|---|---|
| Interface | How users provide goals, constraints, consent, and corrections |
| Policy enforcement | What the agent may see, decide, spend, and do |
| Controller / harness | How calls, tools, retries, checkpoints, and stop rules work |
| Model | Which capability, context, latency, and deployment profile fits |
| Context builder | Which instructions, evidence, history, and tool metadata enter each step |
| Tools | Narrow operations with typed contracts and independent authorization |
| Working state | Current plan, observations, outputs, budgets, and task status |
| Memory | Explicitly selected facts allowed to persist beyond the immediate task |
| Runtime | Isolation, networking, credentials, concurrency, and resource limits |
| Evaluation | Tasks, trials, outcomes, graders, and release thresholds |
| Observability | Traces, state changes, tool calls, policy decisions, outcomes, and cost |
| Human control | Approval, correction, cancellation, escalation, and recovery |
An agent framework may implement parts of this table. It does not make the decisions for you.
Write the loop contract
Before choosing a framework, specify:
Goal:
Success state:
Allowed observations:
Allowed tools and scopes:
Prohibited actions:
Maximum steps / model calls / tool calls:
Token, cost, and wall-clock budget:
Checkpoint frequency:
Conditions requiring approval:
Conditions requiring clarification:
Stop conditions:
Cancellation and compensation behavior:
Final evidence required:
A vague goal plus unrestricted tools is not flexibility; it is an undefined control system.
Context engineering for long-running agents
Anthropic's context-engineering guidance describes context as a finite resource and highlights compaction, structured notes, and subagents for longer horizons.
Apply four rules:
- Keep stable policy separate from transient evidence. System instructions, permissions, and task state should not disappear during summarization.
- Retrieve just in time. Do not preload every document, tool description, or conversation into every step.
- Compact with provenance. Summaries should link to the original event, artifact, or source so critical detail can be recovered.
- Externalize durable state. Store plans, completed steps, decisions, tests, and open questions in a typed task store—not only in conversation tokens.
For context resets, create a handoff packet containing objective, verified state, artifacts, decisions, failed approaches, remaining work, budgets, and next safe action. Anthropic's long-running harness guidance similarly emphasizes context resets plus structured handoffs to preserve coherence.
Tools are interfaces for an unfamiliar caller
An agent selects tools from names, descriptions, schemas, and previous results. Anthropic's tool-design guide recommends treating tool descriptions and specifications as a major performance surface.
Good tools are:
- narrow enough to authorize and understand;
- typed and explicit about units, formats, and identifiers;
- clear about read versus write behavior;
- idempotent where retries are possible;
- concise in their returned context;
- explicit about errors and partial success;
- backed by server-side validation and authorization;
- observable without logging unnecessary sensitive content.
Prefer preview_refund and execute_approved_refund over one broad manage_account tool. The split creates a security boundary and a useful approval point.
Single agent before multi-agent
Multi-agent architecture is justified when at least one is true:
- subtasks require genuinely different tools, permissions, or context;
- independent work can run in parallel and reduce wall-clock time;
- context isolation prevents one large prompt from becoming polluted;
- a separate verifier has independent evidence or executable checks;
- teams or organizations own agents behind stable delegation contracts.
It is not justified merely because software microservices are successful. Each additional agent introduces another stochastic call, context boundary, state transition, failure mode, and cost center.
The peer-reviewed AgentBench study found long-term reasoning, decision-making, and instruction following to be common obstacles in interactive environments. The τ-bench study further shows why repeated reliability and policy following matter in tool-agent-user interactions. Neither result supports adding more agents by default; both support measuring complete trajectories and outcomes.
Reliability math is unforgiving
If a workflow has five required stages and each independently succeeds 95% of the time, the idealized end-to-end success rate is:
0.95^5 ≈ 77%
Real stages are not independent, and recovery can improve the result, but the example exposes why “each agent looked good in a demo” is insufficient. Measure end-to-end success, not the average score of subagents.
Parallelism also changes budgets. Four simultaneous workers may reduce latency but multiply tokens, provider rate-limit pressure, tool contention, and synthesis conflicts.
Safety and control plane
Keep these outside model discretion:
- authentication and authorization;
- tenant and data-policy enforcement;
- spending and time limits;
- allowed network destinations;
- destructive-action approvals;
- idempotency and transactional invariants;
- secrets and credential issuance;
- audit retention;
- emergency stop and task cancellation.
Anthropic's trustworthy-agents principles emphasize human control, secure interactions, transparency, privacy, and alignment. Translate principles into mechanisms: preview modes, staged execution, narrow scopes, visible state, reversible operations, and independent policy checks.
Northstar pattern selection
Northstar wants an agent to resolve delayed-shipment requests.
Candidate A: one autonomous agent
The agent reads account data, contacts logistics, decides compensation, and changes the order. It is flexible but has excessive authority and a large failure surface.
Candidate B: bounded hybrid
- A deterministic router classifies the request and checks identity.
- Account and order data are retrieved through read-only tools.
- A single agent interprets the customer's goal and selects an approved workflow.
- Shipment investigation is delegated to the logistics agent through A2A.
- Code calculates policy-allowed compensation.
- The agent drafts options and evidence.
- The customer or authorized employee approves consequential action.
- A narrow tool executes an idempotent transaction.
- Outcome checks confirm the state before the agent reports completion.
Candidate B uses model judgment where language is ambiguous and deterministic control where money and state change.
Agent design review questions
- Can this be a normal workflow with a model inside one or two steps?
- Which decisions truly require dynamic planning?
- What observable state proves success?
- What prevents a model from acquiring or exercising excess authority?
- Which context must be stable, retrieved, compacted, or persisted?
- What are the maximum depth, fan-out, retries, time, and spend?
- How does a human understand, interrupt, correct, and recover the task?
- Which failures are contained locally, and which cascade?
- Does each extra agent improve measured quality, latency, isolation, or ownership?
- Can the design degrade to a safe workflow when models or tools fail?
Architecture artifact: agent pattern ADR
Include:
- workload and uncertainty characteristics;
- selected control-flow pattern;
- rejected simpler pattern and evidence for rejection;
- model-controlled versus code-controlled decisions;
- context, state, memory, and handoff design;
- tools, identities, and authorization boundaries;
- budgets and stopping rules;
- approval and recovery points;
- end-to-end evaluation and SLOs;
- complexity budget for every additional agent.
Lab
Implement the Northstar delayed-shipment workflow three ways:
- deterministic chain with one model interpretation step;
- single tool-using agent;
- coordinator plus two specialized workers.
Run the same evaluation suite across all three. Compare task success, policy violations, p95 latency, model/tool calls, cost per successful task, failure localization, and operator debugging time. Select the simplest design whose evidence clears the release gates.
Check yourself
- What is the operational difference between a workflow and an agent?
- When is an orchestrator-workers pattern better than a fixed parallel workflow?
- Why should stable policy survive context compaction unchanged?
- Name three valid reasons to introduce another agent.
- Which parts of an agent system must remain outside model discretion?
Further reading
- Anthropic, Building effective agents.
- Anthropic, Effective context engineering for AI agents.
- Anthropic, Writing effective tools for agents.
- Anthropic, Harness design for long-running application development.
- Google Cloud Architecture Center, Choose your agentic AI architecture components.
- Google Cloud Architecture Center, Choose a design pattern for your agentic AI system.
- Liu et al., AgentBench: Evaluating LLMs as Agents, ICLR 2024.
- Yao et al., τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains, 2024 preprint.