For day-to-day agent-assisted work, start with the operating loop. If you are designing an agent workflow or product, the architecture patterns below help you compare options. Choose the detail relevant to your task.
Agent Architecture Patterns
Agent Anatomy
These five parts help explain how an agent works.
| Component | What It Does | Example |
|---|---|---|
| Model | A language model (LLM) generates outputs and may choose which tools to call. | Drafting an answer or selecting a tool |
| Tools | External functions, APIs, or systems the agent can invoke. | Database queries, web search, code execution, file I/O |
| Instructions | Explicit guidelines, scope constraints, and behavioral rules. | System prompts, AGENTS.md, rules files, policy documents |
| Memory | Information retained across interactions, with access and retention limits. | Conversation history, task state, saved project context |
| Retrieval | Looks up information the task needs. | Document search, a knowledge base, project files |
Products combine these parts differently. Check which information and tools your agent actually uses.
Workflows vs. Agents
Anthropic offers a useful distinction between workflows and agents:
| Dimension | Workflows | Agents |
|---|---|---|
| Control | Follows steps defined in advance | Chooses steps as the task proceeds |
| Predictability | The path is defined; results can still vary | Both the path and results can vary |
| Flexibility | Limited to the paths you define | Can choose among permitted actions |
| Best for | Well-defined, repeatable tasks | Open-ended problems with unpredictable steps |
| Cost/latency | Measure calls and time for the workflow | Set limits and measure time and cost |
| Error handling | Defined checks and recovery steps | Feedback, stopping conditions and escalation |
Use a workflow when you can define the steps in advance. Consider an agent when the next step depends on what it finds. In either case, agree how to check the result and when to stop.
Tool Types
These three tool categories help you plan access and checks:
Data tools retrieve context and information. They can expose sensitive data even when they cannot change it. Examples: Query databases, read documents, search the web, pull CRM records.
Action tools change something in a system. Consider who is affected and how a mistake could be corrected. Examples: Send emails, update records, create tickets, issue refunds, deploy code.
Orchestration tools let an agent ask another agent to carry out part of the work. Examples: A “research agent” callable by a “manager agent,” a specialist agent invoked by a triage agent.
Assess each tool’s risk by what it can access, who could be affected, and whether a mistake can be undone. Use that assessment to choose the limits and checks it needs.
Composition Patterns
These eight patterns are options to compare. Start with the simplest approach that meets the task’s needs; additional agents or steps should address a problem you have observed.
Prompt Chaining
Split a task into a fixed sequence of steps. Each model call uses the previous result, with checks between steps.
When to consider it
The task has clear steps. Compare whether the extra calls improve quality enough to justify their time and cost.
Example
Generate marketing copy, then translate it. Write an outline, validate it against criteria, then write the full document.
Routing
Classify the input and direct it to a specialized handler. Each route has its own optimized prompt and tools.
When to consider it
Distinct categories that are better handled separately. Classification can be done accurately.
Example
Customer service — route general questions, refund requests, and technical support to different downstream processes.
Parallelization
Run subtasks simultaneously and aggregate results. Two variants: Sectioning (independent subtasks) and Voting (same task, multiple perspectives).
When to consider it
Subtasks are independent (sectioning) or you need multiple perspectives (voting).
Example
One model processes the user query while another screens for safety. Multiple prompts review code for vulnerabilities; flag if any finds a problem.
Evaluator-Optimizer
One model generates a response and another reviews it. Repeat within agreed limits, and stop for human review when needed.
When to consider it
Clear evaluation criteria exist, and iterative refinement provides measurable improvement.
Example
Literary translation with a critic loop. Complex search tasks requiring multiple rounds of analysis.
Single Agent Loop
One model uses tools and feedback until it finishes, reaches a limit or needs help.
When to consider it
Dynamic decision-making about which tools to call and in what order, but complexity does not warrant splitting across multiple agents.
Example
A coding agent that reads files, writes code, runs tests, and iterates until tests pass.
Orchestrator-Workers
A central LLM dynamically breaks down tasks, delegates to worker LLMs, and synthesizes results. Unlike parallelization, subtasks are not pre-defined.
When to consider it
Complex tasks where you cannot predict the number or nature of subtasks in advance.
Example
A coding product that determines which files need changing and dispatches changes to workers.
Manager (Agents-as-Tools)
A central "manager" agent calls specialized agents as tools. The manager retains control and context, synthesizing outputs into a unified interaction.
When to consider it
You want a single agent maintaining central control and user interaction while delegating specialized work.
Example
A manager agent that calls translator agents for Spanish, French, and Italian, synthesizing all results for the user.
Decentralized Handoff
Agents operate as peers, handing off full execution control to one another based on specialization. No central coordinator.
When to consider it
You don't need central control or synthesis. Each specialized agent can fully take over the interaction.
Example
A triage agent hands off entirely to technical support, sales, or order management. The receiving agent owns the conversation.
Pattern Selection Guide
Use the problem you have observed to choose an option. Compare it with the simplest approach that could meet the task’s needs.
Keep it simple: Add another agent only when it addresses a problem the simpler approach cannot handle well.
The Guardrail Stack
The five layers help you review scope, quality, policy, human decisions and shared oversight. Consider each layer before expanding the work. The controls you need depend on the task and its consequences. See Principle 4: Guardrails Are Non-Negotiable.
Fleet-level controls for managing agents at organizational scale.
| Element | Description |
|---|---|
| Registry | Single source of truth tracking all agents, their capabilities, owners, and status |
| Access control | Role-based permissions determining which agents can access which systems and data |
| Observability | Unified monitoring across all agents — execution traces, cost tracking, error rates, latency |
| Interoperability | Standards for agents to work across platforms and teams (e.g., Model Context Protocol) |
| Audit trail | Record relevant actions, approvals and outcomes, with appropriate access and retention |
| Cost budgeting | Per-agent and per-team token/compute budgets with alerts and hard limits |
Ensure humans retain authority over decisions that agents must not make autonomously.
| Element | Description |
|---|---|
| Architecture | System design, technology choices, data model changes |
| Risk acceptance | Shipping known tradeoffs, accepting technical debt |
| Release timing | When code goes to production |
| Incident response | Rollback decisions, postmortem actions |
| Security-critical changes | Authentication, authorization, encryption |
| Cost commitments | Actions with financial impact above defined thresholds |
Apply the privacy, security and organizational controls the work needs.
| Element | Description | Example |
|---|---|---|
| No secret exposure | Automated secret scanning in pre-commit and CI | Credentials leaking into repositories |
| PII filtering | Limit access to personal information and check outputs for unintended disclosure | Privacy violations in generated content |
| Safety classification | Treat untrusted input carefully; detection can help but may miss attempts to redirect an agent | System exploitation |
| Relevance classification | Flag off-topic or out-of-scope agent behavior | Scope drift and waste |
| Moderation | Content safety checks on agent outputs | Harmful or inappropriate generated content |
| Dependency policy | Block unsafe dependency upgrades or additions | Supply chain attacks |
| Branch policy | No direct pushes to main/protected branches | Unreviewed code reaching production |
Use automated checks and human review to find problems before the work is used.
| Element | Description |
|---|---|
| Formatting & linting | Enforce style consistency (Black, ESLint, Prettier, etc.) |
| Type checking | Static type verification (mypy, TypeScript strict mode) |
| Unit & integration tests | Run relevant tests and add meaningful checks for the change |
| Static analysis | Security scanning, dependency vulnerability checks |
| Coverage thresholds | Review coverage alongside the importance of the cases being tested |
| Design system compliance | Agent-generated UI follows the component library and design tokens |
| Accessibility standards | Automated and manual checks against the relevant accessibility requirements |
Define what the agent may do and enforce the access and execution limits needed for the task.
| Element | Description | Example |
|---|---|---|
| Target | Name the files or systems the agent may use and restrict its permissions accordingly | src/api/users/, payments_table |
| Non-goals | What the agent must NOT change | "Do not modify authentication logic" |
| Acceptance criteria | Concrete definition of "done" | "All tests pass, endpoint returns 200 with valid payload" |
| Allowed dependencies | What the agent may import or call | "No new external packages without approval" |
| Max iterations | Upper bound on agent execution cycles | 20 tool calls, 10 minutes wall time |
The Operating Loop
The Plan-Execute-Verify-Ship-Learn Cycle
Use this loop in an existing task or review. Agree the work, carry it out, check it, decide whether to use it and learn from the result. In software work, automated checks may run in continuous integration (CI) before a pull request (PR) is merged.
Agree the task, its limits and how the result will be checked.
- Product defines: Goal, acceptance criteria, UX requirements
- Engineering defines: Scope, non-goals, risk level, constraints, verification method
- Record who reviews the result, who decides whether it is ready and when to stop
- Match tool access and execution limits to the plan; instructions alone do not enforce them
Agents carry out the agreed task with configured access and execution limits.
- Generate code, tests, documentation, or refactors
- Call tools as needed (data retrieval, API interactions, code execution)
- Iterate within the defined scope (run tests, fix failures, retry)
- Operate within configured iteration limits
- People monitor progress and intervene when the task reaches a limit or needs a decision
Check the work with appropriate automated checks and human review.
- Automated: CI pipeline, static analysis, security scanning, coverage thresholds, policy checks
- Human — Engineering: code review, architecture alignment, edge case consideration
- Human — Product: acceptance review, UX review, copy review, accessibility check
- Review depth scales with risk level
Decide whether to release the work and confirm how to recover from a problem.
- Audit trail: what was generated, by which agent, reviewed by whom
- Recovery: agree a rollback or recovery procedure, including how to handle irreversible changes
- Post-merge monitoring: watch for anomalies in error rates and latency
- Diff review: ensure merged code matches what was reviewed
Use what happened to improve the next task, check or handoff.
- Update instructions when a correction reveals a useful general lesson
- Refine task templates: tighten plans that were ambiguous
- Update evaluation criteria: add checks the verify step missed
- Share useful examples and changes with the team
- Remove instructions or steps that no longer help
Task Classification Matrix
Discuss two questions: boundedness means how clearly the task is defined; risk means what could happen if it goes wrong. The examples below are starting points. Access, data sensitivity and the limits of your checks can change the classification.
Automated verification. Sampling review.
Engineering
- Local API prototypes using synthetic data
- Code formatting, linting, and style fixes
- Documentation and changelog generation
Product
- Copy and microcopy generation within brand guidelines
- Test case generation from acceptance criteria
- Competitive analysis summaries from public data
Automated + human verification.
Engineering
- Frontend component generation and cleanup
- Test generation for existing business logic
- Migration scripts tried on disposable test data
Product
- PRD drafts from user research notes
- User story decomposition from high-level requirements
- Design-to-code translation using design system components
Human-led with agent drafts. Full review.
Engineering
- Security-adjacent feature implementation
- Payment flow modifications
- Data model changes with migration
Product
- User-facing copy with legal/compliance implications
- Onboarding flow changes affecting activation metrics
With human plan review.
Engineering
- Exploratory refactors with clear goals
- Performance optimization within defined bounds
Product
- Feature spec elaboration from brief outline
- Design exploration within existing system
Human review mandatory.
Engineering
- Cross-service integration work
- Complex business logic implementation
Product
- Multi-step workflow redesign
- New feature prototyping within constraints
Agent may draft, human designs and reviews.
Engineering
- Auth system modifications
- Infrastructure security hardening
Product
- Pricing model implementation
- Compliance-critical workflow changes
With agent support for research/exploration.
Engineering
- Technology evaluation and prototyping
- Architecture documentation drafts
Product
- Market research synthesis
- Competitive landscape analysis
Agent provides options, human decides.
Engineering
- Novel architecture decisions
- Performance-sensitive distributed systems
Product
- Product strategy options analysis
- User research synthesis and insight generation
People own these decisions. Scope any agent research or drafting as a separate task with appropriate access and review.
Engineering
- Approving changes during a production incident
- Accepting security risk after an investigation
Product
- Final product strategy and roadmap decisions
- Approving customer-facing claims and brand commitments
- Approving pricing and packaging commitments
The Human-Agent Boundary
Ask: “If the agent gets this wrong, what happens, and how will we notice?”
- For a bounded, reversible task, an agent may execute within agreed permissions and checks.
- If a mistake could affect a customer, agree the human review and approval needed before use.
- For security-sensitive or production changes, a person owns the decision and explicitly limits any agent assistance.
- If the consequences are unclear, investigate them before delegating.
A passing check only covers what it tests. Revisit the boundary when the task, tools or evidence change; you may need more human involvement as well as less.
The role guides discuss responsibilities and learning. Use the Role skills explorer to compare the skills in those guides, or choose a shared-skill practice guide for your role.
Maturity Model
These five levels describe patterns of agent use, drawing on Eledath’s “Levels of Agentic Engineering.” Use them to discuss the checks, skills and support a workflow needs. They do not establish a team’s performance or require every team to move to a higher number.
The Leadership Guide has a short version of these levels for team discussions.
Assisted
AI provides suggestions that developers accept, modify, or reject. The developer drives all decisions and execution.
Dimensions
Signs to discuss
- Developers use AI for suggestions but control all execution
- No structured prompting or context engineering
- No shared rules or templates for AI usage
- AI usage is individual, not team-standardized
Structured
AI operates within structured contexts. Teams use dedicated AI IDEs, maintain rules files, and follow defined prompting patterns.
Dimensions
Signs to discuss
- Team uses AI IDEs with structured context (rules files, project-level instructions)
- Shared task templates exist for common operations
- All AI-generated output goes through standard code review
- Team has basic conventions for when and how to use AI tools
Integrated
Agents use automated feedback during development. CI checks support human review of the result.
Dimensions
Signs to discuss
- Agents iterate based on CI/test feedback without human intervention in the loop
- Useful lessons from completed work inform instructions, templates and checks
- Evaluation coverage is explicitly tracked and improving
- Guardrail stack (all 5 layers) is operational
- Team measures agentic adoption KPIs
Autonomous
Agents operate in the background, working on tasks asynchronously. Humans define tasks and review results.
Dimensions
Signs to discuss
- Agents produce PRs asynchronously (not just in interactive sessions)
- Human review happens after completion, not during execution
- Cost tracking and budgeting is active per agent and per team
- Execution records show relevant tool calls, outputs and failures for review
- Incident response protocol exists for agent-caused failures
Orchestrated
Multiple agents coordinate in parallel, managed by orchestration systems.
Dimensions
Signs to discuss
- An orchestrator coordinates multiple agents working on related tasks
- Agents are shared and reused across teams (internal skills marketplace)
- Fleet-level observability tracks all agents, costs, and outcomes in one view
- Governance framework (registry, access control, audit) is fully operational
- The organization can articulate and enforce policies across the agent fleet
Look at the actual workflow and handoffs. One person’s tool adoption is not a formula for team capability, and a well-controlled simple workflow may be the right choice.
Further Reading
- Building Effective Agents — Anthropic
- A Practical Guide to Building Agents — OpenAI
- AI Agents Whitepaper — Google
- Cloud Adoption Framework — Microsoft
- Agentic Engineering for Software Teams — vibecoding.app
- The 8 Levels of Agentic Engineering — Eledath
- Top Strategic Technology Trends 2026 — Gartner
- Five Stages of Agentic Evolution — Gartner