Skip to content

Practitioner Guide

How we build and operate with AI agents day-to-day.

HELM 1.0.1 · Updated

Table of contents

For day-to-day agent-assisted work, start with the operating loop. If you are designing an agent workflow or product, the architecture patterns below help you compare options. Choose the detail relevant to your task.

Agent Architecture Patterns

Agent Anatomy

These five parts help explain how an agent works.

ComponentWhat It DoesExample
ModelA language model (LLM) generates outputs and may choose which tools to call.Drafting an answer or selecting a tool
ToolsExternal functions, APIs, or systems the agent can invoke.Database queries, web search, code execution, file I/O
InstructionsExplicit guidelines, scope constraints, and behavioral rules.System prompts, AGENTS.md, rules files, policy documents
MemoryInformation retained across interactions, with access and retention limits.Conversation history, task state, saved project context
RetrievalLooks up information the task needs.Document search, a knowledge base, project files

Products combine these parts differently. Check which information and tools your agent actually uses.

Workflows vs. Agents

Anthropic offers a useful distinction between workflows and agents:

DimensionWorkflowsAgents
ControlFollows steps defined in advanceChooses steps as the task proceeds
PredictabilityThe path is defined; results can still varyBoth the path and results can vary
FlexibilityLimited to the paths you defineCan choose among permitted actions
Best forWell-defined, repeatable tasksOpen-ended problems with unpredictable steps
Cost/latencyMeasure calls and time for the workflowSet limits and measure time and cost
Error handlingDefined checks and recovery stepsFeedback, stopping conditions and escalation

Use a workflow when you can define the steps in advance. Consider an agent when the next step depends on what it finds. In either case, agree how to check the result and when to stop.

Tool Types

These three tool categories help you plan access and checks:

Data tools retrieve context and information. They can expose sensitive data even when they cannot change it. Examples: Query databases, read documents, search the web, pull CRM records.

Action tools change something in a system. Consider who is affected and how a mistake could be corrected. Examples: Send emails, update records, create tickets, issue refunds, deploy code.

Orchestration tools let an agent ask another agent to carry out part of the work. Examples: A “research agent” callable by a “manager agent,” a specialist agent invoked by a triage agent.

Assess each tool’s risk by what it can access, who could be affected, and whether a mistake can be undone. Use that assessment to choose the limits and checks it needs.

Composition Patterns

These eight patterns are options to compare. Start with the simplest approach that meets the task’s needs; additional agents or steps should address a problem you have observed.

Pattern 1

Prompt Chaining

Split a task into a fixed sequence of steps. Each model call uses the previous result, with checks between steps.

When to consider it

The task has clear steps. Compare whether the extra calls improve quality enough to justify their time and cost.

Example

Generate marketing copy, then translate it. Write an outline, validate it against criteria, then write the full document.

Pattern Selection Guide

Use the problem you have observed to choose an option. Compare it with the simplest approach that could meet the task’s needs.

StartSingle LLM call with good prompting
Output quality insufficient
ConsiderPrompt Chaining or Evaluator-Optimizer
StartPrompt Chaining
Task decomposition isn't fixed
ConsiderSingle Agent Loop
StartSingle Agent Loop
Tool selection or overlapping tasks cause repeated errors
ConsiderManager or Orchestrator-Workers
StartSingle Agent Loop
Distinct categories with different handling
ConsiderRouting + specialized agents
StartManager pattern
Central agent bottlenecks; specialists need full autonomy
ConsiderDecentralized Handoff

Keep it simple: Add another agent only when it addresses a problem the simpler approach cannot handle well.


The Guardrail Stack

The five layers help you review scope, quality, policy, human decisions and shared oversight. Consider each layer before expanding the work. The controls you need depend on the task and its consequences. See Principle 4: Guardrails Are Non-Negotiable.

Fleet-level controls for managing agents at organizational scale.

ElementDescription
RegistrySingle source of truth tracking all agents, their capabilities, owners, and status
Access controlRole-based permissions determining which agents can access which systems and data
ObservabilityUnified monitoring across all agents — execution traces, cost tracking, error rates, latency
InteroperabilityStandards for agents to work across platforms and teams (e.g., Model Context Protocol)
Audit trailRecord relevant actions, approvals and outcomes, with appropriate access and retention
Cost budgetingPer-agent and per-team token/compute budgets with alerts and hard limits

Ensure humans retain authority over decisions that agents must not make autonomously.

ElementDescription
ArchitectureSystem design, technology choices, data model changes
Risk acceptanceShipping known tradeoffs, accepting technical debt
Release timingWhen code goes to production
Incident responseRollback decisions, postmortem actions
Security-critical changesAuthentication, authorization, encryption
Cost commitmentsActions with financial impact above defined thresholds

Apply the privacy, security and organizational controls the work needs.

ElementDescriptionExample
No secret exposureAutomated secret scanning in pre-commit and CICredentials leaking into repositories
PII filteringLimit access to personal information and check outputs for unintended disclosurePrivacy violations in generated content
Safety classificationTreat untrusted input carefully; detection can help but may miss attempts to redirect an agentSystem exploitation
Relevance classificationFlag off-topic or out-of-scope agent behaviorScope drift and waste
ModerationContent safety checks on agent outputsHarmful or inappropriate generated content
Dependency policyBlock unsafe dependency upgrades or additionsSupply chain attacks
Branch policyNo direct pushes to main/protected branchesUnreviewed code reaching production

Use automated checks and human review to find problems before the work is used.

ElementDescription
Formatting & lintingEnforce style consistency (Black, ESLint, Prettier, etc.)
Type checkingStatic type verification (mypy, TypeScript strict mode)
Unit & integration testsRun relevant tests and add meaningful checks for the change
Static analysisSecurity scanning, dependency vulnerability checks
Coverage thresholdsReview coverage alongside the importance of the cases being tested
Design system complianceAgent-generated UI follows the component library and design tokens
Accessibility standardsAutomated and manual checks against the relevant accessibility requirements

Define what the agent may do and enforce the access and execution limits needed for the task.

ElementDescriptionExample
TargetName the files or systems the agent may use and restrict its permissions accordinglysrc/api/users/, payments_table
Non-goalsWhat the agent must NOT change"Do not modify authentication logic"
Acceptance criteriaConcrete definition of "done""All tests pass, endpoint returns 200 with valid payload"
Allowed dependenciesWhat the agent may import or call"No new external packages without approval"
Max iterationsUpper bound on agent execution cycles20 tool calls, 10 minutes wall time

The Operating Loop

The Plan-Execute-Verify-Ship-Learn Cycle

Use this loop in an existing task or review. Agree the work, carry it out, check it, decide whether to use it and learn from the result. In software work, automated checks may run in continuous integration (CI) before a pull request (PR) is merged.

PlanExecuteVerifyShipLearn
PlanProduct + Engineering

Agree the task, its limits and how the result will be checked.

  • Product defines: Goal, acceptance criteria, UX requirements
  • Engineering defines: Scope, non-goals, risk level, constraints, verification method
  • Record who reviews the result, who decides whether it is ready and when to stop
  • Match tool access and execution limits to the plan; instructions alone do not enforce them

Task Classification Matrix

Discuss two questions: boundedness means how clearly the task is defined; risk means what could happen if it goes wrong. The examples below are starting points. Access, data sensitivity and the limits of your checks can change the classification.

Low Risk
Medium Risk
High Risk
Well-bounded
Semi-bounded
Open-ended
Agent-drivenWell-bounded / Low Risk

Automated verification. Sampling review.

Engineering
  • Local API prototypes using synthetic data
  • Code formatting, linting, and style fixes
  • Documentation and changelog generation
Product
  • Copy and microcopy generation within brand guidelines
  • Test case generation from acceptance criteria
  • Competitive analysis summaries from public data

The Human-Agent Boundary

Ask: “If the agent gets this wrong, what happens, and how will we notice?”

  • For a bounded, reversible task, an agent may execute within agreed permissions and checks.
  • If a mistake could affect a customer, agree the human review and approval needed before use.
  • For security-sensitive or production changes, a person owns the decision and explicitly limits any agent assistance.
  • If the consequences are unclear, investigate them before delegating.

A passing check only covers what it tests. Revisit the boundary when the task, tools or evidence change; you may need more human involvement as well as less.

The role guides discuss responsibilities and learning. Use the Role skills explorer to compare the skills in those guides, or choose a shared-skill practice guide for your role.


Maturity Model

These five levels describe patterns of agent use, drawing on Eledath’s “Levels of Agentic Engineering.” Use them to discuss the checks, skills and support a workflow needs. They do not establish a team’s performance or require every team to move to a higher number.

The Leadership Guide has a short version of these levels for team discussions.

Level 1

Assisted

AI provides suggestions that developers accept, modify, or reject. The developer drives all decisions and execution.

"We use Copilot for suggestions"
Dimensions
CapabilitiesTab completion, inline suggestions, single-turn Q&A, code explanation
ToolsGitHub Copilot, ChatGPT, basic AI-assisted IDE features
Human roleFull control. AI is a passive tool.
Agent autonomyNone. Every output requires explicit human action.
Product dimensionPMs and designers use AI for ad-hoc tasks (drafting docs, brainstorming). No integration with engineering workflows.
Risk profileDepends on the task and data. Reviewing each suggestion can still miss errors.
Signs to discuss
  • Developers use AI for suggestions but control all execution
  • No structured prompting or context engineering
  • No shared rules or templates for AI usage
  • AI usage is individual, not team-standardized
Failure modeAccepting suggestions without checking or understanding them.

Look at the actual workflow and handoffs. One person’s tool adoption is not a formula for team capability, and a well-controlled simple workflow may be the right choice.


Further Reading