Skip to content

Leadership Guide

Help your team work well with AI agents.

HELM 1.0.1 · Updated

Table of contents

Organizational Model

The Structural Shift

When work passes between specialist teams, responsibility can become unclear. Before adding agents, look at one workflow and agree who owns the result and who makes each decision.

One option is a small cross-functional team: people with different skills who share responsibility for a feature from start to finish. The examples below call these teams “pods.” Consider whether this arrangement would help your work before changing your team structure.

Example (Specialist handoffs)
Product
Frontend
Backend
QA
DevOps

Check where handoffs help or delay the work

One option (Cross-functional teams)
Pod A
PMDesignerArchitectEngineersAgent
Feature A
Pod B
PMDesignerArchitectEngineersAgent
Feature B
Pod C
PMDesignerArchitectEngineersAgent
Feature C
01

Workflows: Simplify before adding agents

Start with the result you need. Remove unnecessary steps, then decide where an agent could help. A small change may be enough.

02

Leadership: Agree goals, limits and decisions

Make the expected result, constraints and decision owner clear. Include detailed steps when the task needs them, and give the team room to use its expertise.

03

Skills: Connect expertise across the work

Help people understand the parts of a task they need to review. Involve specialists when a change reaches beyond that knowledge.

04

Learning: Review and improve together

Review working practices when tasks, tools or results change. Share useful lessons and remove instructions that no longer help.

05

Structure: Make handoffs and ownership clear

Agree who owns the result across product, engineering and quality work. Try a different team arrangement when an observed handoff or responsibility problem warrants it.

06

Measurement: Review outcomes, quality and effort

Use delivery counts to understand activity, alongside evidence about user outcomes, quality and effort. Avoid treating agent use or output volume as a measure of a person’s ability.

In this example, a cross-functional team owns a feature through delivery, with agents helping on bounded tasks. Review whether that arrangement would improve your current handoffs.

Roles with Explicit Authority

These nine roles describe responsibilities to assign. The people covering them need authority, time and support. One person may cover several responsibilities; a separate position is a decision about workload and risk. Model hosting and inference responsibilities apply when the product needs that infrastructure.

Engineering Roles

AI Architect

Owns: End-to-end orchestration and structural decisions.

Responsibilities:

  • Selects models and defines which model handles which task
  • Designs data flow from input to output
  • Decides orchestration pattern (single agent, multi-agent, workflow)
  • Defines failure modes and recovery paths
  • Makes the structural decisions the rest of the team builds on

See the Staff / Principal Engineer guide for examples and skills to discuss.

AI Reliability Engineer

Owns: Observability, cost measurement, and failure recovery. The SRE equivalent for AI systems.

Responsibilities:

  • Defines what to measure to know the system works
  • Monitors cost per execution and flags unsustainable patterns
  • Manages failure detection and recovery mechanisms
  • Owns the guardrail stack implementation and enforcement
  • Runs incident response for agent-related failures

See the SRE / DevOps Engineer guide for examples and skills to discuss.

Evaluation Lead

Owns: The approach to checking agent outputs, including criteria, examples and the limits of automated checks.

Responsibilities:

  • Defines “how do we know this is good enough to ship?”
  • Designs eval suites for agent behavior (beyond standard test suites)
  • Sets passing thresholds and quality bars
  • Ensures evaluation runs before every ship decision
  • Tracks quality metrics over time to detect drift

Agree how output evaluation and product testing are covered in your team. See the QA Engineer / SDET guide for examples of both.

Product Engineer

Owns: Feature velocity and integration.

Responsibilities:

  • Runs the agent execution loop for scoped delivery tasks
  • Creates and maintains task templates and agent instructions
  • Integrates agent-generated output into the product
  • Ensures agent output meets product requirements and UX standards
  • Manages the Plan-Execute-Verify-Ship-Learn loop (see Practitioner Guide)

See the Software Engineer guide for examples and skills to discuss.

Platform Engineer

Owns: Shared execution infrastructure, including model hosting and inference when the product requires them.

Responsibilities:

  • Manages compute infrastructure for agent execution
  • Optimizes cost-efficiency at the infrastructure layer
  • Handles latency and reliability of model inference
  • Implements the governance layer (registry, access control, observability)
  • Manages secrets, API keys, and secure agent-to-system connectivity

See the Platform Engineer guide for examples and skills to discuss.

Engineering Manager

Owns: Team capability, learning support and review of delivery results.

Responsibilities:

  • Uses the Maturity Model to discuss working practices and the checks they need
  • Finds workflow bottlenecks and helps people build relevant skills
  • Defines and enforces the team’s operating rhythm around the Plan-Execute-Verify-Ship-Learn loop
  • Reviews handoffs and responsibilities before changing team structure
  • Reviews user outcomes, quality and effort alongside delivery volume
  • Coaches engineers on judgment, review quality, and context engineering
  • Discusses workload, concerns and learning needs with the people affected

See the Engineering Manager guide for examples and skills to discuss.

Product Roles

Product Manager

Owns: Problem definition, acceptance criteria, and product quality.

Responsibilities:

  • Defines what to build and why (the “Plan” phase of the operating loop)
  • Writes acceptance criteria that agents can execute against
  • Reviews agent output for product correctness (does it solve the user’s problem?)
  • Chooses the detail a task needs, including constraints and examples of acceptable results
  • Tracks product outcome metrics alongside delivery metrics

See the Product Manager guide for examples and skills to discuss.

Product Designer

Owns: UX quality, design system, and interaction patterns.

Responsibilities:

  • Maintains the design system that agents generate from (tokens, components, patterns)
  • Reviews agent-generated UI for UX quality and design consistency
  • Defines design tokens and component specifications as agent instructions
  • Designs interactions, reviews generated interfaces and improves shared design guidance
  • Addresses the emerging discipline sometimes called Agent Experience (AX): designing for both human and agent actors

See the Product Designer guide for examples and skills to discuss.

QA Engineer

Owns: Product-level quality from the user’s perspective. Distinct from the Evaluation Lead.

Responsibilities:

  • Translates acceptance criteria into testable assertions
  • Builds evaluation suites that validate product behavior, not just code correctness
  • Monitors quality drift from a user-facing perspective (UX regressions, accessibility, copy errors)
  • Works with the Evaluation Lead on comprehensive quality coverage

Working with the Evaluation Lead: Check whether the agent output meets its instructions and technical requirements, then whether the resulting product works for users. Agree the checks and handoffs so both responsibilities are covered.

Use the role guides to discuss responsibilities, relevant experience and learning needs.

Scaling Path

The team sizes below are illustrations, not staffing requirements. Start with the responsibilities your work needs and review how much time and expertise each takes.

Example: smaller team (7–8 people):

  • 1 AI Architect (leads)
  • 2-3 Product Engineers
  • 1 AI Reliability Engineer (may share duties with the Architect early on)
  • 1 Engineering Manager
  • 1 Product Manager (may be part-time)
  • 1 Product Designer (may be shared across pods early on)

Example: larger team (11–16 people):

  • 1 AI Architect
  • 3-4 Product Engineers (some specializing in different surfaces)
  • 1-2 AI Reliability Engineers
  • 1-2 Evaluation Leads
  • 1 Platform Engineer
  • 1-2 Engineering Managers (one per pod at scale)
  • 1 Product Manager
  • 1-2 Product Designers
  • 1 QA Engineer

Give each responsibility a named owner with enough time and authority to do the work. If one person covers several areas, check that the workload is realistic and that review remains effective.

Decision Rights Matrix

Use the roles below as starting points for assigning decisions. Name one accountable decision owner, the people to consult and any required approvals. Where a row lists several roles, resolve that choice before using the table.

DecisionRoles to involveFurther inputWhat to consider
Model selectionAI Architect + Evaluation LeadProduct EngineerTechnical fit + eval data must align
Orchestration patternAI ArchitectTeamTask structure, integration, cost and failure handling
Cost controlAI Reliability EngineerAI ArchitectToken spend, compute budgets, cost alerts
Eval thresholds (“can we ship?”)Evaluation LeadProduct, ArchitectAgree the checks before reviewing the output
Feature prioritizationProduct Manager + AI ArchitectTeamArchitect says what’s feasible, PM decides what matters
What to build (problem selection)Product ManagerArchitect, TeamPM owns problem definition; engineering owns solution
Acceptance criteriaProduct ManagerEngineering, DesignCriteria must be agent-executable; PM defines, engineering validates feasibility
UX quality standardsProduct DesignerPM, EngineeringDesign system compliance, accessibility, interaction quality
Design system changesProduct DesignerPM, AI ArchitectComponents and tokens agents generate from
Product-level quality thresholdsQA EngineerPM, Evaluation LeadUser-facing quality distinct from agent output correctness
Architecture decisionsAI ArchitectTeamPeople approve architecture and accept its consequences
Security decisionsAI Architect + Reliability EngTeamPeople approve changes and accept security risk
Release decisionsProduct Manager + AI ArchitectReliability EngHuman judgment on production readiness

Involve the people affected and make the final decision owner clear. Record how to raise concerns and resolve disagreements.


Maturity Model

The full five-level maturity model (dimension tables, assessment criteria, prerequisites, and failure modes for each level) lives in the Practitioner Guide. Both guides share this model so practitioners and leaders use a common vocabulary.

Level 1

Assisted

AI provides suggestions that developers accept, modify, or reject. The developer drives all decisions and execution.

"We use Copilot for suggestions"
Read the full workflow descriptions and their limits

Use the levels to discuss workflows, not to score people. One person’s tool adoption does not determine the team’s capability. Check the actual handoffs, review capacity and support needed.

Use what you observe to choose a next step in the Adoption Roadmap and relevant questions in Measurement.


Adoption Roadmap

This four-phase example uses a 180-day schedule to organize a trial and possible expansion. Adapt its dates and activities to your team. Expand when the work and checks justify it; a calendar date or task count does not prove readiness.

Contained Pilot — Days 1–30

Try one workflow and record what changes for delivery and the people doing the work.
Suggested activities
  • Select one repository with moderate complexity
  • Choose a small set of repeatable tasks, such as a local prototype, test or refactor
  • Measure baseline metrics: PR cycle time, change failure rate, test coverage, bug rate
  • Agree scope, restrict access, protect sensitive data and name a reviewer for agent changes
  • Give participants a chance to try relevant tasks with support
  • Document what works, what fails, and what surprises

ProductPM participates in defining task types. Designer reviews agent-generated UI. Baseline product metrics recorded.

Before expanding
  • Baseline metrics recorded for comparison
  • Representative tasks reviewed, with enough evidence to explain the next decision
  • No critical quality incidents from agent output
  • Team can articulate which tasks agents handle well and which they don't
  • Basic rules file created and shared across the team
Checks and precautions
  • A named, capable person reviews each agent change during the pilot
  • Start with low-risk, well-bounded tasks only
  • If a quality incident occurs, pause, investigate and agree what must change before continuing

Measurement and Failure Modes

KPI Dashboard

Choose key performance indicators (KPIs) that help answer a decision your team faces. Review user outcomes, quality and total effort together. The suggested directions below need context; none is a score for an individual.

Lead time

Discuss decrease

Time from issue opened to code merged

PR review time

Discuss decrease

Time from PR opened to approved

Change failure rate

Check stability

% of deployments causing incidents or rollbacks

Rollback frequency

Check stability

Number of rollbacks per deployment period

Escaped defects

Discuss decrease

Bugs found in production per sprint

Test coverage delta

Monitor

Change in test coverage over time

Deployment frequency

Discuss increase

How often the team deploys to production

Read these measures together
Lead time↓ANDDeployment frequency↑ANDChange failure rate↔ANDEscaped defects↓

This pattern may suggest improvement. Check user outcomes and total effort before drawing a conclusion.

Six Failure Modes

These examples describe symptoms, possible explanations and actions to try. Investigate the cause in your own work before choosing a response.

The selected tasks may not address an important problem, or the benefit may be offset by review and rework.

  • Tie every agent workflow to a measurable delivery KPI
  • Require a "so what?" test: if automated, what bottleneck does it remove?
  • Review task selection criteria quarterly

Agent output velocity exceeds the team's review capacity. Often caused by large, unfocused agent PRs.

  • Enforce smaller PR scope (one concern per PR, bounded by task template)
  • Tighten acceptance criteria so PRs are more focused
  • Scale review capacity: train more team members
  • Implement risk-based review: low-risk PRs get sampling-based review
  • Consider review automation for mechanical aspects

The checks may be missing important cases. Compare the reported problems with what is actually tested.

  • Expand evaluation coverage beyond unit tests (integration tests, performance benchmarks, architecture fitness functions)
  • Track post-release defect rate specifically for agent-generated code
  • Implement regular "agent output audits"
  • Monitor change failure rate as an early warning signal

Effective agent interaction patterns are not captured and shared. Knowledge stays in individual heads.

  • Convert individual prompts into shared task templates
  • Maintain team-level rules files (not personal ones)
  • Publish an internal "agentic SOP" with examples
  • Pair programming sessions where skilled users demonstrate approach
  • Use the Learn phase to share useful findings and remove unhelpful instructions

Shared records, ownership or controls may not have kept up with the work. Check where visibility or responsibility is missing.

  • Implement governance layer before cross-team scaling
  • Start with clear ownership, appropriate access, useful activity records and cost tracking
  • Review access controls and audit records before adding more workflows or teams
  • Assign a governance owner (AI Reliability or Platform Engineer)
  • Review governance completeness quarterly

The work may not address the intended user problem. Check the evidence behind the priorities and what happened after release.

  • Tie agent task selection to product outcome metrics
  • Require PM sign-off on every task plan
  • Measure feature adoption and user satisfaction alongside delivery speed
  • Apply "redesign, don't automate" to product discovery, not just delivery

Further Reading