Organizational Model
The Structural Shift
When work passes between specialist teams, responsibility can become unclear. Before adding agents, look at one workflow and agree who owns the result and who makes each decision.
One option is a small cross-functional team: people with different skills who share responsibility for a feature from start to finish. The examples below call these teams “pods.” Consider whether this arrangement would help your work before changing your team structure.
Check where handoffs help or delay the work
Workflows: Simplify before adding agents
Start with the result you need. Remove unnecessary steps, then decide where an agent could help. A small change may be enough.
Leadership: Agree goals, limits and decisions
Make the expected result, constraints and decision owner clear. Include detailed steps when the task needs them, and give the team room to use its expertise.
Skills: Connect expertise across the work
Help people understand the parts of a task they need to review. Involve specialists when a change reaches beyond that knowledge.
Learning: Review and improve together
Review working practices when tasks, tools or results change. Share useful lessons and remove instructions that no longer help.
Structure: Make handoffs and ownership clear
Agree who owns the result across product, engineering and quality work. Try a different team arrangement when an observed handoff or responsibility problem warrants it.
Measurement: Review outcomes, quality and effort
Use delivery counts to understand activity, alongside evidence about user outcomes, quality and effort. Avoid treating agent use or output volume as a measure of a person’s ability.
In this example, a cross-functional team owns a feature through delivery, with agents helping on bounded tasks. Review whether that arrangement would improve your current handoffs.
Roles with Explicit Authority
These nine roles describe responsibilities to assign. The people covering them need authority, time and support. One person may cover several responsibilities; a separate position is a decision about workload and risk. Model hosting and inference responsibilities apply when the product needs that infrastructure.
Engineering Roles
AI Architect
Owns: End-to-end orchestration and structural decisions.
Responsibilities:
- Selects models and defines which model handles which task
- Designs data flow from input to output
- Decides orchestration pattern (single agent, multi-agent, workflow)
- Defines failure modes and recovery paths
- Makes the structural decisions the rest of the team builds on
See the Staff / Principal Engineer guide for examples and skills to discuss.
AI Reliability Engineer
Owns: Observability, cost measurement, and failure recovery. The SRE equivalent for AI systems.
Responsibilities:
- Defines what to measure to know the system works
- Monitors cost per execution and flags unsustainable patterns
- Manages failure detection and recovery mechanisms
- Owns the guardrail stack implementation and enforcement
- Runs incident response for agent-related failures
See the SRE / DevOps Engineer guide for examples and skills to discuss.
Evaluation Lead
Owns: The approach to checking agent outputs, including criteria, examples and the limits of automated checks.
Responsibilities:
- Defines “how do we know this is good enough to ship?”
- Designs eval suites for agent behavior (beyond standard test suites)
- Sets passing thresholds and quality bars
- Ensures evaluation runs before every ship decision
- Tracks quality metrics over time to detect drift
Agree how output evaluation and product testing are covered in your team. See the QA Engineer / SDET guide for examples of both.
Product Engineer
Owns: Feature velocity and integration.
Responsibilities:
- Runs the agent execution loop for scoped delivery tasks
- Creates and maintains task templates and agent instructions
- Integrates agent-generated output into the product
- Ensures agent output meets product requirements and UX standards
- Manages the Plan-Execute-Verify-Ship-Learn loop (see Practitioner Guide)
See the Software Engineer guide for examples and skills to discuss.
Platform Engineer
Owns: Shared execution infrastructure, including model hosting and inference when the product requires them.
Responsibilities:
- Manages compute infrastructure for agent execution
- Optimizes cost-efficiency at the infrastructure layer
- Handles latency and reliability of model inference
- Implements the governance layer (registry, access control, observability)
- Manages secrets, API keys, and secure agent-to-system connectivity
See the Platform Engineer guide for examples and skills to discuss.
Engineering Manager
Owns: Team capability, learning support and review of delivery results.
Responsibilities:
- Uses the Maturity Model to discuss working practices and the checks they need
- Finds workflow bottlenecks and helps people build relevant skills
- Defines and enforces the team’s operating rhythm around the Plan-Execute-Verify-Ship-Learn loop
- Reviews handoffs and responsibilities before changing team structure
- Reviews user outcomes, quality and effort alongside delivery volume
- Coaches engineers on judgment, review quality, and context engineering
- Discusses workload, concerns and learning needs with the people affected
See the Engineering Manager guide for examples and skills to discuss.
Product Roles
Product Manager
Owns: Problem definition, acceptance criteria, and product quality.
Responsibilities:
- Defines what to build and why (the “Plan” phase of the operating loop)
- Writes acceptance criteria that agents can execute against
- Reviews agent output for product correctness (does it solve the user’s problem?)
- Chooses the detail a task needs, including constraints and examples of acceptable results
- Tracks product outcome metrics alongside delivery metrics
See the Product Manager guide for examples and skills to discuss.
Product Designer
Owns: UX quality, design system, and interaction patterns.
Responsibilities:
- Maintains the design system that agents generate from (tokens, components, patterns)
- Reviews agent-generated UI for UX quality and design consistency
- Defines design tokens and component specifications as agent instructions
- Designs interactions, reviews generated interfaces and improves shared design guidance
- Addresses the emerging discipline sometimes called Agent Experience (AX): designing for both human and agent actors
See the Product Designer guide for examples and skills to discuss.
QA Engineer
Owns: Product-level quality from the user’s perspective. Distinct from the Evaluation Lead.
Responsibilities:
- Translates acceptance criteria into testable assertions
- Builds evaluation suites that validate product behavior, not just code correctness
- Monitors quality drift from a user-facing perspective (UX regressions, accessibility, copy errors)
- Works with the Evaluation Lead on comprehensive quality coverage
Working with the Evaluation Lead: Check whether the agent output meets its instructions and technical requirements, then whether the resulting product works for users. Agree the checks and handoffs so both responsibilities are covered.
Use the role guides to discuss responsibilities, relevant experience and learning needs.
Scaling Path
The team sizes below are illustrations, not staffing requirements. Start with the responsibilities your work needs and review how much time and expertise each takes.
Example: smaller team (7–8 people):
- 1 AI Architect (leads)
- 2-3 Product Engineers
- 1 AI Reliability Engineer (may share duties with the Architect early on)
- 1 Engineering Manager
- 1 Product Manager (may be part-time)
- 1 Product Designer (may be shared across pods early on)
Example: larger team (11–16 people):
- 1 AI Architect
- 3-4 Product Engineers (some specializing in different surfaces)
- 1-2 AI Reliability Engineers
- 1-2 Evaluation Leads
- 1 Platform Engineer
- 1-2 Engineering Managers (one per pod at scale)
- 1 Product Manager
- 1-2 Product Designers
- 1 QA Engineer
Give each responsibility a named owner with enough time and authority to do the work. If one person covers several areas, check that the workload is realistic and that review remains effective.
Decision Rights Matrix
Use the roles below as starting points for assigning decisions. Name one accountable decision owner, the people to consult and any required approvals. Where a row lists several roles, resolve that choice before using the table.
| Decision | Roles to involve | Further input | What to consider |
|---|---|---|---|
| Model selection | AI Architect + Evaluation Lead | Product Engineer | Technical fit + eval data must align |
| Orchestration pattern | AI Architect | Team | Task structure, integration, cost and failure handling |
| Cost control | AI Reliability Engineer | AI Architect | Token spend, compute budgets, cost alerts |
| Eval thresholds (“can we ship?”) | Evaluation Lead | Product, Architect | Agree the checks before reviewing the output |
| Feature prioritization | Product Manager + AI Architect | Team | Architect says what’s feasible, PM decides what matters |
| What to build (problem selection) | Product Manager | Architect, Team | PM owns problem definition; engineering owns solution |
| Acceptance criteria | Product Manager | Engineering, Design | Criteria must be agent-executable; PM defines, engineering validates feasibility |
| UX quality standards | Product Designer | PM, Engineering | Design system compliance, accessibility, interaction quality |
| Design system changes | Product Designer | PM, AI Architect | Components and tokens agents generate from |
| Product-level quality thresholds | QA Engineer | PM, Evaluation Lead | User-facing quality distinct from agent output correctness |
| Architecture decisions | AI Architect | Team | People approve architecture and accept its consequences |
| Security decisions | AI Architect + Reliability Eng | Team | People approve changes and accept security risk |
| Release decisions | Product Manager + AI Architect | Reliability Eng | Human judgment on production readiness |
Involve the people affected and make the final decision owner clear. Record how to raise concerns and resolve disagreements.
Maturity Model
The full five-level maturity model (dimension tables, assessment criteria, prerequisites, and failure modes for each level) lives in the Practitioner Guide. Both guides share this model so practitioners and leaders use a common vocabulary.
Assisted
AI provides suggestions that developers accept, modify, or reject. The developer drives all decisions and execution.
Structured
AI operates within structured contexts. Teams use dedicated AI IDEs, maintain rules files, and follow defined prompting patterns.
Integrated
Agents use automated feedback during development. CI checks support human review of the result.
Autonomous
Agents operate in the background, working on tasks asynchronously. Humans define tasks and review results.
Orchestrated
Multiple agents coordinate in parallel, managed by orchestration systems.
Use the levels to discuss workflows, not to score people. One person’s tool adoption does not determine the team’s capability. Check the actual handoffs, review capacity and support needed.
Use what you observe to choose a next step in the Adoption Roadmap and relevant questions in Measurement.
Adoption Roadmap
This four-phase example uses a 180-day schedule to organize a trial and possible expansion. Adapt its dates and activities to your team. Expand when the work and checks justify it; a calendar date or task count does not prove readiness.
Contained Pilot — Days 1–30
Suggested activities
- Select one repository with moderate complexity
- Choose a small set of repeatable tasks, such as a local prototype, test or refactor
- Measure baseline metrics: PR cycle time, change failure rate, test coverage, bug rate
- Agree scope, restrict access, protect sensitive data and name a reviewer for agent changes
- Give participants a chance to try relevant tasks with support
- Document what works, what fails, and what surprises
ProductPM participates in defining task types. Designer reviews agent-generated UI. Baseline product metrics recorded.
Before expanding
- Baseline metrics recorded for comparison
- Representative tasks reviewed, with enough evidence to explain the next decision
- No critical quality incidents from agent output
- Team can articulate which tasks agents handle well and which they don't
- Basic rules file created and shared across the team
Checks and precautions
- A named, capable person reviews each agent change during the pilot
- Start with low-risk, well-bounded tasks only
- If a quality incident occurs, pause, investigate and agree what must change before continuing
Expand Safely — Days 31–60
Suggested activities
- Expand to 2–3 repositories
- Create task templates for each repeatable pattern
- Introduce risk labels (low / medium / high) on every agent task
- Implement full quality guardrail layer (Layer 2)
- Review existing policy controls, including secrets, branch protection and access, for the expanded work
- Start tracking adoption KPIs: % PRs agent-assisted, CI first-pass rate
- Expand rules file based on Phase 1 lessons
ProductPM creates acceptance criteria templates. Designer contributes design tokens and component specs. Begin tracking design compliance rate.
Before expanding
- Task briefs cover the recurring work chosen for this trial
- Risk labeling applied to all agent tasks
- Quality guardrails (Layer 2) fully automated in CI
- Policy controls cover the data and actions in the expanded workflow
- Adoption KPIs tracked weekly
- No increase in change failure rate compared to baseline
Checks and precautions
- Maintain senior review on medium and high risk tasks
- Change review depth only when risk, verification coverage and observed results support it
- Weekly retrospective on agent output quality
Standardize — Days 61–90
Suggested activities
- Write a short standard operating procedure (SOP) covering task planning, checks, ownership and escalation
- Review policy controls for the workflow, including personal-data handling and output checks where applicable
- Add repo-level policy enforcement
- Train all team members on SOP, task templates, and rules files
- Discuss current practices using the five levels and choose a useful improvement
- Establish evaluation framework beyond CI
- Define roles and decision rights
ProductProduct team trained on SOP alongside engineering. Product-specific KPIs added to dashboard. PM owns Plan phase. Designer owns design system compliance.
Before expanding
- Internal SOP is published and accessible to all team members
- The people doing and reviewing the work understand the procedure and know where to ask for help
- Guardrail stack (Layers 1–4) fully operational
- The team can explain its working practices and the next improvement to try
- Evaluation framework exists beyond CI
- Decision rights are documented
Checks and precautions
- Ensure SOP is a living document, not a one-time artifact
- Schedule quarterly SOP reviews
- Assign an SOP owner responsible for updates
Scale — Days 91–180
Suggested activities
- Roll out to additional teams and repositories
- Implement governance layer (Layer 5): agent registry, access control, cross-team observability
- Share useful task examples, instructions and templates across teams
- Establish cost budgeting per team and per agent workflow
- Begin experimenting with Level 4 capabilities (background agents, async PRs)
- Publish organizational metrics dashboard
- Conduct cross-team retrospectives
- Evaluate dedicated role staffing
ProductProduct outcome metrics in org dashboard. Evaluate dedicated QA Engineer staffing. PM templates shared across teams. Design system fully instrumented.
Before expanding
- Multiple teams operating under the same SOP
- Governance layer (Layer 5) operational (minimum: registry + cost tracking)
- Teams can find and use the shared task examples and instructions
- Organizational KPI dashboard published and reviewed weekly
- Change failure rate stable or improved relative to baseline
- Review shows whether expanded use is helping and what needs to change
- Roles and decision rights scaled to match organizational breadth
Checks and precautions
- Check that responsibilities, working practices and support are clear before expanding
- Start Level 4 experiments in a single pod before expanding
- Monitor cost carefully during scale-out; token spend can increase non-linearly
Measurement and Failure Modes
KPI Dashboard
Choose key performance indicators (KPIs) that help answer a decision your team faces. Review user outcomes, quality and total effort together. The suggested directions below need context; none is a score for an individual.
Lead time
Discuss decreaseTime from issue opened to code merged
PR review time
Discuss decreaseTime from PR opened to approved
Change failure rate
Check stability% of deployments causing incidents or rollbacks
Rollback frequency
Check stabilityNumber of rollbacks per deployment period
Escaped defects
Discuss decreaseBugs found in production per sprint
Test coverage delta
MonitorChange in test coverage over time
Deployment frequency
Discuss increaseHow often the team deploys to production
% PRs agent-assisted
MonitorProportion of PRs that involved agent execution
% PRs passing CI first run
Discuss increaseShare of first CI runs that pass; interpretation depends on the checks and task mix
% tasks within SLA
Discuss increaseAgent tasks completed within defined time/iteration bounds
Contribution split
MonitorRatio of agent-assisted vs. fully manual work
Rules file update frequency
MonitorHow often the team's rules and templates are refined
Cost per agent task
Discuss decreaseAverage token/compute spend per completed task
Feature adoption rate
Discuss increase% of users engaging with agent-built features
User satisfaction delta
Check stabilityNPS/CSAT change for agent-assisted releases
Requirement accuracy
Discuss increase% of shipped features matching acceptance criteria on first pass
Design compliance rate
Discuss increase% of agent-generated UI matching design system
Review rejection rate
Monitor% of agent PRs rejected in code review
Post-merge defect rate
Discuss decreaseBugs introduced by agent-generated code found after merge
Evaluation coverage
Discuss increase% of agent output types covered by automated evaluation
Guardrail trigger rate
MonitorHow often guardrails catch issues before merge
This pattern may suggest improvement. Check user outcomes and total effort before drawing a conclusion.
Six Failure Modes
These examples describe symptoms, possible explanations and actions to try. Investigate the cause in your own work before choosing a response.
What to investigate
The selected tasks may not address an important problem, or the benefit may be offset by review and rework.
Actions to try
- Tie every agent workflow to a measurable delivery KPI
- Require a "so what?" test: if automated, what bottleneck does it remove?
- Review task selection criteria quarterly
What to investigate
Agent output velocity exceeds the team's review capacity. Often caused by large, unfocused agent PRs.
Actions to try
- Enforce smaller PR scope (one concern per PR, bounded by task template)
- Tighten acceptance criteria so PRs are more focused
- Scale review capacity: train more team members
- Implement risk-based review: low-risk PRs get sampling-based review
- Consider review automation for mechanical aspects
What to investigate
The checks may be missing important cases. Compare the reported problems with what is actually tested.
Actions to try
- Expand evaluation coverage beyond unit tests (integration tests, performance benchmarks, architecture fitness functions)
- Track post-release defect rate specifically for agent-generated code
- Implement regular "agent output audits"
- Monitor change failure rate as an early warning signal
What to investigate
Effective agent interaction patterns are not captured and shared. Knowledge stays in individual heads.
Actions to try
- Convert individual prompts into shared task templates
- Maintain team-level rules files (not personal ones)
- Publish an internal "agentic SOP" with examples
- Pair programming sessions where skilled users demonstrate approach
- Use the Learn phase to share useful findings and remove unhelpful instructions
What to investigate
Shared records, ownership or controls may not have kept up with the work. Check where visibility or responsibility is missing.
Actions to try
- Implement governance layer before cross-team scaling
- Start with clear ownership, appropriate access, useful activity records and cost tracking
- Review access controls and audit records before adding more workflows or teams
- Assign a governance owner (AI Reliability or Platform Engineer)
- Review governance completeness quarterly
What to investigate
The work may not address the intended user problem. Check the evidence behind the priorities and what happened after release.
Actions to try
- Tie agent task selection to product outcome metrics
- Require PM sign-off on every task plan
- Measure feature adoption and user satisfaction alongside delivery speed
- Apply "redesign, don't automate" to product discovery, not just delivery
Further Reading
- Building AI Agents Without Organizational Chaos — Chrono Innovation
- The Agentic Organization — McKinsey
- Seizing the Agentic AI Advantage — McKinsey
- The State of AI in 2025 — McKinsey
- Agentic AI Strategy — Deloitte
- Human-Agentic Workforce — Deloitte
- State of AI in the Enterprise 2026 — Deloitte
- Agentic AI Enterprise Adoption — Deloitte
- Top Strategic Technology Trends 2026 — Gartner
- Five Stages of Agentic Evolution — Gartner
- The 8 Levels of Agentic Engineering — Eledath
- Agentic Engineering for Software Teams — vibecoding.app
- Agentic AI for PMs — IdeaPlan
- From UX to AX — The Atlantic
- Five Product Shifts — cases.media