Working with agents
Agent-assisted delivery needs checks on the generated work and on the resulting product. Plan those checks with the people who will use their results.
Agent output quality asks whether the generated work follows its instructions and meets the relevant checks. Product quality asks whether the result works for users, including its behavior, clarity and accessibility.
HELM names these responsibilities Evaluation Lead and QA Engineer. They need clear ownership, but they need not be separate jobs in every team. Decide based on workload, expertise and risk.
Experience to discuss
Discuss how someone chooses tests, investigates conflicting signals and checks the experience of using a product. Use a relevant agent-generated example and ask what automated checks would miss and how they would handle that gap.
Responsibilities and skills
Purpose
Make sure the team checks both agent output and the resulting product, with clear ownership and follow-up for problems.
Responsibilities to discuss
- Evaluation Lead: Design evaluation suites for agent behavior that go beyond conventional unit and integration tests
- Evaluation Lead: Agree what acceptable output looks like for each workflow
- Evaluation Lead: Track quality over time and investigate changes
- Evaluation Lead: Make the agreed checks part of the release decision
- Evaluation Lead: Automate useful checks and identify where human review is still needed
- Evaluation Lead: Partner with the AI Architect on what "correct" means per task type
- QA Engineer: Translate acceptance criteria into testable assertions that reflect real user outcomes
- QA Engineer: Build or curate suites that stress UX regressions, accessibility, copy, and interaction quality
- QA Engineer: Monitor product-side drift: issues that clear agent evaluation but still violate user expectations
- QA Engineer: Coordinate with the Evaluation Lead so coverage spans both agent output and end-to-end product behavior
- QA Engineer: Review agent-generated UI for design-system fit, accessibility, and interaction quality
- QA Engineer: Check whether the result works for users during the Verify phase
Skills for this work
These role-specific skills complement the five shared competencies. Choose examples relevant to the work and support people as they practice.
- Evaluation design — Choose meaningful examples and criteria for judging outputs that may have more than one acceptable answer.
- Drift detection — Notice changes in quality over time, including problems individual test runs can miss.
- Product judgment — Assessing experience quality beyond functional pass/fail.
- Statistical thinking — Understand the limits of a sample and explain what the available evidence supports.
- Automation at scale — Build useful automated checks and plan the human review effort they leave.
- Cross-functional communication — Turning quality signals into concrete, prioritized feedback for engineering and product.
Compare the skills in this guide and choose an area to discuss or practice.
Signals that need more context
Use work examples alongside these signals; none is a complete measure of someone’s ability.
- Test execution counts without examples of choosing and investigating meaningful checks
- Depth in a single framework without judgment about what to automate and why
- Quality defined only as absence of defects, ignoring intent and experience
- Assumptions that all code is human-written, reviewed at human cadence, and stable between releases
- Release checks without involvement in planning and reviewing the work
Questions to discuss
Adapt these example questions to the role and the person’s opportunities to do the work.
- Evaluation design — An agent generates API endpoints. Design an evaluation suite that decides whether the output is production-ready. What do you measure beyond tests passing?
- Drift detection — CI is green on agent PRs, but customer bug reports are up ~15% month over month. How do you investigate?
- Product quality — An agent-built checkout passes all functional tests. What do you still verify? (Probe UX, accessibility, copy, edge cases, trust.)
- Threshold setting — For a task with no single right answer, how do you define "good enough"? Walk through your framework.
- Process design — For a team at Maturity Level 3, design the quality workflow. Where does evaluation run? Where does product QA run? How do they hand off and escalate?
An example day
Illustrative scenario. Use this example to discuss how the responsibilities fit together.
Output review: A generated change passes its tests but mishandles an input missing from the test data. You add that case, explain the expected behavior and ask for a correction.
Product review: A walkthrough reveals an unclear error message and a keyboard trap. You describe the user impact, agree fixes and add checks before release. The team reviews what both kinds of testing missed.
Related HELM guidance
The Evaluation Lead and QA Engineer responsibilities contribute to Layer 2: Quality and the Verify phase. Agree how your team covers both kinds of review.
Use Silent Quality Drift and the KPI Dashboard to investigate cases where test results and user experience disagree.