Skip to content
Roles/engineering

Software Engineer

Build, review and maintain software with AI agents

Related job titles
Frontend EngineerBackend EngineerFull-Stack Engineer
HELM responsibilitiesProduct Engineer

HELM 1.0.1 · Updated

Table of contents

Working with agents

Agents can help draft implementations, refactor code and prepare tests. Engineers still need to understand the problem, the code and the system it runs in.

Before delegating, define the task and decide how to check it. Review the result for requirements, integration problems and failures that the agent’s tests may miss.

Keep practicing the coding and debugging skills needed to understand those changes. Take responsibility for what you accept, release and maintain.

Experience to discuss

Programming, API, database and framework experience can help someone understand and check an agent’s work. Ask for examples of planning a task, reviewing code and investigating a defect.

Choose examples relevant to the team and discuss the reasoning behind the work, alongside the result.

Responsibilities and skills

Purpose

Deliver useful software with agents while keeping the code understandable, tested and fit for the product.

Responsibilities to discuss

  • Define task plans with clear scope, acceptance criteria, and constraints
  • Run the Plan-Execute-Verify-Ship-Learn loop for scoped delivery tasks
  • Review agent-generated code for architectural alignment, edge cases, and security
  • Create and maintain task templates and agent instructions (rules files)
  • Integrate agent output into the product, ensuring it meets UX and product standards
  • Use test results to guide another attempt, then review the change
  • Share useful lessons through examples, tests and task templates
  • Collaborate with PM and Design to translate acceptance criteria into agent-executable plans

Skills for this work

These role-specific skills complement the five shared competencies. Choose examples relevant to the work and support people as they practice.

  • Context engineering — Give agents relevant context, clear constraints and examples of the result you need.
  • Cross-layer code review — Understand how a change affects frontend, backend and infrastructure. Involve the right expertise when a part needs closer review.
  • Architectural judgment — Recognize when a change conflicts with the architecture or creates security, reliability or maintenance problems.
  • Quality evaluation at volume — Keep review thorough as output increases. Prioritize meaningful checks and keep the workload manageable.
  • Product awareness — Check whether the change solves the user’s problem as well as meeting technical requirements.
  • System thinking — Trace how a change affects dependencies, interfaces and behavior in use.
Role skills explorer

Compare the skills in this guide and choose an area to discuss or practice.

Explore

Signals that need more context

Use work examples alongside these signals; none is a complete measure of someone’s ability.

  • Whiteboard algorithm challenges disconnected from how the team actually ships
  • Arbitrary years-of-experience gates tied to specific framework versions
  • Raw speed of manual typing or line count as a proxy for seniority
  • Syntax recall without examples of applying or checking that knowledge
  • Specialist knowledge without evidence of working across relevant boundaries

Questions to discuss

Adapt these example questions to the role and the person’s opportunities to do the work.

  • Code review exercise — Evaluate an agent-generated pull request for correctness, security, and fit with the system's architecture.
  • Task planning — Given a feature request, produce a plan that bounds agent work: scope, non-goals, acceptance criteria, and risk level.
  • System design — Discuss decisions agents should not make alone: boundaries, ownership, failure modes, and evolution of the architecture.
  • Judgment scenarios — Present flawed agent output and ask what is wrong, and how you would change instructions or rules to prevent recurrence.
  • Debugging — Diagnose a subtle defect in agent-generated code that satisfies tests but fails under realistic edge conditions or integration pressure.

An example day

Illustrative scenario. Use this example to discuss how the responsibilities fit together.

You review three changes prepared by an agent. One is ready. Another breaks sign-in in a case the tests missed, and the third addresses the wrong requirement. You request corrections and add a test for the sign-in failure.

For the next task, you agree the scope and expected behavior with product and design. You give the agent that context, review its result and share an example of the requirement that was previously misunderstood.

Related HELM guidance

Use Plan-Execute-Verify-Ship-Learn to organize the task and the Task Classification Matrix to discuss its scope and risk. Principle 3 keeps the responsibility clear: people decide what to accept and release.