Skip to content

Evidence and Sources

HELM combines proposed practices, provider guidance, practitioner commentary and enterprise research. Publication status describes availability; it does not establish validation. Last evidence review: .

This register covers the baseline’s major recommendations, numerical claims and categorical statements. No HELM pilot is currently labeled a measured result. Sources listed only as further reading are explicitly marked unreviewed. Conceptual corrections are tracked separately from evidence labeling.

Evidence labels

Proposed recommendation
HELM guidance requiring application and validation in context. A related citation does not validate the full recommendation.
Author observation
A documented observation by the HELM author, with context, date and limitations.
Participant report
A participant’s attributed account, with permission and context; not independently measured.
Measured result
A result with a baseline, method, observation period, task mix and limitations.
Externally supported claim
The linked source was inspected and supports the specific statement. The source’s method and applicability still limit the claim.
Unverified claim
The wording, attribution, number or inference has not been substantiated. Do not treat it as established evidence.
Illustrative example
A teaching scenario or proposed target, not an observed implementation or measured result.

Claim register

Simplicity First

Proposed recommendation · recommendation · reviewed 2026-09-17

Anthropic supports choosing simpler designs based on observed engineering experience. HELM’s progression is guidance; the assertion that these principles underpin every successful implementation is unverified.

Guidance: /foundation#principle-1-simplicity-first

Follow-up: C-07.

Redesign, Don’t Automate

Proposed recommendation · recommendation · reviewed 2026-09-17

Workflow redesign is supported as a recommendation in enterprise commentary. It is not evidence that every workflow requires redesign or that most failures have a single cause.

Guidance: /foundation#principle-2-redesign-dont-automate

Follow-up: C-01.

14% agentic deployment readiness

Unverified claim · survey · reviewed 2026-09-17

The exact 14% figure, population and explanation were not located on the inspected landing page. Full report/primary-claim verification is outstanding; treat this baseline statistic as unverified.

Guidance: /foundation#principle-2-redesign-dont-automate

Follow-up: C-08.

Nearly 80% adoption and limited bottom-line impact

Externally supported claim · survey · reviewed 2026-09-17

The June 2025 report states both proportions and cites its March 2025 survey. These are enterprise self-reports. The following “because” explanation is the report’s interpretation, not demonstrated causation.

Guidance: /foundation#principle-2-redesign-dont-automate

Agents Execute, Humans Are Accountable

Proposed recommendation · recommendation · reviewed 2026-09-17

HELM’s normative allocation of accountability and its human/agent ownership table. Related provider guidance is not validation of every categorical boundary.

Guidance: /foundation#principle-3-agents-execute-humans-are-accountable

Follow-up: C-06.

Guardrails Are Non-Negotiable

Proposed recommendation · recommendation · reviewed 2026-09-17

The five-layer stack is HELM’s synthesis. Its completeness and mandatory scaling threshold have not been validated by a HELM pilot.

Guidance: /foundation#principle-4-guardrails-are-non-negotiable · /practitioners#the-guardrail-stack

Follow-up: C-06.

Over 40% canceled or failed by 2027

Unverified claim · forecast · reviewed 2026-09-17

A forecast, not an observed failure rate. The cited trends URL returned 403; the exact forecast, publication year and attribution to insufficient risk controls remain unverified.

Guidance: /foundation#principle-4-guardrails-are-non-negotiable

Follow-up: C-08.

Structure Over Tooling

Proposed recommendation · recommendation · reviewed 2026-09-17

Clear ownership is a recommendation. The baseline claim that most failures are structural, and that the anecdote is the default trajectory, has no prevalence evidence.

Guidance: /foundation#principle-5-structure-over-tooling

Follow-up: C-07.

Week-three/week-five/week-six quotation

Externally supported claim · commentary · reviewed 2026-09-17

The quotation is present in the linked article. Its presentation as “one team’s experience” is not substantiated by case details; it is consultancy commentary rather than a documented participant report.

Guidance: /foundation#principle-5-structure-over-tooling

Follow-up: C-08.

One in five with mature agent governance

Externally supported claim · survey · reviewed 2026-09-17

The 2026 landing page states this finding. Sample: 3,235 leaders in 24 countries at AI-leading organizations, surveyed August–September 2025. This is self-reported governance maturity, not a universal population estimate.

Guidance: /foundation#principle-5-structure-over-tooling

Team-Wide Adoption Over Individual Mastery

Proposed recommendation · recommendation · reviewed 2026-09-17

The multiplayer effect is practitioner commentary. The exact claim that team maturity equals its least-adopted critical-path member is an unvalidated HELM extrapolation.

Guidance: /foundation#principle-6-team-wide-adoption-over-individual-mastery · /practitioners#maturity-model · /leadership#maturity-model

Follow-up: C-03.

Cowork in 10 days and Block’s 100+ skills

Unverified claim · commentary · reviewed 2026-09-17

Both numbers appear as secondary accounts in Eledath’s article. Primary company evidence and the causal claim that team-wide adoption explains the outcomes have not been verified.

Guidance: /foundation#principle-6-team-wide-adoption-over-individual-mastery

Follow-up: C-08.

Five-component agent anatomy

Proposed recommendation · definition · reviewed 2026-09-17

A HELM teaching taxonomy. “Every agent” and “five core components” are not established universal requirements. Model names are examples, not evaluated recommendations.

Guidance: /practitioners#agent-anatomy

Follow-up: C-07.

Workflows versus agents

Externally supported claim · definition · reviewed 2026-09-17

Anthropic explicitly distinguishes predefined code paths from model-directed processes. The claim it was “first articulated” there and every cost/predictability comparison are not established by this source review.

Guidance: /practitioners#workflows-vs-agents

Follow-up: C-07.

Tool types and risk labels

Proposed recommendation · recommendation · reviewed 2026-09-17

Classification and risk assignments are proposed guidance. Read-only access can expose sensitive information; the blanket low-risk wording is pending correction.

Guidance: /practitioners#tool-types

Follow-up: C-06.

Eight composition patterns and selection guide

Proposed recommendation · recommendation · reviewed 2026-09-17

The providers describe related patterns. HELM’s eight-pattern set and simplest-to-most-complex ordering are a synthesis, not a validated universal ordering.

Guidance: /practitioners#composition-patterns

Plan-Execute-Verify-Ship-Learn

Proposed recommendation · recommendation · reviewed 2026-09-17

A proposed repeatable operating practice. No HELM implementation outcome is claimed; the 1.1 pilot will test practical utility and total effort.

Guidance: /practitioners#the-operating-loop

Task matrix and human-agent boundary

Proposed recommendation · recommendation · reviewed 2026-09-17

The nine-cell matrix, examples and CI-based boundary are HELM guidance. The matrix is not a validated risk classifier; permission, consequence and verification gaps are tracked for correction.

Guidance: /practitioners#task-classification-matrix · /practitioners#the-human-agent-boundary

No external source attached; this is explicitly provisional or illustrative.

Follow-up: C-06.

Five-level maturity model

Proposed recommendation · recommendation · reviewed 2026-09-17

HELM adapts eight individual-practice levels into five team levels. The resulting dimensions, thresholds and failure modes have not been validated as a maturity instrument.

Guidance: /practitioners#maturity-model · /leadership#maturity-model

Follow-up: C-02.

Vertical pods and six organizational shifts

Proposed recommendation · recommendation · reviewed 2026-09-17

Proposed organizational guidance. Existing roles, cross-functional work and workflow constraints require context; an agent does not by itself establish a need to reorganize.

Guidance: /leadership#organizational-model

Follow-up: C-01.

Nine authorities, dedicated positions and scaling counts

Proposed recommendation · recommendation · reviewed 2026-09-17

Role authorities and the 7–8/12–16 staffing examples are HELM proposals. The source focuses on runtime AI products; the staffing counts are not empirically established minima.

Guidance: /leadership#roles-with-explicit-authority · /leadership#scaling-path

Follow-up: C-04.

Decision Rights Matrix

Proposed recommendation · recommendation · reviewed 2026-09-17

The matrix is a proposed governance tool. Paired owners conflict with the surrounding single-owner rule; the accountable-owner/approval distinction requires a versioned correction.

Guidance: /leadership#decision-rights-matrix

Follow-up: C-05.

180-day adoption roadmap

Illustrative example · example · reviewed 2026-09-17

The days, task counts, review percentages and exit criteria are proposed planning targets. They are not a measured time-to-adoption benchmark or validated safety thresholds.

Guidance: /leadership#adoption-roadmap

No external source attached; this is explicitly provisional or illustrative.

KPI dashboard and success formula

Proposed recommendation · recommendation · reviewed 2026-09-17

HELM’s metric definitions, directions and formula are recommendations, not validated causal indicators. Review rejection and coverage require contextual interpretation; no increase/decrease target establishes individual performance.

Guidance: /leadership#kpi-dashboard

No external source attached; this is explicitly provisional or illustrative.

Follow-up: C-09.

Six failure modes and mitigations

Proposed recommendation · recommendation · reviewed 2026-09-17

Diagnostic hypotheses and example symptoms, including 3x delivery speed. No prevalence, causal certainty or measured effectiveness is established.

Guidance: /leadership#six-failure-modes

No external source attached; this is explicitly provisional or illustrative.

Role transformations, competencies and hiring guidance

Proposed recommendation · recommendation · reviewed 2026-09-17

The eight role guides and universal competency list are HELM proposals. Statements about traditional work, necessary role changes and abandoned hiring criteria require contextual validation. Day-in-the-life and interview examples are illustrative.

Guidance: /roles

No external source attached; this is explicitly provisional or illustrative.

Follow-up: C-07.

3–10x output and 10x execution claims

Unverified claim · commentary · reviewed 2026-09-17

No source, measurement method or comparison is attached to these role-description multipliers. They must not be interpreted as promised or observed productivity gains.

Guidance: /roles/engineering-manager · /roles/product-manager

No external source attached; this is explicitly provisional or illustrative.

Follow-up: C-07.

Competency explorer percentages

Illustrative example · example · reviewed 2026-09-17

Percentages count editorially classified competencies in each role guide. They do not measure a person’s skill gap, readiness, proficiency or proportion of work.

Guidance: /roles/competency-map

No external source attached; this is explicitly provisional or illustrative.

Source register

_From UX to AX_

Cited by: leadership.

Registered; not yet inspected. No access date or claim verification is asserted. Applicability remains unassessed; this reference is further reading rather than verified support.

History: 2026-09-17 — inventoried from the baseline bibliography; original URL retained.

_Five Product Shifts_

Cited by: leadership.

Registered; not yet inspected. No access date or claim verification is asserted. Applicability remains unassessed; this reference is further reading rather than verified support.

History: 2026-09-17 — inventoried from the baseline bibliography; original URL retained.

_Cloud Adoption Framework_

Cited by: foundation, practitioners.

Registered; not yet inspected. No access date or claim verification is asserted. Applicability remains unassessed; this reference is further reading rather than verified support.

History: 2026-09-17 — inventoried from the baseline bibliography; original URL retained.

_A Practical Guide to Building Agents_

Cited by: foundation, practitioners.

Registered; not yet inspected. No access date or claim verification is asserted. Applicability remains unassessed; this reference is further reading rather than verified support.

History: 2026-09-17 — inventoried from the baseline bibliography; original URL retained.

Linked claims: Agents Execute, Humans Are Accountable, Guardrails Are Non-Negotiable, Five-component agent anatomy, Tool types and risk labels, Eight composition patterns and selection guide

_Agentic Engineering for Software Teams_

Cited by: foundation, leadership, practitioners.

Registered; not yet inspected. No access date or claim verification is asserted. Applicability remains unassessed; this reference is further reading rather than verified support.

History: 2026-09-17 — inventoried from the baseline bibliography; original URL retained.

_Building Effective Agents_

Cited by: foundation, practitioners.

Source inspected. Attempted 2026-09-17; accessed 2026-09-17.

Applicability: Provider engineering experience with LLM workflows and agent applications.

Supports simple composable patterns and the workflow/agent distinction. It does not validate HELM’s complete eight-pattern ordering or universal success language.

Source history
  • 2026-09-17: Initial claim-level review; existing URL retained.

Linked claims: Simplicity First, Agents Execute, Humans Are Accountable, Guardrails Are Non-Negotiable, Five-component agent anatomy, Workflows versus agents, Eight composition patterns and selection guide

_The 8 Levels of Agentic Engineering_

Cited by: foundation, leadership, practitioners.

Source inspected. Attempted 2026-09-17; accessed 2026-09-17.

Applicability: March 10, 2026 practitioner article, updated March 11, about AI-assisted coding.

Contains eight levels, a multiplayer-effect example, Cowork in 10 days and a secondary account of Block’s 100+ skills. Does not establish the exact least-member formula or a causal explanation for company outcomes.

Source history
  • 2026-09-17: Initial claim-level review; existing URL retained.

Linked claims: Team-Wide Adoption Over Individual Mastery, Cowork in 10 days and Block’s 100+ skills, Plan-Execute-Verify-Ship-Learn, Five-level maturity model

_Scaling AI Requires New Processes_

Cited by: foundation.

Registered; not yet inspected. No access date or claim verification is asserted. Applicability remains unassessed; this reference is further reading rather than verified support.

History: 2026-09-17 — inventoried from the baseline bibliography; original URL retained.

_Building AI Agents Without Organizational Chaos_

Cited by: foundation, leadership.

Source inspected. Attempted 2026-09-17; accessed 2026-09-17.

Applicability: January 15, 2026 consultancy commentary about building AI products.

The week-three/week-five/week-six wording appears in the article. No named team, study method or measured case is supplied. Staffing recommendations address AI-product construction and cannot establish universal coding-agent staffing requirements.

Source history
  • 2026-09-17: Initial claim-level review; existing URL retained.

Linked claims: Structure Over Tooling, Week-three/week-five/week-six quotation, Vertical pods and six organizational shifts, Nine authorities, dedicated positions and scaling counts, Decision Rights Matrix

_Human-Agentic Workforce_

Cited by: leadership.

Registered; not yet inspected. No access date or claim verification is asserted. Applicability remains unassessed; this reference is further reading rather than verified support.

History: 2026-09-17 — inventoried from the baseline bibliography; original URL retained.

_Agentic AI Strategy_

Cited by: foundation, leadership.

Registered; not yet inspected. No access date or claim verification is asserted. Applicability remains unassessed; this reference is further reading rather than verified support.

History: 2026-09-17 — inventoried from the baseline bibliography; original URL retained.

_Agentic AI Enterprise Adoption_

Cited by: leadership.

Registered; not yet inspected. No access date or claim verification is asserted. Applicability remains unassessed; this reference is further reading rather than verified support.

History: 2026-09-17 — inventoried from the baseline bibliography; original URL retained.

_State of AI in the Enterprise 2026_

Cited by: foundation, leadership.

Source inspected. Attempted 2026-09-17; accessed 2026-09-17.

Applicability: Survey of 3,235 leaders at AI-leading organizations in 24 countries, August–September 2025.

The 2026 landing page reports one in five with mature autonomous-agent governance. The 14% deployment-readiness statement was not located on that page. Findings are self-reported, not a causal test of HELM.

Source history
  • 2026-09-17: Initial claim-level review; existing URL retained.

Linked claims: Redesign, Don’t Automate, 14% agentic deployment readiness, Structure Over Tooling, One in five with mature agent governance

_Five Stages of Agentic Evolution_

Cited by: leadership, practitioners.

Registered; not yet inspected. No access date or claim verification is asserted. Applicability remains unassessed; this reference is further reading rather than verified support.

History: 2026-09-17 — inventoried from the baseline bibliography; original URL retained.

_Top Strategic Technology Trends 2026_

Cited by: foundation, leadership, practitioners.

Access blocked. Attempted 2026-09-17; accessed not retrieved.

Applicability: Technology-trends article; suitability for the specific cancellation forecast remains unverified.

HTTP 403 during review. The forecast, attribution year and causal wording need a directly inspectable primary source.

Source history
  • 2026-09-17: Retrieval blocked; retained original citation and recorded correction C-08. No silent source replacement.

Linked claims: Over 40% canceled or failed by 2027

_Agentic AI for PMs_

Cited by: leadership.

Registered; not yet inspected. No access date or claim verification is asserted. Applicability remains unassessed; this reference is further reading rather than verified support.

History: 2026-09-17 — inventoried from the baseline bibliography; original URL retained.

_AI Agents Whitepaper_

Cited by: practitioners.

Registered; not yet inspected. No access date or claim verification is asserted. Applicability remains unassessed; this reference is further reading rather than verified support.

History: 2026-09-17 — inventoried from the baseline bibliography; original URL retained.

_The Agentic Organization_

Cited by: foundation, leadership.

Registered; not yet inspected. No access date or claim verification is asserted. Applicability remains unassessed; this reference is further reading rather than verified support.

History: 2026-09-17 — inventoried from the baseline bibliography; original URL retained.

_Seizing the Agentic AI Advantage_

Cited by: foundation, leadership.

Source inspected. Attempted 2026-09-17; accessed 2026-09-17.

Applicability: June 13, 2025 enterprise report citing a March 2025 AI survey.

The report states nearly eight in ten use gen AI and a similar proportion report no significant bottom-line impact. Workflow redesign is its interpretation/recommendation, not a controlled causal result.

Source history
  • 2026-09-17: Initial claim-level review; existing URL retained.

Linked claims: Redesign, Don’t Automate, Nearly 80% adoption and limited bottom-line impact, Vertical pods and six organizational shifts

_The State of AI in 2025_

Cited by: leadership.

Registered; not yet inspected. No access date or claim verification is asserted. Applicability remains unassessed; this reference is further reading rather than verified support.

History: 2026-09-17 — inventoried from the baseline bibliography; original URL retained.

Review and correction policy

Maintainer: Firas Kafri. Review major claims quarterly, and sooner when a source changes or contradictory evidence appears. Preserve the old URL and record replacement rationale and date. Changes to guidance go through the compatibility policy and correction register. Modification dates record editorial changes; evidence-review dates record source inspection.

To report an error, open an issue with the claim ID, proposed correction and inspectable source. Keep private company and employee information out of public reports.