Enterprise Adoption and Factory Maturity Model
Enterprises often buy agent capability faster than they build the operating system required to govern it. A maturity model must describe observable organizational capability, not enthusiasm, model intelligence, or the percentage of code gen
A rapid review of the chapter’s existing Quick Read, principles, definitions, lessons, and review material.
Reconstruct and defend this chapter’s architecture.
Reconstruct the architecture, name each boundary, and defend the tradeoffs.
Open the source exercise
Map a 5,000-engineer enterprise across the eight maturity dimensions. Select one product corridor for Level 2, show a 90-day adoption plan, define baselines and promotion gates, and respond to a CIO who wants autonomous production deployment in quarter one.
3. Enduring Principle
Assess capability by domain and scope
Use maturity levels per repository, workflow, risk class, and environment. The enterprise level is the lowest material capability needed for the claimed use case—not the highest demo achieved anywhere.
| Level | Operating mode | Required proof |
|---|---|---|
| 0 — Human execution | AI suggests; humans perform and own all actions | Secure tool use and human review |
| 1 — Assisted execution | Human initiates bounded deterministic or generative tasks | Traceable inputs/outputs, conventional CI, no delegated authority |
| 2 — Delegated execution | Human approves a WorkOrder; factory implements; humans review material outputs | Persistent state, bounded authority, isolated Attempts, deterministic controls, review-ready PR |
| 3 — Governed autonomy | Factory may plan/execute within policy; independent validators prove quality; humans handle material risk | Versioned Plans, policy engine, independent evidence, recovery, audit, calibrated trust |
| 4 — Conditional autonomy | Policy may authorize merge/deployment for qualified low-risk scopes | Artifact provenance, production verification, rollback, SLOs, mature exception and demotion controls |
| 5 — Trusted factory operation | Humans govern policies and portfolios rather than routine individual work | Sustained evidence across production outcomes, continuous control validation, governed learning, rapid quarantine |
Level 5 is an operating condition, not a permanent badge. A critical event can demote or quarantine a scope immediately.
Evaluate multiple dimensions
Score evidence in at least these dimensions: specification, orchestration, quality/evidence, security/identity, delivery/operations, governance/roles, observability/economics, and learning/change control. Numeric scores support trends; bands communicate decisions. Hard gates override averages.
Advance through use-case corridors
Do not transform the entire SDLC at once. Choose a repeatable, valuable, reversible corridor such as governed issue to validated PR. Establish baseline lead time, change failure, human effort, wait time, satisfaction, and control escapes. Prove Level 2, stabilize it, then shadow Level 3 decisions before enforcement.
Recommended progression:
- Observe: instrument existing work and establish baseline.
- Assist: use agents within human-operated workflows.
- Delegate: authorize bounded WorkOrders and retain full review.
- Shadow govern: compute policy/evidence decisions without changing outcomes.
- Enforce: block one narrow transition on objective evidence.
- Conditionally automate: permit low-risk transitions with rollback.
- Scale: expand only after sustained outcome evidence.
Change the organization, not only the tooling
Product owners improve intent and criteria. Platform teams provide paved execution and evidence paths. Security encodes policy and reviews exceptions. Quality engineers design assurance systems and adversarial validation. Staff engineers own invariants and architecture. Leaders measure customer value, risk, and cognitive load—not generated lines.
Adoption requires training, role clarity, psychological safety for reporting failures, and explicit accountability. Approval fatigue is reduced by evidence-centered reviews and exception routing, not by deleting accountability.
Promote and demote through evidence
For Level 2 to Level 3, the working default is at least 100 successful WorkOrders over 30 days, at least 99% independent validation success, zero critical policy/security violations, zero unauthorized actions, and explicit human promotion. These are starting policy values, not universal science. Riskier domains should require more.
Failures decay in scoring but never disappear from audit history. Security bypass, fabricated evidence, unauthorized action, tampering, or high-impact validation escape triggers immediate demotion or quarantine pending review.
8. Personal notes and lessons learned
- Capability is not autonomy; evidence and policy convert capability into authority.
- Maturity must be scoped. One safe documentation bot does not make a trusted factory.
- Shadow mode is a learning phase only when disagreement and calibration are measured.
- The credible executive message is “we expand proven corridors,” not “AI transforms everything.”
9. Interview questions
- How would you assess an enterprise that has widespread Copilot use but no governed WorkOrders?
- What proof is needed before moving from delegated execution to governed autonomy?
- How do you prevent a maturity model from becoming checkbox theater?
- Which roles change most under an AI Software Factory?
- When should autonomy be demoted immediately?
Review this chapter.
Challenge a claim, boundary, missing failure mode, unclear term, or unsupported evidence statement.
- Claim
- Boundary
- Failure
- Evidence