Enterprise Adoption and Factory Maturity Model
Enterprises often buy agent capability faster than they build the operating system required to govern it. A maturity model must describe observable organizational capability, not enthusiasm, model intelligence, or the percentage of code gen
Reconstruct and defend this chapter’s architecture.
Reconstruct the architecture, name each boundary, and defend the tradeoffs.
Open the source exercise
Map a 5,000-engineer enterprise across the eight maturity dimensions. Select one product corridor for Level 2, show a 90-day adoption plan, define baselines and promotion gates, and respond to a CIO who wants autonomous production deployment in quarter one.
1. The problem
Enterprises often buy agent capability faster than they build the operating system required to govern it. A maturity model must describe observable organizational capability, not enthusiasm, model intelligence, or the percentage of code generated by AI.
2. Why the problem exists
Autonomy depends on specification quality, platform reliability, policy, evidence, security, delivery, and accountable roles. These capabilities mature unevenly. A team can have excellent coding agents and weak release governance, or strong CI and no durable WorkOrder authority. One label for the whole enterprise hides the limiting constraint.
3. Enduring Principle
Assess capability by domain and scope
Use maturity levels per repository, workflow, risk class, and environment. The enterprise level is the lowest material capability needed for the claimed use case—not the highest demo achieved anywhere.
| Level | Operating mode | Required proof |
|---|---|---|
| 0 — Human execution | AI suggests; humans perform and own all actions | Secure tool use and human review |
| 1 — Assisted execution | Human initiates bounded deterministic or generative tasks | Traceable inputs/outputs, conventional CI, no delegated authority |
| 2 — Delegated execution | Human approves a WorkOrder; factory implements; humans review material outputs | Persistent state, bounded authority, isolated Attempts, deterministic controls, review-ready PR |
| 3 — Governed autonomy | Factory may plan/execute within policy; independent validators prove quality; humans handle material risk | Versioned Plans, policy engine, independent evidence, recovery, audit, calibrated trust |
| 4 — Conditional autonomy | Policy may authorize merge/deployment for qualified low-risk scopes | Artifact provenance, production verification, rollback, SLOs, mature exception and demotion controls |
| 5 — Trusted factory operation | Humans govern policies and portfolios rather than routine individual work | Sustained evidence across production outcomes, continuous control validation, governed learning, rapid quarantine |
Level 5 is an operating condition, not a permanent badge. A critical event can demote or quarantine a scope immediately.
Evaluate multiple dimensions
Score evidence in at least these dimensions: specification, orchestration, quality/evidence, security/identity, delivery/operations, governance/roles, observability/economics, and learning/change control. Numeric scores support trends; bands communicate decisions. Hard gates override averages.
Advance through use-case corridors
Do not transform the entire SDLC at once. Choose a repeatable, valuable, reversible corridor such as governed issue to validated PR. Establish baseline lead time, change failure, human effort, wait time, satisfaction, and control escapes. Prove Level 2, stabilize it, then shadow Level 3 decisions before enforcement.
Recommended progression:
- Observe: instrument existing work and establish baseline.
- Assist: use agents within human-operated workflows.
- Delegate: authorize bounded WorkOrders and retain full review.
- Shadow govern: compute policy/evidence decisions without changing outcomes.
- Enforce: block one narrow transition on objective evidence.
- Conditionally automate: permit low-risk transitions with rollback.
- Scale: expand only after sustained outcome evidence.
Change the organization, not only the tooling
Product owners improve intent and criteria. Platform teams provide paved execution and evidence paths. Security encodes policy and reviews exceptions. Quality engineers design assurance systems and adversarial validation. Staff engineers own invariants and architecture. Leaders measure customer value, risk, and cognitive load—not generated lines.
Adoption requires training, role clarity, psychological safety for reporting failures, and explicit accountability. Approval fatigue is reduced by evidence-centered reviews and exception routing, not by deleting accountability.
Promote and demote through evidence
For Level 2 to Level 3, the working default is at least 100 successful WorkOrders over 30 days, at least 99% independent validation success, zero critical policy/security violations, zero unauthorized actions, and explicit human promotion. These are starting policy values, not universal science. Riskier domains should require more.
Failures decay in scoring but never disappear from audit history. Security bypass, fabricated evidence, unauthorized action, tampering, or high-impact validation escape triggers immediate demotion or quarantine pending review.
4. Tradeoffs
Maturity models can become compliance theater. Prevent this by requiring retained proof and by assessing actual workflows, not policy documents. Uniform enterprise standards improve control but can suppress local learning; define mandatory boundaries centrally and allow teams to experiment inside them. Early automation can show fast savings while increasing hidden review and incident costs, so measure total human and operational effort.
5. Current Mission Control Implementation
Mission Control expresses several Level 2 foundations: governed Missions/Plans/WorkOrders, Tasks and Attempts, leases, execution manifests, receipts, approval records, policy concepts, model routing, and operator views. Study branches add stronger sandbox and publication controls.
It has not yet earned a product-wide Level 3 claim. QC adapters are mocked, release automation is shadow mode, policy configuration has blocked the golden path, and the complete browser-initiated flow lacks accepted retained evidence. The honest rating is capability-specific: architecture and domain model approach Level 2/3 design, while the supported end-to-end operating proof remains below that claim.
6. Future Vision
Create a maturity evidence dashboard by controlled repository and workflow. It should show prerequisites, last accepted proof, incidents, exceptions, autonomy ceiling, and next gate. The first promotion target remains Governed Issue to Validated Pull Request; deployment autonomy waits for signed artifact identity, production observation, and rollback proof.
7. Versioned references
- Mission Control local HEAD:
a49064875d0711253d74029e3066cc74c7c1c2a5; staged-only work is excluded from maturity claims - Mastery sources: operational autonomy, governance, quality, economics, and golden-path lab chapters
- External canon: DORA delivery metrics; NIST SSDF outcome-based practices; Team Topologies; Google SRE
8. Personal notes and lessons learned
- Capability is not autonomy; evidence and policy convert capability into authority.
- Maturity must be scoped. One safe documentation bot does not make a trusted factory.
- Shadow mode is a learning phase only when disagreement and calibration are measured.
- The credible executive message is “we expand proven corridors,” not “AI transforms everything.”
9. Interview questions
- How would you assess an enterprise that has widespread Copilot use but no governed WorkOrders?
- What proof is needed before moving from delegated execution to governed autonomy?
- How do you prevent a maturity model from becoming checkbox theater?
- Which roles change most under an AI Software Factory?
- When should autonomy be demoted immediately?
10. Whiteboard exercise
Map a 5,000-engineer enterprise across the eight maturity dimensions. Select one product corridor for Level 2, show a 90-day adoption plan, define baselines and promotion gates, and respond to a CIO who wants autonomous production deployment in quarter one.
11. Hands-on lab
Assess Mission Control’s golden path using retained evidence only. Assign a level to each dimension, cite exact proof and gaps, identify the limiting dimension, and propose the smallest 30-day promotion experiment. Present the assessment twice: a technical review and a board-level risk/value narrative.
Review this chapter.
Challenge a claim, boundary, missing failure mode, unclear term, or unsupported evidence statement.
- Claim
- Boundary
- Failure
- Evidence