0% read on this device
Browse the curriculum

Start Here

Vision

First Principles

Operating Model

Domain Model

Agent Factory

Runtime Architecture

AI Engineering

Autonomous Workflows

Verification & Delivery

Factory Platform

Quality Engineering

Security & Governance

Case Studies

Labs

Interview Practice

Research Journal

Reference

Curriculum/Operating Model/A rapid review of the chapter’s existing Quick Read, principles, definitions, lessons, and review material.
Operating Model7 min readchapter

Factory Economics and Operating Metrics

Agent activity is easy to measure and easy to mistake for value. Tokens, sessions, generated lines, tool calls, and pull request volume can all rise while customer outcomes slow, defects increase, and engineers spend more time recovering or

Status: Draft for studyRisk: highLifecycle: intent · plan · verify · learnContent reviewed 2026-08-31Maturity guide →
Claim boundaryThis chapter references implementation evidence. Inspect its evidence boundary before treating a claim as proven.
study mode

A rapid review of the chapter’s existing Quick Read, principles, definitions, lessons, and review material.

Whiteboard exercise

Reconstruct and defend this chapter’s architecture.

Reconstruct the architecture, name each boundary, and defend the tradeoffs.

governancehuman authority7 min chapter
Open the source exercise

Draw a metric tree from customer outcome down to delivery, factory, and runtime diagnostics. Add a team that doubles PR throughput while review time and change failure rise. Explain why the factory has not improved and which bottleneck to address.

3. Enduring Principle

Measure validated customer value, not activity

The primary lead-time clock starts when business intent becomes a governed Mission. It stops when the change is deployed, independently validated in production or a production-equivalent environment, and the expected customer outcome is confirmed. Merge time is an intermediate measure, not the outcome.

The first three executive measures are:

  1. Lead Time to Validated Customer Value — elapsed time from governed intent to confirmed outcome.
  2. Change Failure Rate — the proportion of deployments that cause rollback, hotfix, emergency intervention, customer regression, reliability or security incident, or SLA/SLO violation within a default seven-day observation window.
  3. Engineering Leverage — valuable outcomes per unit of scarce engineering capacity without increasing cognitive load or coordination cost.

These form a constraint system. Speed without quality is rework. Quality without speed is delay. Throughput without human sustainability is hidden debt.

Build a metric hierarchy

Business outcomes: adoption, revenue, retention, risk reduction, or customer problem solved.

Delivery outcomes: lead time, deployment frequency, change failure, recovery time, accepted WorkOrders, and outcome confirmation.

Factory effectiveness: autonomous completion, first-pass validation, recovery success, evidence completeness, approval latency, review time, and cost per accepted outcome.

Operational diagnostics: tokens, tool calls, model latency, queue depth, lease expiry, retries, provider errors, and context size.

Diagnostic metrics explain outcomes. They are not the outcomes.

Define engineering leverage carefully

No single number proves leverage. Use a balanced evidence set:

  • reduced lead time with stable or improved change failure rate;
  • increased throughput of accepted, validated work;
  • reduced human implementation hours per work item;
  • reduced waiting and coordination time;
  • more time spent on architecture, product, and customer problems;
  • stable or lower review and recovery burden; and
  • improved developer satisfaction and perceived control.

The objective is more customer value per engineer, not more commits per engineer.

Measure flow and attention

Break lead time into queue, planning, approval, execution, validation, review, deployment, and outcome-observation time. This reveals whether the factory accelerates work or moves the bottleneck.

Human attention is a constrained resource. Track number of interventions, decision latency, time per approval, evidence inspection time, false alarms, and repeated requests. An autonomy system that consumes more senior attention than it returns has negative leverage.

Attribute full cost

Cost per accepted WorkOrder should include model and token spend, tools, infrastructure, CI, storage, human implementation, review, recovery, rework, incidents, and allocated platform operation. Estimate uncertainty explicitly.

Compare marginal cost and marginal value. A more expensive model can be cheaper overall if it reduces retries and review. A cheaper run that fails validation is inventory, not value.

Use cohorts and baselines

Compare similar repositories, risk bands, change types, and autonomy levels. Establish a stable pre-factory baseline and use medians and percentiles rather than averages alone. Avoid causal claims from simultaneous organizational, tooling, and product changes without an experimental design.

8. Notes and lessons learned

Mission Control now has materially more than activity counters: it retains governed lineage and exposes provenance-aware effectiveness projections. The mastery challenge is to preserve that distinction. Current operational records support flow and diagnostic decisions; complete factory ROI still requires sustained intent-to-production outcome evidence and full cost attribution.

9. Interview and discussion questions

  1. When exactly does lead time start and stop?
  2. Why is time to merge insufficient?
  3. How do you define a change failure?
  4. What proves engineering leverage without surveilling developers?
  5. How would you calculate cost per accepted WorkOrder?
  6. Which current Mission Control metrics are proxies?
  7. How would you establish causal confidence in a factory rollout?
Evidence boundary

Curriculum maturity is not implementation proof.

This chapter defines architecture or practice. It does not by itself prove a corresponding production implementation.

CurriculumDraft for studyImplementation evidenceNot asserted hereInspect evidence map →
External review

Review this chapter.

Challenge a claim, boundary, missing failure mode, unclear term, or unsupported evidence statement.

  • Claim
  • Boundary
  • Failure
  • Evidence