The AI Software Factory Guide
On this page9 sections
How to design, build, prove, operate, and improve an engineering system in which humans define intent and accept risk while bounded agents plan, implement, validate, and recover — and independent evidence decides what advances.
Intent → Plan → Define Agent → Execute through Harness → Apply Skills → Evaluate → Improve → Deliver Software
This eight-stage value stream is the primary reader model. The supporting architecture has six areas: Intent, Harness, Capability, Model, Trust, and Learning, surrounded by adoption. Chapter 2 teaches how the two fit together.
Read the guide front to back, or enter at the part that matches your question.
The factory in one line, stage by stage
Click a stage for a concise contract brief, then follow its links to the canonical chapters for technical depth.
- Stage 1 · Builder Intent
- Stage 2 · Plan
- Stage 3 · Define Agent
- Stage 4 · Execute through Harness
- Stage 5 · Apply Skills
- Stage 6 · Evaluate
- Stage 7 · Improve
- Stage 8 · Deliver Software
Front matter
- How to read this guide
- What this guide covers — the coverage map: stack boundaries, vocabulary areas, capability areas by priority, and where each lives
Part I — Understand
- Why software engineering is changing
- The factory in one view
- First principles: trust, evidence, and authority
Part II — Design
- The human–agent operating model
- Authoritative records: from company to release
- Intent and specification engineering
- Governance, policy, and risk-proportional approval
- Economics, metrics, and human attention
- Tokenomics and factory economics
- Multi-repository design and coordinated delivery
Part III — Build
- The Agent Factory
- Skills as packages
- Control plane, orchestrator, and execution plane
- Durable execution: tasks, attempts, leases, and recovery
- Coding harnesses and agent protocols
- Harness engineering
- Development environments, sandboxes, and compute
- Agent architecture: loop, MCP, tools, context, and memory
- Data, knowledge, and semantic engineering
- Context engineering
- Models and capability selection
- Routing and the escalation ladder
- Agent and loop engineering
- Loop engineering patterns and defaults
- The 12-layer production AI agent stack
- Autonomous engineering workflows
Part IV — Prove
- Quality and evidence architecture
- Testing strategy for agentic change
- Evaluation engineering
- Evals as factory assets
- Quality contracts, proof packages, and certificates
- CI/CD, progressive delivery, and production verification
- Security: identity, secrets, threats, and supply chain
Part V — Operate
- The factory as a platform
- Observability, telemetry, and forensics
- Resilience, incidents, and the control tower
- Control surfaces, event contracts, and storage
- Enterprise adoption and the infrastructure landscape
Part VI — Improve
- Production feedback, automated review, and the agentic merge queue
- Governed learning
- Meta-loops and the closed-loop factory
- Mission Control as a living case study
- Mastering the factory
- Where this is going
Appendices (reference)
- A. Canonical glossary
- B. Mission Control case studies
- C. Research canon
- D. Coverage and maturity · Changelog · Reviewer guide
- E. Software architecture and system design study guide
- F. Principles to have cold
- G. Mission Control operator surfaces
The v1 curriculum chapters are preserved unchanged in
archive/guide-v1/.