Governed Continuous Learning and Recursive Improvement
A factory that never learns repeats failures and requires permanent manual tuning. A factory that changes its prompts, policies, workflows, evaluations, or authority automatically can become unpredictable. Continuous learning must improve t
A rapid review of the chapter’s existing Quick Read, principles, definitions, lessons, and review material.
Reconstruct and defend this chapter’s architecture.
Reconstruct the architecture, name each boundary, and defend the tradeoffs.
Open the source exercise
Draw a failed production change becoming a deduplicated suggestion, evaluation, governed WorkOrder, canary, and promoted workflow rule. Add malicious source content, evaluator disagreement, and a quality regression.
3. Enduring Principle
Automate observation and proposal; govern promotion
The factory may continuously collect, normalize, deduplicate, and analyze signals. It may propose a prompt, skill, workflow, policy, evaluator, model route, or architectural change. Promotion to active behavior requires explicit human review based on independent evaluation.
Keep the learning loop inside the delivery hierarchy
An accepted recommendation becomes a governed Mission or WorkOrder with scope, criteria, budget, risk, and owner. It does not mutate production configuration directly. The same execution, validation, evidence, and release rules apply to factory self-improvement as to customer software.
Separate three loops
Inner loop: implement and test one change.
Outer loop: validate, review, release, and observe the outcome.
Meta loop: detect patterns across outcomes and propose changes to the factory system.
The meta loop has greater leverage and therefore requires stronger evidence and promotion control.
Treat research and memory as untrusted inputs
Every source needs identity, retrieval time, content hash, classification, license or usage constraint, sensitivity, and provenance. Extracted claims need supporting evidence, confidence, contradictions, and lifecycle. Source content cannot change instructions, invoke tools, or grant authority.
Evaluate against baselines and quality floors
An improvement experiment defines baseline, candidate, comparable cases, primary metric, quality floor, risk stop, budget, observation window, and rollback. Faster or cheaper is not improvement when validation, security, reliability, or human attention worsens.
Calibrate autonomy from outcomes
Sustained validated outcomes may make a factory eligible for greater autonomy. Critical violations, fabricated evidence, security escapes, or repeated failure should demote or quarantine it automatically. Promotion remains human-owned.
8. Notes and lessons learned
Recursive improvement should be ordinary governed engineering applied to the factory itself. The recursive object changes; accountability does not.
9. Interview and discussion questions
- What may the factory learn automatically?
- Why must promotion remain human-owned?
- How do inner, outer, and meta loops differ?
- How do you prevent memory or research poisoning?
- What evidence would justify an autonomy increase?
- When should learning cause immediate rollback or quarantine?
Review this chapter.
Challenge a claim, boundary, missing failure mode, unclear term, or unsupported evidence statement.
- Claim
- Boundary
- Failure
- Evidence