0% read on this device
Browse the curriculum

Start Here

Vision

First Principles

Operating Model

Domain Model

Agent Factory

Runtime Architecture

AI Engineering

Autonomous Workflows

Verification & Delivery

Factory Platform

Quality Engineering

Security & Governance

Case Studies

Labs

Interview Practice

Research Journal

Reference

Curriculum/AI Engineering/A focused view of boundaries, contracts, state, authority, failure paths, and tradeoffs drawn from this chapter.
AI Engineering16 min readchapterQuick Read

Evaluation Science and Controlled Experimentation

Add experimental rigor to repeatable agent evaluation and promotion decisions.

Status: Review readyRisk: highLifecycle: verify · learnContent reviewed 2026-08-30Maturity guide →
Claim boundaryThis is curriculum guidance. It does not by itself prove a production implementation.
architecture mode

A focused view of boundaries, contracts, state, authority, failure paths, and tradeoffs drawn from this chapter.

Whiteboard exercise

Reconstruct and defend this chapter’s architecture.

Reconstruct the architecture, name each boundary, and defend the tradeoffs.

agent runtimemodelscontext16 min chapter
Open the source exercise

Design an experiment comparing two agent configurations on 200 tasks across three risk classes. Add grader disagreement, a security hard-gate failure, cost improvement, and one segment regression. Make the promotion decision.

4. Tradeoffs and alternatives

Large datasets increase coverage and maintenance. Small curated sets are explainable and easier to overfit. Human judgment captures nuance and costs time. Model graders scale and share failure modes with evaluated systems. Online experiments improve realism and require strict user, data, and risk controls.

5. Current Mission Control Implementation

The current curriculum covers representative cohorts, baselines and candidates, trace replay, criterion-level receipts, canaries, model routing, evaluation datasets, and promotion. It does not yet specify contamination controls, grader calibration, agreement, repeated-trial analysis, statistical decision rules, shadow experiments, or a full adversarial-evaluation program.

External review

Review this chapter.

Challenge a claim, boundary, missing failure mode, unclear term, or unsupported evidence statement.

  • Claim
  • Boundary
  • Failure
  • Evidence