Evaluation Science and Controlled Experimentation
Add experimental rigor to repeatable agent evaluation and promotion decisions.
A focused view of boundaries, contracts, state, authority, failure paths, and tradeoffs drawn from this chapter.
Reconstruct and defend this chapter’s architecture.
Reconstruct the architecture, name each boundary, and defend the tradeoffs.
Open the source exercise
Design an experiment comparing two agent configurations on 200 tasks across three risk classes. Add grader disagreement, a security hard-gate failure, cost improvement, and one segment regression. Make the promotion decision.
4. Tradeoffs and alternatives
Large datasets increase coverage and maintenance. Small curated sets are explainable and easier to overfit. Human judgment captures nuance and costs time. Model graders scale and share failure modes with evaluated systems. Online experiments improve realism and require strict user, data, and risk controls.
5. Current Mission Control Implementation
The current curriculum covers representative cohorts, baselines and candidates, trace replay, criterion-level receipts, canaries, model routing, evaluation datasets, and promotion. It does not yet specify contamination controls, grader calibration, agreement, repeated-trial analysis, statistical decision rules, shadow experiments, or a full adversarial-evaluation program.
Review this chapter.
Challenge a claim, boundary, missing failure mode, unclear term, or unsupported evidence statement.
- Claim
- Boundary
- Failure
- Evidence