0% read on this device
Browse the curriculum

Start Here

Vision

First Principles

Operating Model

Domain Model

Agent Factory

Runtime Architecture

AI Engineering

Autonomous Workflows

Verification & Delivery

Factory Platform

Quality Engineering

Security & Governance

Case Studies

Labs

Interview Practice

Research Journal

Reference

Curriculum/Autonomous Workflows/Complete source chapter
Autonomous Workflows15 min readchapterQuick Read

Repository Onboarding and Codebase Intelligence

Establish the evidence required before an autonomous workflow may operate on a repository.

Status: Review readyRisk: highLifecycle: intent · plan · executeContent reviewed 2026-08-30Maturity guide →
Claim boundaryThis is curriculum guidance. It does not by itself prove a production implementation.
Quick Read

The chapter in one pass.

~3 min
  • Purpose: Establish the evidence required before an autonomous workflow may operate on a repository.
  • Best for: Platform teams, repository owners, security engineers, and agent engineers.
  • Prerequisites: Authoritative Delivery Hierarchy and Development Environments.
  • Reading time: 15 minutes.
  • You will learn: How to discover repository structure, ownership, instructions, dependencies, tests, data sensitivity, environments, and risk before admission.
  • Keep three ideas: registration is not readiness; indexes are derived and expiring; and uncertainty must narrow authority.
Whiteboard exercise

Reconstruct and defend this chapter’s architecture.

Reconstruct the architecture, name each boundary, and defend the tradeoffs.

orchestrationworkflow15 min chapter
Open the source exercise

Onboard a service with a shared schema repository, a generated client, an undocumented deployment job, and sensitive test data. Mark authoritative sources, derived indexes, owner decisions, blockers, and the first workflow you would safely admit.

1. The problem

An agent can clone a repository and still lack the information needed to change it safely. It may not know the authoritative build command, code owners, generated files, service dependencies, migration rules, test data boundaries, deployment path, or local instructions. Repository discovery performed independently during every run is slow and inconsistent. Stale discovery is worse: it creates confidence from facts that no longer match the target commit.

2. Why the problem exists

Repository knowledge is scattered across source, configuration, documentation, CI, deployment systems, package registries, ownership systems, and human memory. Monorepos contain several products with different rules. Multi-repository systems hide contracts and release order outside any single checkout. Some facts are safe to infer; others require an accountable owner.

3. Enduring Principle

Admit repositories through a readiness process

Onboarding creates a versioned Repository Readiness Record rather than a permanent “connected” flag. The record covers:

DimensionRequired understanding
Identity and ownershipCanonical repository, default branch, accountable owner, code owners, support contacts
InstructionsGoverning repository instructions, precedence, exceptions, generated-code rules
ArchitectureComponents, boundaries, entry points, data flows, external services, critical invariants
DependenciesPackages, services, schemas, repositories, runtime and release order
Build and testToolchains, setup, commands, test topology, fixtures, flaky suites, expected duration
DeliveryCI, artifacts, environments, deployment, feature flags, migrations, rollback
Security and dataClassification, secrets, network needs, licenses, sensitive paths, threat boundaries
Factory fitEligible workflows, tools, sandboxes, agents, budgets, verification, approval levels

Separate authoritative declarations from derived intelligence

Owner, data classification, permitted workflows, and release authority require declared sources. Symbols, call graphs, ownership suggestions, test impact, and architecture summaries may be derived. Every derived view records source commit, method, coverage, confidence, and expiry.

Build a codebase intelligence pipeline

Useful indexes include lexical and symbol search, dependency and ownership graphs, build targets, test-to-code mapping, API and schema inventories, historical change hotspots, incidents, architecture decisions, and documentation. Retrieval must preserve source and commit lineage.

Let uncertainty reduce scope

Missing owners, nonreproducible builds, unknown deployment paths, unclassified data, or absent tests should block high-risk autonomous change. The repository may still be eligible for read-only analysis or documentation proposals. Readiness is granular by workflow and risk class.

4. Tradeoffs and alternatives

Deep onboarding costs time and becomes stale. Incremental discovery tied to changed areas reduces cost but must not skip critical global controls. Human-authored architecture is more intentional; generated maps are more current. Preserve both and surface conflicts.

Embedding every repository can improve semantic search and create privacy, cost, and freshness problems. Use hybrid retrieval only where evaluations show it improves target tasks.

5. Current Mission Control Implementation

The current domain includes repository registration, configuration, workspace manifests, multi-repository coordination, environments, preflight, policy, and context packages. These establish important authority and runtime boundaries.

The published curriculum does not yet demonstrate a complete repository onboarding pipeline, readiness record, codebase indexing lifecycle, owner attestation, drift detection, or workflow-specific admission based on discovered capabilities. This chapter defines that missing front door.

6. Future Vision

Adding a repository should launch a read-only discovery workflow that produces an explainable readiness packet. Owners approve material facts and choose eligible workflows. Subsequent commits trigger targeted refresh. Before each WorkOrder, preflight verifies that required readiness evidence remains fresh for the affected scope.

7. Versioned references

8. Notes and lessons learned

Repository onboarding is not administrative setup. It is the first assurance case: evidence that the factory understands enough of the target to grant a particular kind of execution authority.

9. Interview and discussion questions

  1. Which repository facts may be inferred, and which require an owner?
  2. How should readiness expire?
  3. What blocks code modification but permits read-only analysis?
  4. How should conflicting instructions be resolved?
  5. What changes when one product spans several repositories?

10. Whiteboard exercise

Onboard a service with a shared schema repository, a generated client, an undocumented deployment job, and sensitive test data. Mark authoritative sources, derived indexes, owner decisions, blockers, and the first workflow you would safely admit.

11. Hands-on lab

Run read-only discovery against a disposable repository. Produce a readiness record, instruction hierarchy, component map, dependency graph, build/test inventory, data classification questions, and eligible-workflow recommendation. Invalidate one source commit and prove the readiness status becomes stale rather than silently remaining approved.

External review

Review this chapter.

Challenge a claim, boundary, missing failure mode, unclear term, or unsupported evidence statement.

  • Claim
  • Boundary
  • Failure
  • Evidence