The AI-SDLC loop, layer by layer.

One loop, four depths. Each control adds a layer: the idea, the platform harness, the mechanics, then the full engineering view.

Agents draft. People decide. Agents draft. People decide. The platform enforces. The AI-SDLC: one loop that work moves around, over and over. The same loop, now showing the rails a platform must provide to make the harness real. 1 The plan People pick the goal and set the guardrails. People set the outcome and the constraints. Every funded bet carries kill criteria. A signal never becomes committed work automatically. Scope locks before agents draft. 2 Agents draft Agents draft in a sandbox AIAIAI Agents write and test the change in a safe practice space. Repo → CI → artifact registry Ephemeral preview + isolated test data API / CLI paths; no console-only steps Sandbox: no production write, deploy, or secret-read capability. Capability is restricted before review. 3 A person decides A human reviews the work. Nothing ships without a yes. Review the diff + evidence manifest. Nothing reaches production without an explicit human authorization. Manifest binds: image digest · contract digest · IaC revision · test-data hash Self-approval is structurally forbidden: requester and authorizer must differ. LIVE LIVE, gated The change goes live, carefully. Real users, real stakes. reserved-tenant canary · real users, real stakes human promotion · deterministic rollout controller 4 Learn What happened feeds the next plan. Traces follow every change into production and back. Privacy-safe signals feed the plan and tighten the rails. Deploy markers · per-route SLOs Learning must change a durable record, not end as a retro. One loop. One loop. Two layers. The moving work is the AI-SDLC: plan → draft → decide → live → learn. The platform rails are the harness: access, proof, and control around every step. ACCESS Workload identity, short-lived credentials, least privilege PROOF Contract + policy tests in CI Digest-bound evidence manifest CONTROL Infra + policy as code Agents: zero prod capability THE RAILS RUN THE WHOLE WAY They are not a checkpoint added at the end. An agent stays inside identity, test, policy, and audit boundaries the entire time. What is an AI-SDLC? A way of building software where AI agents do most of the drafting, people make the decisions, and the system itself enforces the rules, so speed never outruns safety. What is the harness? The machine that makes that safe: the encoded rules, guardrails, tests, and locks that let agents do real work without being able to break things. The harness is the AI-SDLC made physical. What is agentic engineering? Engineers direct AI agents that do the hands-on work: setting intent, steering, reviewing, and owning the final approval. Think driver's ed: the learner (the AI) really drives, the instructor (a person) has their own brake. Here, the instructor never leaves. What moves through the loop? One change at a time: a PR-sized unit of work with its own preview, its own evidence manifest, and its own decision. The loop governs the change. The operating model above it governs the organization: pods, control plane, authority. What must the platform provide? Self-service projects, previews, and CI a team runs alone Proof gates, observability, canary + reconciliation Policy-enforced production controls, captured as code Assembled from managed primitives; custom only where nothing native fits. What must never be lost? Self-service deploy; fast ship → observe → reverse Branch previews and a fast local loop Agent-drivable API / CLI workflows end to end Parity is measured, not declared: preview, production, rollback, and review speed. PROMOTION MECHANICS · ONE CHANGE, END TO END Draft in sandbox CI proof gates Evidence manifest Human authorization Canary Observe Learning recorded Self-approval is forbidden: the requester and the authorizer must be different identities, enforced by the platform, not by a policy memo. Autonomy is granted per action, from observe-only to execute-and-promote, and no level can change its own governing policy. KEEPING GATES HONEST rejection rate pairs with material change rate · ground truth is escaped defects per surface · start with three instrumented gates, not twenty An educational companion to the case study above. Every mechanism shown here runs in the reference implementation: one synthetic case, one product surface, one signal end to end, on synthetic data only.
The harness, one level down. The engineering view: the components a harness is assembled from, and what a managed platform provides versus what the team must build. Each component lists what is native (a managed primitive exists to build on) and what the team must build on top of it. The enforcement layer is always a build. THE AGENT PATH, MECHANICALLY 1 · Plan Humans set the outcome, the owners, and the constraints before any agent drafts. Every funded bet carries a baseline and kill criteria. A signal never becomes committed work automatically. 2 · Draft Agents run the full issue-to-PR loop in an ephemeral preview with an isolated database per branch. Sandbox: no production write, deploy, or secret-read capability. Repo → CI → artifact registry. Every path is API / CLI; nothing depends on a console click. 3 · Decide A human reviews the diff plus a condensed evidence manifest bound to: image digest · contract digest · IaC revision · environment · test-data hash Agents never deploy to production. Requester and authorizer must be different identities, checked structurally at the gate. 4 · Live Only a deterministic rollout controller holds the narrow production capability. Human promotion, then a reserved-tenant canary: real users, real stakes. Agents may only request pause or rollback. 5 · Learn Traces follow every change across runtimes, privacy-safe by policy. Deploy markers, per-route SLOs and alerts. Signals feed the next plan and tighten the rails: a durable record changes, or nothing was learned. What runs on it Any product workload with real stakes. The reference implementation runs exactly one: a fictional organization, a single product surface, one signal end to end. The harness is deliberately product-agnostic: the loop governs the change, whatever the change happens to be. HARNESS COMPONENTS · WHAT YOU BUILD VS WHAT A MANAGED PLATFORM GIVES YOU Identity + access Workload identity · short-lived credentials · least privilege Build:a delegated actor context, per-hop reauthorization, machine-to-machine scopes and actor types. Native:the workload-identity primitive only. The enforcement layer has nothing native. Secrets Central secret-manager integration Build:least-privilege wiring per service and per agent, privacy-safe policy patterns, no agent secret-read in production. Native:the secret-manager primitive. Previews An ephemeral environment per branch Build:the entire preview surface: isolated database or schema per preview, seeded and masked data, a local loop that matches CI. Native:little that is turnkey. Previews are usually the biggest build. CI/CD + templates Pipelines · artifact registry · templates Build:golden-path service templates and one-command or one-PR deploy paths that agents can invoke themselves. Native:the artifact-registry primitive and pipeline runners. Proof gates Deterministic CI: unit, integration, contract, performance, and security checks Build:CI that fails early on contract drift, missing tenant predicates, or sensitive data in telemetry. A digest-bound manifest per change. Open frontier: eval-in-CI for agent changes. Observability Logs, metrics, distributed traces, correlation IDs Build:instrumentation is never inherited; it is a build. Privacy-safe by policy, deploy markers, per-route SLOs, error-to-issue automation. Native:the primitives and dashboards. Canary + reconciliation Progressive rollout, traffic splitting, rollback Build:the deterministic rollout controller, reconciliation with checksums, and the reserved-tenant canary wiring. Native:managed canary mechanics; the controller and its evidence trail are yours. Policy + IaC Infrastructure as code · policy as code Build:domain guardrails enforced by construction, not inspection. Every production change captured as code. Native:org-level policy engines; the domain rules are always yours. WHAT KEEPS A GATE HONEST The tripwire pair Gate rejection rate is gameable at least five ways: split the change, reject cosmetically, self-reject, gate-shop, or submit a deliberately weak first draft. It means something only when paired with material change rate, and the ground truth is escaped defects per surface. 30% rejection with 5% material change is theater with extra steps. Start with three instrumented gates, not twenty: merge, design-system conformance, data-model change. The autonomy ladder Autonomy is granted per action and per surface: observe-only → suggest → execute in sandbox → execute-and-promote. Freedom is earned where an agent has proven safe, and revoked after drift. A ratchet that only turns one way is not governance. No autonomy level can change its own governing policy. The failure path is deny-and-continue with a recovery route, never a dead end. Engineering companion to "Agents draft. People decide. The platform enforces." Every mechanism described here is exercised by the reference implementation on synthetic data. Nothing on this board is a deployment claim; it is the operating model made teachable, executable, and inspectable.