The way of working is the hard part
When agents draft most of the work, production stops being the bottleneck and trustworthy judgment becomes the scarce resource. I designed the operating model that protects it, expressed in one sentence and proven three times: agents draft, people decide, the platform enforces.

Agents draft. People decide. The platform enforces.
The method
An AI-native SDLC where the gates are structural, earned, and revocable, never advisory.
The model
Outcome pods own outcomes. A thin control plane enforces. Agile drops to a compatibility layer.
The proof
A running reference where an invalid release physically cannot advance.
Leadership asked how fast agents could take the team. I argued that speed was the wrong deliverable.
Once agents can draft most of the work, the reflex question is how much faster everything can go. My position was the opposite, and it became the spine of everything that followed: the speed is the known, easy part. The hard part is building the instrumentation and the way of working that let a small team use that speed safely.
Producing more output per week is a solved problem. Building the conditions under which agents can do most of the building, without quietly outrunning accountability, is not. That is the real deliverable, and nobody hands it to you.
We stopped citing Replit and started citing ourselves
The move that earns trust is turning the method on your own work before anyone else has to.
Every cautionary deck cites the same tale: an agent deleted a production database during a code freeze, because the freeze was written as an instruction rather than enforced as a rule. It is a good story, and it is someone else's. So I ran an adversarial review against my own repository. Two independently prompted agents, briefed to attack the work rather than defend it, found that a trust gate I believed was enforced was only advisory. The validator returned findings tagged block, but the sign path enforced nothing, the button was never disabled, and no override was recorded. The gate meant to catch exactly this had sat unsigned for five weeks. Finding it in my own code is the proof that the method works: adversarial review against real evidence catches what self-assessment cannot.

The argument went through six rounds of draft, challenge, and harden. Across the fact-checking pass, 38 external sources were checked and all 38 held, 31 quoted passages were verified word for word, and one fabricated quote was caught and removed before anything shipped.
A gate is only real if the system enforces it. Five rules hold the loop up.
Restrict capability before you rely on review
Review is a probabilistic filter with a known miss rate. Capability restriction is deterministic. Allowlists, protected paths, and hard-disabled auto-merge come first.
A gate must be structural, not advisory
Hooks over instructions, CI checks over checklist items, disabled buttons over warning labels.
A gate that is never rejected is not a gate
If nothing is ever turned back, the gate is a ritual. Users approve most permission prompts, so an unenforced prompt is not protection.
Autonomy is earned per surface and revocable
A ratchet that only turns one way is not governance. Agents earn freedom where they have proven safe, and lose it after drift.
The failure path must be recoverable
Deny and continue with a recovery route, never a dead end.
The loop, explorable end to end
The model is a working artifact, not a diagram of good intentions.
A live briefing, embedded in full: move through the loop, the platform harness, and the mechanics one level down.
It is easy to install a gate that produces the feeling of safety and none of the substance.
Gate rejection rate is a tripwire, never a target and never a dashboard number, because it is gameable at least five ways: split the change, reject cosmetically, self-reject, gate-shop, or submit a deliberately weak first draft. It means something only when paired with material change rate. A gate with 30% rejection and 5% material change is theater with extra steps, and the ground truth is escaped defects per surface. Start with three instrumented gates, not twenty: merge, design-system conformance, data-model change.
When artifacts are cheap, a role is defined by the gate it owns, not the artifact it produces.
Design
Guards the concept gate, design-system conformance, and trust design for the surfaces where AI drafts and people approve.
Product
Owns the problem gate, the scope lock, and the kill decision.
Research
Owns the evidence chain, protecting provenance from source to synthesis.
Every gate
Names the artifact that proves it ran. A row with no evidence artifact is the row that fails an audit.
One loop. Three layers. Explicit authority.
The loop governs one change. The operating model governs the whole organization.

Outcome pods are the operating model
Small cross-functional human teams own a measurable outcome from evidence through learning, not a backlog of features.
A thin control plane enables it
A link-and-verify layer across the tools you already have. It preserves evidence, decisions, policies, and learning, and structurally blocks incomplete high-risk work while keeping human authority explicit.
Agile becomes a compatibility layer
Roadmaps and sprints stay where they help the surrounding org coordinate, but they no longer define how work actually moves.
Underneath sits a loop from sense to learn, where each stage carries an explicit gate: a signal never becomes committed work automatically, a funded bet always carries a baseline and kill criteria, delivery completion is not outcome completion, and learning must change a durable record rather than end as a retro. Autonomy is granted per action, from observe-only to execute-and-promote, and no level can change its own governing policy.
The memorable moment is a release that cannot move
A model on a page persuades no one who has seen enough models on pages. So I built it.
The reference implementation runs a single synthetic case, a fictional organization called Example Health and one product surface. It follows one signal through five chapters and lets you inspect it at four depths of evidence. Every operational fact on screen is recomputed from a canonical, hash-chained ledger, not typed into a mockup. A proposing agent submits a release and approves its own work. The system blocks it, structurally, and names exactly why. Then a different human, a separate identity from the requester, authorizes it correctly, a simulated canary runs, and the system records a durable learning change.


The demo prefers a caught failure over a fabricated success metric. That preference is the design.
A gate is only real when the system can stop the work.
The reference was verified the way it preaches.
Every workstream close passed through a named review director and two fresh, independent challengers working from the source read-only, with a binary pass or fail and any serious finding forcing a correction and a full rerun. A passing test suite did not count as a pass. The gate exists because it once caught us: an early workstream was closed on a generic review, and the belated director-led challenge found two real authority defects it had missed. The record shows the gate biting repeatedly, with multiple workstreams failing a first challenge and clearing only after correction. On the boundary that matters most, the suite removed each controlling field from an authorized release packet in turn and confirmed the system rejected all 2,356 mutation attempts. Rigor here is not a claim. It is a reproducible event log.
The honesty is not a disclaimer at the bottom. It's a design element carried at the record level.
Every fact wears a provenance label: real and local when it is recomputed from actual repository bytes, synthetic for the fictional case data, simulated for the promotion, canary, and observation. Nothing is ever labeled real and public, because nothing here touches a real customer, a clinical decision, or a production system. This is an independent reference implementation, not a deployment, a clinical system, a compliance attestation, or evidence of a real-world result. It exists to make the operating model teachable, executable, and inspectable, and to earn the one thing that comes next: a measured production pilot that succeeds only if it proves a better decision-and-learning system. A faster build that produces rubber-stamp review is a failed pilot.
Agents took the artifacts. What they handed back is a sharper version of the jobs we always claimed to have.
When production is cheap, the work that remains is the work that was always the point: choosing the right outcome, designing the gate that protects the decision, and owning the standard of proof. That is not a smaller job for a design leader. It's the whole job, finally made explicit.