← All work
Practitioner case study · Multi-agent factory

Inside my multi-agent design operation

I personally design, run, and verify a multi-agent software factory: agents that build in parallel against a frozen contract, other agents that attack their work, and a fixer that answers every finding in writing. The proof is not a slide. It is a public, timestamped repo anyone can install in 60 seconds.

I run a team of AI agents that build, attack, and correct each other, and I verify the result myself

Role
Sr. Director, Product Design, AI-Native
Practice
Multi-agent software factory
Artifact
ds-canon, public and installable
The run
2026-08-04, one afternoon

An operating model, not a prompt

The orchestrator plans and verifies. Delegates build in parallel against a frozen contract. Adversarial challengers attack the code, the running server, and the docs. A fixer answers every finding in writing. Nothing merges on an agent's word alone.

Speed that came with more scrutiny

Under 48 minutes from git init to a deployed public repo, three builders in parallel, then 28 adversarial findings, a real server-freeze BLOCKER caught, and the test suite grown from 33 to 78 before anything shipped.

A public artifact, not a demo

The whole run produced ds-canon, an MCP server published to npm the same day, with the unedited factory log and every challenge and disposition committed to the repo so anyone can audit it.

The operating model

The thing I actually built is not the code. It is the way the code gets built, attacked, and verified when the workforce is a set of AI agents that will happily manufacture confidence.

The model has four moving parts and one rule. The orchestrator, which I direct, plans the work and verifies the result. Delegates build in parallel, each against a frozen contract written before any of them starts. Adversarial challengers, running different models than the builders, attack the finished work: the source code, the running server as a black box, and the documentation. A fixer answers every finding in writing, fixing or rejecting each one with a reason. The rule underneath all of it: nothing merges on an agent's word alone. Every line is written by one agent and attacked by a different one, and the orchestrator re-verifies the fixer independently rather than taking its word.

That is the discipline I care about. AI makes it cheap to produce work and cheap to sound sure. The only thing that keeps speed honest is making every claim survive an attack from something that did not write it.

The run, timed

On 2026-08-04 I ran this model end to end to build ds-canon, a small MCP server that serves a design system as a queryable source of truth. I let the numbers below stand as the case study because every one of them is traceable to the public repo and its timestamped factory log. Nothing here is reconstructed after the fact.

01

Parallel build on a frozen contract

Three builders worked at once, each in its own ownership zone. Two of them coded against a contract the third was still implementing. Their code typechecked and built clean on first integration, and the one real cross-builder defect that existed was caught immediately by an integration test, not by later debugging.

02

Adversarial attack from two directions

Two challengers filed 26 and 2 findings against a fully green test suite: one reading the source for correctness and contract honesty, one attacking the running server black-box with hostile inputs. Green tests hide things. Readers and attackers find them.

03

One real BLOCKER, caught and re-attacked

The black-box attacker found a genuine ship-stopper the whole suite missed: a single oversized not-found lookup froze the entire server through an unbounded fuzzy-match pass. It was fixed, then I re-ran the attack with a 2MB hostile payload. The server replied in 143 ms.

It answered all 28 findings individually, each marked FIXED, REJECTED, or DEFERRED with its reasoning, and grew the test suite from 33 to 78 tests in the process. The same attacker also certified the three README demo paths flawless, in writing. Under 48 minutes of wall clock ran from git init to a deployed public GitHub repo, and the package was published to npm the same day. A human team does not hit that number, and the interesting part is not the speed. It is that the speed arrived with more scrutiny, not less.

The spec behind the attack

The reason a challenger produces a real BLOCKER instead of a compliment is that I write the charter before I let one agent review another agent's work. It fixes the stance, the severity language, the evidence a finding has to carry, and the reply the fixer owes in return. This is the spec, paraphrased from the one that ran on 2026-08-04.

CHALLENGER CHARTER

STANCE      Adversarial. You did not write this code.
            Your job is to break it, not to praise it.

SEVERITY    BLOCKER  ship-stopper: crash, freeze, data loss, security
            MAJOR    fix before a hiring engineer sees it
            MINOR    a real defect, not urgent
            NIT      polish

EVIDENCE    Every finding carries, or it does not count:
            - file:line
            - a concrete failure scenario (inputs to wrong output)
            - the fix you would make

VERDICT     A written FIX FIRST list. Nothing merges on it alone.

FIXER REPLY  Every finding answered in writing, exactly one of:
            FIXED     what changed, where, and the new test
            REJECTED  why it is not a defect
            DEFERRED  why, and what it is traded against

The spec I write before letting one agent review another agent's work. The severity taxonomy and the evidence bar are what turn "looks fine" into a reproducible finding, and the reply contract is what keeps a fix from merging on assertion.

The broader practice

This run is one visible instance of how I work every day. The same operating model sits behind the rest of my practice, generically: a persistent memory vault so context survives across sessions, session logs so every decision has a trail, verification loops whose whole job is to catch fabrication before it reaches anyone, and a delegation economics where the cheapest capable model wins the task and the expensive one is spent only on planning, judgment, and synthesis.

The verification loops are the part I trust least by default and lean on most. In separate work the same discipline caught a fabricated quote and pulled it before an executive-bound export. That is the point of all of it. The agents are cheap and fast; the scrutiny is the deliverable.

Memory and logs

A persistent vault carries context between sessions, and append-only logs keep a trail of what was decided and why, so a run can be audited later rather than reconstructed from memory.

Verification over trust

Independent checks re-run the work rather than believing the agent that produced it. Fabrications get caught here, before they ship, because catching them is the loop's only job.

Side by side excerpt: challenger C2's BLOCKER finding about an unbounded fuzzy-match freeze, and fixer F1's written FIXED disposition describing the two-layer fix and the 143 ms re-attack result.

A real finding and its written disposition, excerpted from the factory run. Both documents are committed unedited in the public repo.

The factory is not a demo. It shipped a public artifact you can install in 60 seconds.

The agents are fast. The scrutiny is the deliverable.

The artifact

Everything above is verifiable because the run left a public trail. The code is on GitHub and on npm, and the factory log, the two challenger reports, and the fixer's line-by-line dispositions are committed inside the repo, unedited. You do not have to take my account of it. You can read the working artifacts of the agents challenging each other, and you can install the result yourself.

Model

Build, attack, fix, verify

An operating model where agents build in parallel against a frozen contract, other agents attack the result, a fixer answers in writing, and the orchestrator verifies rather than trusts.

Run

Fast, and more scrutinized

Under 48 minutes to a deployed repo, 28 adversarial findings dispositioned, one server-freeze BLOCKER caught and re-attacked, and the suite grown from 33 to 78 tests.

Proof

A repo, not a story

ds-canon shipped to GitHub and npm the same day, with every challenge and disposition public. The proof of the practice is a thing you can install, not a claim you have to believe.

Recognition

Peer-nominated, on the record

The practice this page describes earned Tebra's AI Disruptor IMPACT award for Q2 2026, announced publicly by the company in July 2026.