← All writing
AI-native

Inside the agent factory: why I don't trust green tests

Green tests measure what you imagined. Adversarial review hunts what you did not. Those are different jobs, and the interesting bugs live in the gap.

Inside the agent factory abstract editorial cover showing parallel builders, adversarial challengers, and a review gate

A test suite that is all green after the first build is not evidence that the code is correct. It is evidence that the code does what the person who wrote the tests expected it to do. Those are not the same claim.

I got a direct demonstration of that gap on August 4, 2026, the same afternoon I built ds-canon, a design-system-of-record MCP server. This is the making-of story: how the thing got built, tested, attacked, and fixed in one sitting, with the actual findings and timestamps from the build log.

Freeze the contract before anyone writes code

The run started at 14:43:43 with scaffolding and a package install, done by 14:44:47. Before any feature code got written, I froze two things: the TypeScript type contract, src/types.ts, that every downstream piece of code would treat as non-negotiable, and a fixture spec pinning the exact entities the demo and test suite would depend on: 40 tokens, 9 components, space.inset.md consumed by exactly Button, Card, Field, and Modal, color.accent.legacy consumed by exactly Banner and LegacyButton, and so on.

That sounds like overhead. It is the opposite. Three separate agents were about to write code in parallel with no communication between them. The only thing that would let their work integrate on the first try was agreeing, in writing, on the shape of the data before any of them started.

Three builders, one afternoon, zero cross-talk

Builders launched at 14:45:55, all three running in separate terminal panes:

  • B1 built the fixtures and the loader that validates and indexes them.
  • B2 built the eight tools and the MCP server that exposes them.
  • B3 built the test suite, writing tests against the frozen contract before B1's and B2's code had even run together.

B2 and B3 were coding against a contract, not against each other's actual output. That is the bet: if the contract is precise enough, three agents can build in parallel and their work will fit without a coordination meeting.

Integration started at 14:51:22: typecheck, build, then run B3's suite against B1 and B2's code for the first time. It went green at 14:52:31, nine minutes after the builders launched, but not on the first try blind. B3's integration test caught a real cross-builder defect immediately: the typed not_found result B2's tools returned violated the declared output schema at the SDK validation layer. One seam between two builders' work, one fix, suite green.

Nine minutes to green is fast. It is also the exact moment green tests stop telling you anything new, because the tests that just went green are the tests B3 thought to write. What they do not know they did not cover does not show up as a failure. It does not show up at all.

Two attackers, two different targets

Challengers launched at 14:52:42, in parallel, each explicitly chartered to be adversarial:

  • C1 attacked the source code: correctness, contract fidelity, protocol discipline, and whether the repo would hold up to a staff-level reviewer's read.
  • C2 attacked the running server as a black box: hostile inputs, concurrency, protocol robustness, and whether the README's three demo queries were flawless.

Both finished at 15:00:08, seven and a half minutes later. Across the two challengers, they filed 28 findings total against a fully green 33-test suite. Of those, 1 was a BLOCKER and 6 were MAJOR.

C2 found the BLOCKER: an unbounded fuzzy-match routine that could freeze the whole server on one oversized lookup. get_token, get_component, and find_usages all fall back to a suggestion when a name does not match anything. The implementation compared edit distance against every candidate name with nothing capping input length.

C2 measured it against the real server: 50,000 characters took 374ms, 200,000 took 1,370ms, and 500,000 took 3,490ms. At 10MB, the MCP client's 60-second timeout fired before the server responded. Because the server runs on one Node event loop, the work did not just slow one request. It blocked every call on the connection.

The green suite never touched this because nobody had written a test that sent an oversized name. That is not a gap in diligence. It is the structural limit of tests: they check the cases someone thought of.

Six MAJOR findings hiding under green

C1 filed six findings at MAJOR severity, each a case where the implementation quietly disagreed with its contract or promise while the suite kept passing:

  • dependentCount counted deprecated dependents alongside active ones, even though the contract and tool description both promised active entities.
  • find_usages was only one hop deep, so a token consumed through an alias chain would not appear as a dependent of the token it ultimately resolves to.
  • check_token_drift resolved exact-value ties by load order. The fixture used 16px for both spacing and type, so a font size could be “fixed” toward a spacing token.
  • The declared and documented lang parameter was validated and then never read. CSS, JSX, and text produced byte-identical output.
  • The loader's validation was shallower than the schemas enforced downstream, so malformed fixtures could load and then fail with an opaque SDK error on first use.
  • The loader's roughly fifteen error branches had exactly one branch under test.

These are the class of defects a green suite can miss: not “the code crashes,” but “the code does something subtly different from what it claims to do,” where the claim and behavior only diverge in a case nobody's test constructed.

Answer every finding, then verify independently

The fix pass launched at 15:00:21 and closed at 15:19:47, run by one fixer working from both challenge reports. Every one of the 28 findings received a written disposition, FIXED, REJECTED, or DEFERRED, with a reason in factory/challenges/fix-log.md. The fixer could not quietly pick the convenient subset.

That freeze got defense in depth. The three tools' input schemas now cap name or entity at 256 characters, rejecting oversized input at the protocol boundary, and the fuzzy-match function has an independent length guard as a backstop. Other inputs were deliberately left uncapped because C2 had measured them at scale and confirmed they were linear and safe.

I did not take the fixer's word for it. The build, full suite, and original freeze attack were re-run independently with a 2MB hostile payload. It came back in 143 milliseconds. The test suite grew from 33 to 78 tests as the findings were answered, because every fixed finding needed proof that it stayed fixed.

Documentation got challenged too

Once the code was fixed, a docs writer produced the README and contributing guide from behavior it had observed by running the server. A third challenger followed every command literally and read the whole thing again for voice and unsupported claims. That phase ran from 15:20:09 to 15:28:19. Documentation matching the software is not guaranteed just because its author understands the software.

The number that matters is 28

Git init to a deployed repository took under 48 minutes. That is fast, but it is not the number I would point to. The number I would point to is 28: the findings a fully green suite was sitting on top of, including 1 BLOCKER that could freeze the server and 6 MAJOR disagreements between implementation and contract.

Every finding came from an agent running a different review job, with the explicit task of trying to break the source or the server. Every disagreement remains unedited in the repo's factory/ directory, because the argument only holds if you can inspect it.

The lesson is not “test more.” Green is a measurement of the tests you wrote, not the system you built, and the two only match if you go looking for the ways they do not. I built the tests. I built the code. But I do not get to be the only one checking whether they agree.

Explore the ds-canon repo and full factory log, or read the operating method as a case study: inside my multi-agent design operation.

Jay Trainer

Jay Trainer

Design Leader

Design executive focused on AI-native healthcare workflows and UX research, with product design leadership and design systems, plus human-in-the-loop product development.