← All writing
AI-native · Design systems

ds-canon: your design system, queryable by agents

Agents are writing most of the UI code now. A design system only humans can read is one they will not consult, they will approximate, and the approximation will look plausible enough to ship.

ds-canon abstract product-window editorial cover

Every design system I have shipped eventually decays into tribal knowledge: the token sheet drifts from the code, and the one person who remembers why the legacy blue is still allowed in one component moves on to another project.

Now add agents. An agent gets asked to build a secondary button, and it does what agents do: it looks at a few components for pattern, picks something plausible for the color it needs, color.brand.secondary or space.sm, and ships code that compiles and looks close enough that nobody catches it in review. The token does not exist. It never did. The agent was not lying. It was guessing with total confidence, because guessing with total confidence is what a language model does when there is nothing authoritative in front of it to check against.

A design system only humans can read gets ignored

That is the failure mode worth naming directly. A design system that lives in a Figma file, a Notion page, and one senior designer's memory is a design system only humans can read. Humans could always route around the gaps in that: ask in Slack, guess, get corrected in review. Agents cannot ask in Slack. They cannot tell tribal knowledge from an actual constraint. If the system of record is not something a program can query, agents will not consult it. They will approximate it, and the approximation will ship.

So I built the thing I wanted my own agents to be able to ask.

What I built

ds-canon is a small, read-only MCP server that puts a design system's tokens, components, conventions, and deprecations behind eight queries an agent, or a person, can call directly: what tokens exist, what a specific token is and who still depends on it, what components exist, what a specific component looks like end to end, what breaks if you change something, what is deprecated, and what the house conventions are and why. An eighth tool, check_token_drift, matches literal values in a code snippet against the token source, so a hardcoded hex or pixel value gets flagged even when nobody thought to ask.

The tool that matters most is find_usages, because it answers the question a design system actually needs to answer before anyone touches a shared value: what breaks. Here is the real transcript, verbatim from the server's bundled fixture system, when I asked what changing a spacing token would affect:

{
  "entity": "space.inset.md",
  "usages": [
    { "dependent": "Button", "dependentKind": "component", "relation": "consumes token" },
    { "dependent": "Card", "dependentKind": "component", "relation": "consumes token" },
    { "dependent": "Field", "dependentKind": "component", "relation": "consumes token" },
    { "dependent": "Modal", "dependentKind": "component", "relation": "consumes token" }
  ]
}
find_usages on space.inset.md, real output from the bundled Nimbus DS fixtures. Four components, named exactly. No spelunking through a component library to find every place 12px got typed in by hand.

That is blast radius, in one call. It is the difference between a design system as documentation and a design system as an answer an agent can act on before it breaks something.

Why I directed an agent factory

I could have hired this out or waited for a vendor to get around to it. Instead, I directed an agent factory I built and operated solo. The job of a design leader now includes proving the infrastructure works, not just specifying it in a doc someone else implements months later. Enablement is running code, not a slide, and I want to be someone who can still ship the code and not only the opinion about it.

I did not write ds-canon alone in the ordinary sense either. I ran it through a small agent factory in a single afternoon: three builders working in parallel from a frozen, typed contract, one on the fixtures and loader, one on the tools and server, one on the test suite, none of them seeing the others' code until integration. Then two adversarial challengers whose only job was to attack what the builders shipped: one reading the source for correctness and honesty, one attacking the running server itself with hostile input and concurrency. Both filed real findings against a codebase with a fully green test suite, including a semantics bug in a usage count and one genuine blocker, a fuzzy-suggestion lookup that could freeze the whole server on an oversized bad query.

A fixer answered every finding in writing, fixed or rejected each one with a stated reason, and grew the test suite from 33 tests to 78. I did not take its word for it either: I re-ran the build, the suite, and the freeze attack myself, with a 2MB hostile payload, before calling it done. Git init to a public GitHub repo took under 48 minutes. The dispute log, the prompts, and the timestamps are public in the repo's factory/ directory, because the point of building this way is not the speed. It is that the speed came with more scrutiny than a solo build gets, not less. The method is its own case study: inside my multi-agent design operation.

Point it at your own system

Nimbus DS, the fixture system ds-canon ships with, exists so the demo runs with nothing to clone. Your own design system is the point. Set one environment variable, DS_CANON_FIXTURES, to a local path holding your tokens.json, components.json, and conventions.md, and the same eight queries run against your system instead: read-only, over a local clone, one config entry per person who wants it. If you already export tokens from Style Dictionary or Tokens Studio in W3C DTCG format, that file works here with no transformation.

Install

It runs over stdio. There is no separate service to deploy, and nothing to clone to try it. From npm:

npx -y ds-canon

For Claude Code, one line:

claude mcp add ds-canon -- npx -y ds-canon

The repo is at github.com/jtrainer357/ds-canon, and the package is on npm at npmjs.com/package/ds-canon. Point an agent at it and ask what's deprecated, or what breaks if you touch a shared value. AI proposes. The system of record answers.

Jay Trainer

Jay Trainer

Design Leader

Design executive focused on AI-native healthcare workflows and UX research, with product design leadership and design systems, plus human-in-the-loop product development.