← All writing
AI-native

How my agent system works in one dated comparison

A July 2026 comparison of four specific agent configurations, with a closer look at persistent context and structured challenge, plus evidence gates.

Abstract editorial comparison of four agent configurations side by side, the newest highlighted with a blue ring, over a greeked dated timeline along the bottom.
Scope of this comparison

This article describes a system I built in January 2026 and the public documentation I reviewed in July 2026. Agent products change quickly. The comparison is limited to those configurations at that time, and the behavior I describe is what I observed in specific runs, not a universal claim about every possible setup.

The short version

I wanted an agent system that could preserve context, examine a decision from more than one assigned perspective, and stop at a human decision gate when the evidence was weak.

The configuration used five roles across research, strategy, design, delivery, and engineering. Each role had a written brief and persistent state. For work that benefited from comparison, I ran separate instances with isolated context, then asked another instance to synthesize the reports and surface conflicts. For higher-stakes decisions, I added an advocate and a challenger before making the call myself.

The result I was optimizing for was not universal superiority. It was a narrower operating behavior: keep the reasoning inspectable, make disagreement visible, and prevent a confident draft from becoming an accepted decision without support.

What I compared

I reviewed the public documentation for Hermes Agent, OpenClaw, and Agent Zero. Those tools emphasized different jobs: persistent assistance, channel-connected agent access, and autonomous task execution. My configuration emphasized evidence review and explicit human adjudication inside product work.

My January configurationHermes AgentOpenClawAgent Zero
Primary shapeFive role-based perspectives coordinated through written artifactsPersistent assistant with tools, memory, and delegated subagentsSelf-hosted gateway connecting agents to channels and toolsGeneral-purpose autonomous agent with subordinate agents
PersistenceRole briefs, state, and decision records stored outside the modelPersistent files and configurable memoryWorkspace, session, and memory featuresProject context and memory features
Coordination reviewedSeparate role reports followed by synthesis and human adjudicationParent-to-child delegationSession and agent routingHierarchical task delegation
Decision controlEvidence check, challenge pass, then a human decisionOperator-configured approvals and tool boundariesOperator-configured policies, approvals, and routingOperator-defined goals, prompts, and controls

This is a comparison of emphasis and configuration, not a benchmark. Each of the other systems can be extended, and a different operator could configure the same products differently.

How the system worked

One role for focused work

For a focused task, one instance loaded a role brief and the context relevant to that assignment. The role constrained what the instance should notice and what it was responsible for returning. This was still one model responding to one context, so I treated it as a focused perspective, not an independent expert.

Separate runs for comparison

For scope checks, critiques, risk reviews, and planning, I ran role assignments separately so one report could not simply echo another. A synthesis pass compared the results and called out contradictions. In the runs I evaluated, that separation produced useful differences in emphasis and concern. It did not guarantee cognitive diversity or genuine independence.

Challenge before commitment

For a consequential decision, one run argued for the proposal and another tried to break it. I reviewed both positions, checked their evidence, and made the decision. The value came from the structure of the review, not from assuming that separate model instances would disagree on their own.

The evidence gate

A simple generic mechanism:

  1. A proposal must point to evidence that meets the threshold set for that decision.
  2. If the evidence is weak, missing, or contradictory, the proposal cannot move directly to accepted.
  3. A challenge pass tests the claim, its assumptions, and the cited support.
  4. A person reviews the evidence and the challenge, then accepts, changes, or rejects the proposal.
  5. The decision and its rationale remain attached to the work so they can be inspected later.

A cleared public example is ds-canon. Its factory record shows builders working from a frozen contract, challengers filing reproducible findings, a fixer answering each finding in writing, and a final verification pass rerunning the work. A green test suite was evidence, but it was not the only gate. Findings needed a concrete failure scenario and a proposed fix before they counted, and fixes needed to survive another check before the artifact shipped.

That is the pattern I keep: generation creates a candidate, evidence supports or weakens it, challenge looks for what the first pass missed, and a human owns the consequential decision.

What I observed

Persistent context reduced repeated setup

Keeping role briefs and decision records outside the model made it easier to resume work without rebuilding the entire context from memory. The tradeoff was upkeep. Stale context could mislead a run just as easily as missing context.

Isolation made conflicts easier to see

Separate runs sometimes emphasized different risks or interpreted the same brief differently. The synthesis pass made those differences visible. Sometimes the reports still converged, and sometimes the differences came from the prompts rather than the models. I treated disagreement as a useful signal to inspect, not proof of independent reasoning.

Evidence gates slowed the right moment

The gate added friction at the point where a proposal became a decision. That was intentional. Generating options could stay fast, while acceptance required support, a challenge, and a person willing to own the call.

Limitations

  • The operator remains the bottleneck. The system only challenges the decisions I route through it, and I still judge the evidence.
  • Role prompts can create theater. Different labels do not guarantee different reasoning. Useful separation has to be evaluated in the outputs.
  • Persistent context can decay. Written state needs review or it becomes an authoritative-looking source of stale assumptions.
  • This was a custom configuration. It required conventions and discipline that packaged agent tools provide differently.
  • The comparison expires. Hermes Agent, OpenClaw, Agent Zero, and my own practice continue to change.

What held up

The durable lesson was not that one architecture wins. It was that an agent system becomes more useful for product decisions when it preserves the basis for a claim, separates generation from challenge, and makes a human decision explicit. Those mechanisms can be implemented in many tools. The important question is whether the operating record shows what was proposed, what evidence supported it, what challenged it, and who decided.

Where this is headed

This January 2026 system was an early version of my practice. I have continued to change the implementation, but the same questions remain: what should persist, what should be challenged, what can automation enforce, and where must a person decide?

Jay Trainer

Jay Trainer

Design Leader

Design executive focused on AI-native healthcare workflows and UX research, with product design leadership and design systems, plus human-in-the-loop product development.