VNX 1.7: What I Promised in June, and What the Framework Enforces Now

On June 2 I open-sourced VNX, my framework for Claude Code multi-agent orchestration, with a list of what worked. Version 1.7.1 shipped on October 3. I went back to that list line by line.

Three things are now enforced: one door for every dispatch, a review gate that can't be skipped silently, and a merge that needs green CI on the exact commit. Only the review gate was on the June list. One thing on that list was never true. And the self-learning loop I listed as "burning in" turned out to be measuring nothing at all.

Two things set VNX apart, and both are still the point. The first is governance: every dispatch leaves a receipt, and nothing merges without evidence. The second is how it pays for models. It drives the official CLIs on their own subscription logins instead of buying API credits. Of 839 routing decisions between September 15 and October 5, 763 went to a CLI on a subscription: Claude Code (734), Codex (15) and Kimi (14). The other 76 ran on GLM and DeepSeek on API credits, all before September 25. Since then those two serve as review fallback and for nightly analysis.

What I promised in June, and the claim that was wrong

The June launch post listed what worked "today". Reading it back against the repository history, four lines don't hold.

"Every dispatch produces an append-only, hash-chained NDJSON receipt." The chain existed, behind a flag that was off. It has never been switched on. Today none of my six local ledgers carries a single chained line. I come back to this below, because it is the one claim I have to correct.

"Today VNX is open source: 1.0." That day the newest tag was a release candidate. Version 1.0.0 and the PyPI package came on July 2.

"Per-dispatch git worktree isolation." The README of that same day said the setting defaulted to off. It has been on by default since August 11.

"Codex and Gemini review dispatches." Gemini was dropped as a required gate in July and removed as a reviewer in September. The default review stack is now Codex and Kimi, both on a subscription.

Two columns comparing VNX on June 2 and VNX 1.7.1: single dispatch door, review gate, CI gate, ledger hash chain, learning loop and worker lane, with the June state and the October state for each
June 2 against October 3. Orange marks what is enforced now.

One door for every dispatch

In June there was no single way in. A dispatch could start from a file, a script or a hand-typed command, and each path checked different things. The first door code landed on June 15.

Since June 24 the door is the default. Since 1.6.7, vnx dispatch <file.md> is refused outright. A dispatch that writes code but has no review gate attached is refused before anything spawns. So is a dispatch to a dead lane, to a model a provider constraint blocks, or with a project id that points at a local path.

There is a documented way back: VNX_DISPATCH_LEGACY=1 reopens the old path. I keep it, because a rollback switch you can't use is worse than one you have to explain. One known side door remains, in the plan gate's panel, and its fail switch is still off. That one is on my list.

A review verdict only counts on the commit it read

Until version 1.6.0 in August, the review gate could be skipped. The changelog entry that closed it is blunt: it closes "the gap where 95 PRs merged over four days with zero review-gate runs." Those were my own PRs. Nothing had stopped them.

Now a writing dispatch without a gate and a merge without a verdict are both refused. And a verdict only counts under strict conditions:

  • It is bound to the exact commit under review. A PASS on an older head doesn't carry over to a newer one.
  • It must cover the whole diff. A review that could only read part of the diff is booked as a partial review, and that blocks.
  • REVISE counts as a rejection. It used to be easy to read a REVISE as "mostly fine".
  • A reviewer outage counts as a missing review, not as a rejection. The obligation stays open until a reviewer delivers.

The seats can still be overruled, with VNX_OVERRIDE_<CHECK> and a written reason that lands in the audit trail. The gates are the mechanism behind the async quality gates I wrote about in March. The difference since June is that they can no longer be quietly bypassed.

No merge without green CI on that exact commit

Until early August, my CI ran 18 of 933 test files. The rest existed and were never checked on a pull request. Widening CI exposed 213 red tests nobody had seen.

Today CI runs the full suite minus a published exclusion list of 119 test files in 1.7.1, about 9 percent of the suite, each with a written reason. A merge needs a green CI run on the exact head commit of the pull request. Branch protection lives in a YAML file in the repository, with a drift check against what GitHub actually enforces.

On the VNX repository itself, GitHub enforces it: 15 required checks, including the review gate, with admins included and no force pushes. My merge door refuses a pull request that weakens that file unless it carries an explicit flag and a reason.

The limit: on repositories without that branch protection, the merge door is discipline, not enforcement. Nothing stops a raw gh pr merge there.

Hashing: what is fingerprinted, and what still isn't chained

This is the claim I have to correct.

In June I wrote that every receipt was hash-chained, tamper-evident by chain. The code for that existed, and since then it has grown: SHA-256 over canonical JSON, a verify command, epoch rotation, an external chain-origin anchor checked in CI. An architecture decision in July even approved switching it on by default. The step that would have done that never ran.

The flag VNX_CHAIN_RECEIPTS is off. None of my six project ledgers contains a chained line. My own audit tool reports them as "unchained", and counts that as green.

What does work is hashing in other places:

  • At the door. Every dispatch instruction gets a SHA-256 fingerprint before a permit is issued. Before delivery the file is hashed again, and delivery is refused if it changed. The last 200 routing decisions in my main project all carry that hash.
  • On review verdicts. A verdict must carry a contract hash and the exact commit hash it judged.
  • On plan approvals. That ledger is chained, and its 124 entries verify.

In June I also wrote that every receipt carries the instruction hash. On the receipts themselves that went missing: 31 of 1,822 completed receipts since mid-September carry it. It held at the door, where it is enforced, and not on the receipt, where it was best effort.

So the accurate description of the ledger in 1.7.1 is this. It is append-only NDJSON, and a health check reconciles it every six hours against what was dispatched. It is not cryptographically tamper-evident. Turning the chain on takes a one-time epoch seal per ledger and then a flag. I haven't done it yet.

Learning: measurable now, not smarter

In June I listed "a self-learning loop that consolidates past review findings" as opt-in and burning in.

On September 9 I parked the intelligence layer, with a reason in the register: until the governance ledger is falsifiable.

Two weeks later I found out what it had been doing. The instruments were dead. One table held 0 outcome rows after 6,919 pattern injections, because a timezone error escaped the instrument's own error handling. The consolidation job had skipped all 55 of its scheduled runs, so its table held 0 rows. The confidence table held one row, from June. The loop had been running for months and recording nothing I could use.

In 1.7.0 the learning loop came back in shadow mode. It runs every night, computes which patterns help and which fail, writes a report, and changes nothing. Persisting anything takes an explicit flag. Proposals, rules and edits only take effect after I approve them. A placebo arm exists so a future measurement has something to compare against.

One part did keep running: known failure patterns are added to a worker's prompt as advice. Since September 9 the injector ran 1,122 times in my main project. 895 of those runs added at least one pattern. It changes the context a worker sees. It does not change rules, routing or gates.

That is less than what I promised in June. It is also the first version where I can tell whether learning works at all.

What I didn't promise, but needed most

The changes I'd point to first weren't on the June list.

Every dispatch gets one computed outcome. Before 1.7.0 a dispatch could end without one and close silently. Now an open outcome surfaces to the orchestrator, and only an explicit reject can file it away.

Workers run headless only. The tmux worker lane was removed in September. With it went a whole class of failure: the instruction that never reached the terminal.

No more fabricated success. The Claude tmux lane used to write a report itself when the worker delivered none, with all four required headings, filled from the git log. That report always passed validation. I measured what it did to the numbers: 100.0% contract-complete with the synthesis, 90.9% without. The metric was measuring the repair, not the delivery. The synthesis now writes a format that fails validation on purpose.

The orchestrator has limits too. At 500,000 tokens of context a hook stops it from starting new work until it rotates. And it can no longer spawn its own subagents. One project had logged 280 subagent calls across 13 orchestrator sessions.

Receipts say who actually ran. The lane's real identity and the resolved model are on the receipt, so a model alias that moves under you shows up in the data.

On June 2 the repository had 20 architecture decisions. Today it has 38. The central ledger across my projects holds 41,139 receipts, and the suite has 24,559 test functions over 1,332 files.

What is still open

  • A safety net when the orchestrator dies mid-dispatch. If it stops before the governance step, the dispatch has no report. That is the highest-priority open item for 1.7.2.
  • Watching a worker while it runs. I see liveness through a lock, not the worker's own loop.
  • Default off: the receipt chain, cost-aware routing tiers, the roadmap autopilot, worker permission enforcement and learning persistence.
  • Parallel feature tracks. Still designed, not built. There is no wave scheduler and no merge lease.
  • One maintainer. That hasn't changed since June.

If you run Claude Code agents for your own work, the auth rules for Claude's API credits decide which wallet pays for them. And the Hermes Agent post explains why VNX drives the official CLI instead of reusing its login.

VNX is open source on GitHub. If you find a June claim I missed, or a gate that is weaker than I describe here, I'd like to hear it.

If you want this kind of evidence trail around the agents in your own company, that is the core of my AI governance work.

Vincent van Deth

AI Strategy & Architecture

I build production systems with AI — and I've spent the last six months figuring out what it actually takes to run them safely at scale.

My focus is AI Strategy & Architecture: designing multi-agent workflows, building governance infrastructure, and helping organisations move from AI experiments to auditable, production-grade systems. I'm the creator of VNX, an open-source governance layer for multi-agent AI that enforces human approval gates, append-only audit trails, and evidence-based task closure.

Based in the Netherlands. I write about what I build — including the failures.

Comments

Your email address will not be published. Comments are reviewed before publication.

Loading comments...