Research 014: the bus on NATS — decide now, build in the lab, cut over once after the core

Measured: AMQP is spoken in three places of the mesh's own code and in none of the
sdk or the modules; the predecessor's world is AMQP and retiring. NATS answers every
guarantee the bus relies on, durability via JetStream. Recommended: not underneath
the migration, not after it either — in parallel, one rehearsed rollout.
This commit is contained in:
2026-09-23 23:14:39 +02:00
parent 8638ba3a4f
commit b59907ee08
@@ -0,0 +1,107 @@
---
status: active
initiated: 2026-09-23
touches:
- 02-DECISIONS/0002-nodes-communicate-over-a-broker.md
- 02-DECISIONS/0033-the-substrate-is-a-store-and-a-broker.md
- 02-DECISIONS/0039-the-broker-client-lives-in-the-runtime-not-the-sdk.md
- 02-DECISIONS/0041-events-are-a-relationship.md
- 02-DECISIONS/0042-the-shape-of-an-event-on-the-wire.md
- 02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md
- 02-DECISIONS/0078-the-store-and-broker-are-modules.md
- 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
---
# 014 — The bus on NATS: replace the broker, and when
**The question, asked mid-migration:** the operator wants the mesh's bus — today an AMQP broker —
replaced by NATS, with everything the broker does today. Do we finish the migration on AMQP and
move to NATS after, or go to NATS directly?
**The short answer: decide NATS now, build it in the lab in parallel, and cut the mesh's bus over
in one rehearsed rollout — after the migration's core is done and never underneath it.** "First
everything on AMQP" is already the state and costs nothing more: the predecessor's broker was
merged into the mesh's tonight, and every module converted from here on targets the sdk's broker
contract, which names no protocol. The only code that speaks AMQP is the mesh's own, in three
places, and it is swapped once.
## What was measured
**Where AMQP is spoken** (non-test files): the controller's `internal/link` (5 files, ~2.4k lines
with tests) and the builder's main; the host's `internal/link` (~1.5k); the tool runtime's
`broker-amqp.ts` (~600 with its main). **The sdk speaks none** — its `Broker` is `request`,
`handle`, `publish`, `subscribe`, `close` ([ADR 0039](../../02-DECISIONS/0039-the-broker-client-lives-in-the-runtime-not-the-sdk.md)
put the client in the runtime for exactly this). **Every module's tools and events go through the
sdk**; one module (`amqp-ping`, a probe) talks AMQP on purpose.
**What the bus carries:** a control queue (`control`, `.upgrades`, `.catchup`), one queue per node
(`node.<name>`), `builds`, the enrolment/report/alive/built flows, a topic exchange of events
(`mesh.events.<module>.<event>`, dead-lettered to `mesh.events.dead`), an RPC exchange (`mesh.rpc`)
and per-tool service queues (`serve.<module>.<tool>`), plus an MQTT exchange.
**What it relies on:** TLS for the bus a node enrols over; prefetch with reject/nack so the control
queue **holds messages unacknowledged while the store restarts** and retries
([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)); publisher confirms and
mandatory routing; a dead-letter exchange; **per-module accounts scoped by `emits`/`consumes`**
([ADR 0043](../../02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md)),
minted through the broker's management API; the broker as a module holding the `mesh-broker` seat
([ADR 0078](../../02-DECISIONS/0078-the-store-and-broker-are-modules.md)).
**The predecessor's world**, now on the same broker: ~50 AMQP connections per machine from four
machines, 595 queues, four vhosts, six users. All of it AMQP, all of it retiring module by module.
NATS speaks no AMQP: that world cannot move; it does not need to.
## How NATS answers each of those
| The mesh relies on | NATS | Note |
|---|---|---|
| topic routing keys | subjects with wildcards | same shape (`mesh.events.>`) |
| RPC and tool invocation over an exchange + reply queue | request/reply, native | simpler than today |
| competing consumers | queue groups | same |
| durable control queue, hold-unacked-and-retry, catch-up | JetStream: streams, durable consumers, ack/nak with delay, replay | the 0083 guarantee moves to JetStream; core NATS alone is at-most-once and would not do |
| dead-letter | max-deliver + advisories, or a stream fed from them | different mechanism, same effect |
| TLS bus | TLS | same |
| accounts scoped by emits/consumes, vhosts | accounts (isolation) with users and per-subject publish/subscribe permissions | stronger than today; the controller writes an auth config the host declares and the server reloads, instead of calling a management API |
| management API | `nats` CLI and an HTTP monitoring endpoint | no vhost concept — accounts instead |
| MQTT | built in | same |
| a module seat `mesh-broker` | unchanged — the seat is the server, the module changes | ADR 0079 |
| multi-node | clusters and leaf nodes | not needed now; a leaf per node is a later question |
Nothing the mesh needs is missing. The differences are in the shape of durability (JetStream must
be declared, streams and consumers are objects) and of accounts (configuration, not API calls).
## The cost, honestly
The bus is the mesh's nervous system. Moving it means: a new `nats` module in the catalogue taking
the `mesh-broker` seat; the controller's and the host's link packages rewritten; the tool runtime's
client swapped behind the unchanged contract; enrolment, reports, builds and the guard's ports
re-derived; the lab beds that prove the bus (store window, enrolment, upgrades) re-run on the new
one; and eight decisions amended or superseded. Weeks, not days — and none of it can be done
halfway on a live mesh: the controller, every host and every tool runtime move together.
## The sequencing question, answered
1. **NATS first, directly.** Stalls the migration for the length of the build; the mesh's bus
changes under a half-migrated node; the predecessor's clients still need AMQP, so a second
broker runs anyway, plus a bridge for whatever crosses. Rejected.
2. **Finish on AMQP, NATS after.** Wastes nothing — no module written from here on speaks AMQP —
but leaves the decision unmade while modules are written, and the runtimes idle. Adequate.
3. **Decide NATS now; build it in the lab in parallel; cut the mesh's bus over in one rehearsed
rollout after the core is migrated.** The predecessor's clients never notice: their broker is
the one the mesh adopted, kept as a compatibility module with an end date — the day the last
AMQP client is gone. **Recommended.**
## What this leaves open
- The **decision itself**, as a record: the bus is NATS; the AMQP broker becomes the predecessor's
compatibility broker and retires with the last AMQP client. Written when the operator says so.
- **JetStream's shape for the control plane**: one stream per concern (control, nodes, builds,
events) or one with subjects; retention; what the store window guarantee looks like as ack-wait
and nak-delay. Measured in the lab, not designed on paper.
- **Accounts as configuration**: the controller writes users and permissions into a file the host
declares, reloaded on change — which is the [ADR 0102](../../04-ISSUES/102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md)
discipline applied from the start — or the JWT/operator model. The first is simpler and matches
how the mesh already writes everything.
- **The MCP bridge** the operator asked for the same day is written against the sdk's contract, so
it moves with the bus and is not written twice.
- Whether a **leaf node per machine** replaces the hub-and-spoke bus later — out of scope here.