7.0 KiB
status, initiated, touches
| status | initiated | touches | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| active | 2026-09-23 |
|
014 — The bus on NATS: replace the broker, and when
The question, asked mid-migration: the operator wants the mesh's bus — today an AMQP broker — replaced by NATS, with everything the broker does today. Do we finish the migration on AMQP and move to NATS after, or go to NATS directly?
The short answer: decide NATS now, build it in the lab in parallel, and cut the mesh's bus over in one rehearsed rollout — after the migration's core is done and never underneath it. "First everything on AMQP" is already the state and costs nothing more: the predecessor's broker was merged into the mesh's tonight, and every module converted from here on targets the sdk's broker contract, which names no protocol. The only code that speaks AMQP is the mesh's own, in three places, and it is swapped once.
What was measured
Where AMQP is spoken (non-test files): the controller's internal/link (5 files, ~2.4k lines
with tests) and the builder's main; the host's internal/link (~1.5k); the tool runtime's
broker-amqp.ts (~600 with its main). The sdk speaks none — its Broker is request,
handle, publish, subscribe, close (ADR 0039
put the client in the runtime for exactly this). Every module's tools and events go through the
sdk; one module (amqp-ping, a probe) talks AMQP on purpose.
What the bus carries: a control queue (control, .upgrades, .catchup), one queue per node
(node.<name>), builds, the enrolment/report/alive/built flows, a topic exchange of events
(mesh.events.<module>.<event>, dead-lettered to mesh.events.dead), an RPC exchange (mesh.rpc)
and per-tool service queues (serve.<module>.<tool>), plus an MQTT exchange.
What it relies on: TLS for the bus a node enrols over; prefetch with reject/nack so the control
queue holds messages unacknowledged while the store restarts and retries
(ADR 0083); publisher confirms and
mandatory routing; a dead-letter exchange; per-module accounts scoped by emits/consumes
(ADR 0043),
minted through the broker's management API; the broker as a module holding the mesh-broker seat
(ADR 0078).
The predecessor's world, now on the same broker: ~50 AMQP connections per machine from four machines, 595 queues, four vhosts, six users. All of it AMQP, all of it retiring module by module. NATS speaks no AMQP: that world cannot move; it does not need to.
How NATS answers each of those
| The mesh relies on | NATS | Note |
|---|---|---|
| topic routing keys | subjects with wildcards | same shape (mesh.events.>) |
| RPC and tool invocation over an exchange + reply queue | request/reply, native | simpler than today |
| competing consumers | queue groups | same |
| durable control queue, hold-unacked-and-retry, catch-up | JetStream: streams, durable consumers, ack/nak with delay, replay | the 0083 guarantee moves to JetStream; core NATS alone is at-most-once and would not do |
| dead-letter | max-deliver + advisories, or a stream fed from them | different mechanism, same effect |
| TLS bus | TLS | same |
| accounts scoped by emits/consumes, vhosts | accounts (isolation) with users and per-subject publish/subscribe permissions | stronger than today; the controller writes an auth config the host declares and the server reloads, instead of calling a management API |
| management API | nats CLI and an HTTP monitoring endpoint |
no vhost concept — accounts instead |
| MQTT | built in | same |
a module seat mesh-broker |
unchanged — the seat is the server, the module changes | ADR 0079 |
| multi-node | clusters and leaf nodes | not needed now; a leaf per node is a later question |
Nothing the mesh needs is missing. The differences are in the shape of durability (JetStream must be declared, streams and consumers are objects) and of accounts (configuration, not API calls).
The cost, honestly
The bus is the mesh's nervous system. Moving it means: a new nats module in the catalogue taking
the mesh-broker seat; the controller's and the host's link packages rewritten; the tool runtime's
client swapped behind the unchanged contract; enrolment, reports, builds and the guard's ports
re-derived; the lab beds that prove the bus (store window, enrolment, upgrades) re-run on the new
one; and eight decisions amended or superseded. Weeks, not days — and none of it can be done
halfway on a live mesh: the controller, every host and every tool runtime move together.
The sequencing question, answered
- NATS first, directly. Stalls the migration for the length of the build; the mesh's bus changes under a half-migrated node; the predecessor's clients still need AMQP, so a second broker runs anyway, plus a bridge for whatever crosses. Rejected.
- Finish on AMQP, NATS after. Wastes nothing — no module written from here on speaks AMQP — but leaves the decision unmade while modules are written, and the runtimes idle. Adequate.
- Decide NATS now; build it in the lab in parallel; cut the mesh's bus over in one rehearsed rollout after the core is migrated. The predecessor's clients never notice: their broker is the one the mesh adopted, kept as a compatibility module with an end date — the day the last AMQP client is gone. Recommended.
What this leaves open
- The decision itself, as a record: the bus is NATS; the AMQP broker becomes the predecessor's compatibility broker and retires with the last AMQP client. Written when the operator says so.
- JetStream's shape for the control plane: one stream per concern (control, nodes, builds, events) or one with subjects; retention; what the store window guarantee looks like as ack-wait and nak-delay. Measured in the lab, not designed on paper.
- Accounts as configuration: the controller writes users and permissions into a file the host declares, reloaded on change — which is the ADR 0102 discipline applied from the start — or the JWT/operator model. The first is simpler and matches how the mesh already writes everything.
- The MCP bridge the operator asked for the same day is written against the sdk's contract, so it moves with the bus and is not written twice.
- Whether a leaf node per machine replaces the hub-and-spoke bus later — out of scope here.