Files
hq/01-RESEARCH/014-the-bus-on-nats/00-overview.md
T

7.1 KiB

status, became, initiated, touches
status became initiated touches
graduated 02-DECISIONS/0106-the-bus-is-nats.md 2026-09-23
02-DECISIONS/0002-nodes-communicate-over-a-broker.md
02-DECISIONS/0033-the-substrate-is-a-store-and-a-broker.md
02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md
02-DECISIONS/0041-events-are-a-relationship.md
02-DECISIONS/0042-the-shape-of-an-event-on-the-wire.md
02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md
02-DECISIONS/0078-the-store-and-broker-are-modules.md
02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md

014 — The bus on NATS: replace the broker, and when

The question, asked mid-migration: the operator wants the mesh's bus — today an AMQP broker — replaced by NATS, with everything the broker does today. Do we finish the migration on AMQP and move to NATS after, or go to NATS directly?

The short answer: decide NATS now, build it in the lab in parallel, and cut the mesh's bus over in one rehearsed rollout — after the migration's core is done and never underneath it. "First everything on AMQP" is already the state and costs nothing more: the predecessor's broker was merged into the mesh's tonight, and every module converted from here on targets the sdk's broker contract, which names no protocol. The only code that speaks AMQP is the mesh's own, in three places, and it is swapped once.

What was measured

Where AMQP is spoken (non-test files): the controller's internal/link (5 files, ~2.4k lines with tests) and the builder's main; the host's internal/link (~1.5k); the tool runtime's broker-amqp.ts (~600 with its main). The sdk speaks none — its Broker is request, handle, publish, subscribe, close (ADR 0039 put the client in the runtime for exactly this). Every module's tools and events go through the sdk; one module (amqp-ping, a probe) talks AMQP on purpose.

What the bus carries: a control queue (control, .upgrades, .catchup), one queue per node (node.<name>), builds, the enrolment/report/alive/built flows, a topic exchange of events (mesh.events.<module>.<event>, dead-lettered to mesh.events.dead), an RPC exchange (mesh.rpc) and per-tool service queues (serve.<module>.<tool>), plus an MQTT exchange.

What it relies on: TLS for the bus a node enrols over; prefetch with reject/nack so the control queue holds messages unacknowledged while the store restarts and retries (ADR 0083); publisher confirms and mandatory routing; a dead-letter exchange; per-module accounts scoped by emits/consumes (ADR 0043), minted through the broker's management API; the broker as a module holding the mesh-broker seat (ADR 0078).

The predecessor's world, now on the same broker: ~50 AMQP connections per machine from four machines, 595 queues, four vhosts, six users. All of it AMQP, all of it retiring module by module. NATS speaks no AMQP: that world cannot move; it does not need to.

How NATS answers each of those

The mesh relies on NATS Note
topic routing keys subjects with wildcards same shape (mesh.events.>)
RPC and tool invocation over an exchange + reply queue request/reply, native simpler than today
competing consumers queue groups same
durable control queue, hold-unacked-and-retry, catch-up JetStream: streams, durable consumers, ack/nak with delay, replay the 0083 guarantee moves to JetStream; core NATS alone is at-most-once and would not do
dead-letter max-deliver + advisories, or a stream fed from them different mechanism, same effect
TLS bus TLS same
accounts scoped by emits/consumes, vhosts accounts (isolation) with users and per-subject publish/subscribe permissions stronger than today; the controller writes an auth config the host declares and the server reloads, instead of calling a management API
management API nats CLI and an HTTP monitoring endpoint no vhost concept — accounts instead
MQTT built in same
a module seat mesh-broker unchanged — the seat is the server, the module changes ADR 0079
multi-node clusters and leaf nodes not needed now; a leaf per node is a later question

Nothing the mesh needs is missing. The differences are in the shape of durability (JetStream must be declared, streams and consumers are objects) and of accounts (configuration, not API calls).

The cost, honestly

The bus is the mesh's nervous system. Moving it means: a new nats module in the catalogue taking the mesh-broker seat; the controller's and the host's link packages rewritten; the tool runtime's client swapped behind the unchanged contract; enrolment, reports, builds and the guard's ports re-derived; the lab beds that prove the bus (store window, enrolment, upgrades) re-run on the new one; and eight decisions amended or superseded. Weeks, not days — and none of it can be done halfway on a live mesh: the controller, every host and every tool runtime move together.

The sequencing question, answered

  1. NATS first, directly. Stalls the migration for the length of the build; the mesh's bus changes under a half-migrated node; the predecessor's clients still need AMQP, so a second broker runs anyway, plus a bridge for whatever crosses. Rejected.
  2. Finish on AMQP, NATS after. Wastes nothing — no module written from here on speaks AMQP — but leaves the decision unmade while modules are written, and the runtimes idle. Adequate.
  3. Decide NATS now; build it in the lab in parallel; cut the mesh's bus over in one rehearsed rollout after the core is migrated. The predecessor's clients never notice: their broker is the one the mesh adopted, kept as a compatibility module with an end date — the day the last AMQP client is gone. Recommended.

What this leaves open

  • The decision itself, as a record: the bus is NATS; the AMQP broker becomes the predecessor's compatibility broker and retires with the last AMQP client. Written when the operator says so.
  • JetStream's shape for the control plane: one stream per concern (control, nodes, builds, events) or one with subjects; retention; what the store window guarantee looks like as ack-wait and nak-delay. Measured in the lab, not designed on paper.
  • Accounts as configuration: the controller writes users and permissions into a file the host declares, reloaded on change — which is the ADR 0102 discipline applied from the start — or the JWT/operator model. The first is simpler and matches how the mesh already writes everything.
  • The MCP bridge the operator asked for the same day is written against the sdk's contract, so it moves with the bus and is not written twice.
  • Whether a leaf node per machine replaces the hub-and-spoke bus later — out of scope here.