Files
hq/03-DESIGN/01-to-be/25-the-bus-on-nats.md
T
jschoubben 3f9b316015 Design 25: a scoped inbox needs allow_responses, or nothing can answer
Found composing the first real configuration. Scoping every inbox to its
owner is right and leaves a responder unable to reply, because the answer
goes to the caller's inbox. The fix is not a wider grant but the server's
own allow_responses: one reply to the subject of a message the user actually
received. Without it every tool call times out while the permission list
looks correct.
2026-09-26 20:58:34 +02:00

447 lines
31 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
layer: to-be
status: in-progress
code:
- mesh-controller internal/link (to be replaced)
- mesh-host internal/link (to be replaced)
- mesh-tools src/broker-amqp.ts (to be replaced)
- mesh-catalog modules/nats (to be written)
- mesh-sdk src (the protocol's NATS binding, step 3)
updated: 2026-09-26
decisions:
- 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0117-the-bus-is-the-only-broker.md
- 02-DECISIONS/0116-the-bus-is-built-in-five-steps.md
- 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md
- 02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md
- 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
- 02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md
- 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
- 02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md
---
# 25. The bus on NATS
**Status: proposed — a design to be reviewed before any code.** This is the architecture
[ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md) asks for. It says what rides the bus,
under which subject, with which guarantee, under whose account; how a node joins; how a person
reaches a tool; how the mesh moves from the bus it has to this one; and how each claim is checked.
Prose and diagrams only; no configuration is pasted.
## 1. What the bus is for
The bus carries five kinds of traffic today, and this design keeps the five, renaming nothing a
module can see:
| Traffic | Today | Guarantee it needs |
|---|---|---|
| **control** — a node's report, its heartbeat, a build's outcome, an enrolment | queues `control`, `.upgrades`, `.catchup` | nothing lost while the store restarts; retried; in order per node |
| **declarations** — the controller tells a node what to be | queue `node.<name>` | the node gets the newest; a stale one is never applied |
| **builds** — the controller asks the build machine to build | queue `builds` | at least once, one builder at a time |
| **events** — a module says something happened | topic exchange `mesh.events`, keys `<module>.<event>` | delivered to every consumer that declared it; dead-lettered when it cannot be |
| **tools** — one module or person asks another's tool a question | exchange `mesh.rpc`, per-tool service queues `serve.<module>.<tool>` | one answer, from one server, or a timeout |
The sdk's contract — `request`, `handle`, `publish`, `subscribe`, `close` — is the whole surface a
module sees, and it does not change ([ADR 0039](../../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md)).
## 2. Subjects
NATS addresses everything by subject. The mesh's subject space is one tree, and every account's
permissions are expressed as which branches of it that account may publish to and subscribe from.
```
mesh.control.<node>.report a node's report (JetStream: CONTROL)
mesh.control.<node>.alive heartbeat (core, no persistence)
mesh.control.enrol an enrolment request (JetStream: CONTROL)
mesh.control.built a build's outcome (JetStream: CONTROL)
mesh.node.<node>.declare a declaration for a node (JetStream: NODES, last-per-subject)
mesh.build.request work for the build machine (JetStream: BUILDS, work queue)
mesh.events.<module>.<event> an event (JetStream: EVENTS)
mesh.tools.<module>.<tool> a tool invocation (core request/reply)
mesh.ask.<node>.<command> the controller's command api (core request/reply)
```
Two things this buys over the exchanges: **request/reply is native** — a tool call is one
`request` on `mesh.tools.<module>.<tool>` answered by whichever runtime serves it (a queue group per
tool, so several nodes may serve one tool); and **a declaration is last-per-subject** — the NODES
stream keeps only the newest message on `mesh.node.<node>.declare`, so a node that was away gets
exactly the current declaration and nothing older. That is the wire-level answer to
[issue 107](../../04-ISSUES/107-a-declaration-carries-no-order/00-report.md): the stream's sequence
*is* the order, and a node that sees sequence n refuses n−1 by construction.
**A reply-to travelling through a JetStream stream is carried in the payload, never in the
transport `Reply` field.** Revision, first review: core NATS request/reply sets the requester's
ephemeral inbox as the message's `Reply` field, and a plain responder answers it directly — but a
message a JetStream consumer delivers has already had that field claimed for the consumer's own
ack address (`$JS.ACK.<stream>.<consumer>...`), so by the time the controller (§3's CONTROL
consumer) sees the message, `Reply` names where *it* must ack, not where the original caller is
waiting. `mesh.control.enrol` is the case that matters: a synchronous-feeling caller waiting on an
ephemeral inbox, over a subject the store-window guarantee may legitimately delay by several
`nak` cycles — exactly the combination that would otherwise deliver the answer to a caller who
has long since timed out and unsubscribed. So every CONTROL message that expects an answer states
its reply subject as an ordinary field of its own payload; the controller reads it from there and
publishes the answer to it explicitly, never via `Respond()`. Nothing else in this design routes
a reply through a stream — tools and heartbeats stay on core NATS, where `Reply` means what it has
always meant.
## 3. Streams, and the guarantees they carry
Core NATS is at-most-once. Everything the mesh must not lose lives in a JetStream stream:
| Stream | Subjects | Retention | Why |
|---|---|---|---|
| CONTROL | `mesh.control.>` except `alive` | work queue, one consumer (the controller), explicit ack | the store-window guarantee ([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)): the controller `nak`s with a delay while its store is away and the message is redelivered; nothing is dropped |
| NODES | `mesh.node.>` | last per subject | one declaration per node, always the newest |
| BUILDS | `mesh.build.>` | work queue, explicit ack | at least once; a builder that dies mid-build has its message redelivered |
| EVENTS | `mesh.events.>` | limits (age, size), durable consumer per subscribing module | a subscriber that was down catches up; after `max-deliver` attempts the advisory feeds `mesh.events.dead` (its own small stream) |
Tool calls and heartbeats stay on core NATS: a lost heartbeat is the next heartbeat; a lost tool
call is a timeout the caller already handles.
Streams and consumers are objects the controller creates at genesis and asserts on start; a module
declares nothing about them. The controller is the only writer of stream definitions.
## 4. Accounts
[ADR 0043](../../02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) says a
module's account may publish only what it `emits` and consume only what it `consumes`. NATS
expresses this exactly, per subject, and better than a vhost could:
- **One NATS account for the mesh.** Accounts in NATS isolate subject spaces entirely; the mesh
is one space, so it is one account. The predecessor's compatibility broker is not on this bus at
all.
**One account is a choice with a cost, stated plainly on revision:** none of NATS's own
isolation is free here, because there is only the one subject space, for everyone. Two
consequences that a single-account mesh must therefore grant on purpose, not by omission:
- **A durable consumer needs permission to ack, or it never really consumes.** Acking a
JetStream delivery is a publish to that consumer's own ack-reply address
(`$JS.ACK.<stream>.<consumer>.>`), a different subject from anything the consumer subscribes.
A module's user is therefore granted publish on `$JS.ACK.EVENTS.<module>.>` as well as its
emits — scoped to the one consumer name the controller derives for that module, so a module
can ack only its own deliveries. Without this, first review found, every message it receives
would be redelivered forever: refused by the permission list it already has.
- **A reply inbox needs a subject nothing else can guess or enumerate.** With one account,
inbox privacy is the permission list or it is nothing — there is no second account backing
it up. So no user is ever granted a bare `_INBOX.>`. Each user's inbox subject is derived
from its own identity (`_INBOX.<module>.<node>.>`, or `_INBOX.person.<name>.>`), and its
permissions name only that one prefix, for the reply to any request it makes and nothing
wider. First review found the account note without this and read it as "any user may
subscribe any inbox" — which was accurate against the text as it stood.
**And a scoped inbox needs `allow_responses`, or nothing can answer.** Revision, found while
composing the first real configuration: the rule above scopes each user's inbox to itself,
which is right — and leaves a responder unable to reply, because the answer goes to the
*caller's* inbox, which the responder has no permission for. The two ways out are granting
every responder `_INBOX.>`, which is exactly the blanket grant this bullet refuses, or NATS's
own `allow_responses`: the server permits one reply to the reply-subject of a message the user
actually received, within a TTL, and nothing else. So authority to answer is bounded by having
been asked, and only principals that serve something are granted it — a pure consumer gets
nothing. Without this the scoping is not merely incomplete: every tool call in the mesh times
out, and the permission list looks correct while it happens.
- **One user per module per node**, as today, with publish permissions
`mesh.events.<module>.<event>` for each emit, `mesh.tools.<module>.>` to serve its tools, its
own ack-reply subject for each durable consumer it holds, and its own inbox prefix; subscribe
permissions for each consumed event's subject, its tool subjects, and that same inbox prefix.
Nothing else. A module that tries to publish outside its emits is refused by the server, not by
convention.
- **The controller's user** owns `mesh.control.>`, `mesh.node.>`, `mesh.build.>` and the streams.
**A host's user** may publish its own `mesh.control.<node>.>` and subscribe its own
`mesh.node.<node>.declare` — and nothing of any other node's.
- **A person's user** (§7) is a module-shaped user with permissions on the tool subjects it may
invoke, issued and revoked by the controller like any account.
**Accounts are configuration, not API calls.** The controller composes the server's user list and
permissions into a file the host declares. **How that file reaches the running server is §5's,
not this one's** — revision, first review: an earlier draft said "reloads" and cited a precedent
that does not apply to a container (see §5). No management API, no credential travelling through a
management call, and the [issue 102](../../04-ISSUES/102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md)
discipline from the first day: an address or a permission is read where it is used, never stored
with a port. Passwords are minted and sealed exactly as today; the file holds bcrypt hashes.
Alternative considered and not taken: the operator/JWT model (`nsc`), where accounts are signed
tokens resolved by the server. It is the right model for a multi-tenant NATS; the mesh is one
tenant, already has a sealing key and a controller that writes files, and would gain a second
signing hierarchy for nothing.
## 5. The broker as a module
`nats` is a catalogue module claiming the seat `mesh-broker`
([ADR 0079](../../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md): the seat
is the server, and the server changes). It declares one container (a single binary; **JetStream on a host
directory bind, not a named volume** — revision, second review:
[issue 115](../../04-ISSUES/115-a-named-docker-volume-is-invisible-and-one-flag-from-gone/00-report.md)
is resolved, and converted the store, the broker and two others away from named volumes for the
reason it names; the bus's own data is not the place to reintroduce one), its listening ports —
**the client port, which carries TLS itself** rather than standing beside a plaintext one as the
AMQP broker's 5671/5672 pair did, and the monitoring endpoint on loopback — and a configuration
file the controller composes (accounts, permissions, TLS, JetStream).
**How that file's changes reach the running server, corrected on revision.** First review: the
earlier draft named `reload-on` as the mechanism, citing the container runtime's own trust file as
precedent. `reload-on` is real, but it is a **service** field
([mesh-host declaration.go](https://git.novox.be/novox/mesh-host), `Service.ReloadOn` —
`docker.service` is reloaded via systemd, which is what the cited precedent actually does). A
**container** resource has no reload field at all — only `restart-on`, and a container's
`restart-on` is documented, exactly, to mean *recreate*. Declared as the earlier draft had it,
either the field is silently meaningless on a container resource or — if read as the nearest real
equivalent — every account, permission, or key change recreates the bus's own server: every
connection dropped, every in-flight JetStream ack lost, mid-flight the moment a module is added,
reassigned, or a person's access changes. For the one resource everything else depends on, that
is not an edge case; it is the common case.
**The fix asks nothing new of the host.** `nats-server` already reloads its own configuration
live on `SIGHUP` — accounts, permissions, everything in §4 — without dropping a connection; this
is the server's own documented capability, not something built for the mesh. So the composed
configuration file is mounted into a **directory** resource, not directly — a directory's contents
are not compared for change the way [issue 103](../../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md)'s
fix made a directly-mounted file's content, so a rewritten file inside it is not, on its own, a
reason to recreate the container. The image's own entrypoint watches that one file and sends
`nats-server` its own process `SIGHUP` when it changes — self-contained, inside the module, the
same place `modules/gitea/token.ts` keeps its own state rather than asking the host to model it.
The host's only job is what it already does for any directory resource: keep the file's content
current. Nothing is declared as `reload-on` or `restart-on` for this resource at all.
Its guard is the same rule as the AMQP broker's: the monitoring port is refused from anything but
the private network. It is raised at genesis like the store, adopted as a module in the same
phase. The predecessor's AMQP broker remains a module of its own, `lavinmq-compat`, with a
retirement condition: no client connected for a period the operator sets.
**And with one purpose, which it did not have when this was written.** Revision, second review
([ADR 0117](../../02-DECISIONS/0117-the-bus-is-the-only-broker.md)): two modules of the new mesh
required a broker *of their own* from it — a private vhost per consumer, the analog of
database-per-login, which is a different thing from the mesh's bus and was never examined here.
**There is one bus, and no module is handed a broker as a resource.** A module's messaging is
subjects on the bus under its own account, scoped by its `emits` and `consumes`; a module that
wants a queue of its own has a subject nothing else may publish to and a durable consumer, both
from its declaration. What it does not get is a server of its own. The `amqp` interface retires
with the compatibility broker instead of gaining a successor, and the two modules move to the bus
in step 4.
## 6. Joining: the enrolment handshake
Unchanged in shape, changed in transport. A node that has a token connects to the bus over TLS
with the **enrolment user** — a user that may publish `mesh.control.enrol` and subscribe its own
`_INBOX.enrol.<token-id>.>` and nothing else — publishes its request (the claim of the token, its
keys, its proof, its own reply subject as §2 now requires, and the found tunnel from
[ADR 0105](../../02-DECISIONS/0105-the-mesh-adopts-the-predecessors-tunnel-in-place.md)), and waits
on that inbox. The controller spends the token, records the node, composes the node's own user
into the server's configuration, and — reading the reply subject from the request's payload, never
from the transport `Reply` field the CONTROL consumer has already claimed for its own ack —
answers with the credentials sealed to the node's sealing key. The node reconnects as itself. The
enrolment user's permissions are what make a leaked token useless for anything but enrolling: it
cannot read a declaration or hear an event, and it cannot subscribe any inbox but the one its own
token derives.
## 7. A person's client
The operator asked for the mesh's tools from a workstation, and for it designed here rather than
bridged. It is three things:
1. **A person's account**: `operator issue <name>` on the controller creates a user whose
permissions are the tool subjects it may invoke — `mesh.tools.>` for an administrator, a list
for anyone else — and nothing on control, nodes or builds. It is issued, sealed to the person's
own key, and revoked, like a module's.
2. **A client that speaks the bus**: a small program on the workstation that connects as that user
over TLS, lists tools by asking the catalogue (`mesh.tools.mesh-catalog.catalog_tools`), and
turns each tool into a call — as an MCP server for an agent, and as a command line for a person.
It uses the sdk's `Broker` contract on the NATS runtime, so it is the same code path a module's
tools use, not a second protocol.
3. **Reachability**: the workstation reaches the bus over the private network once it is a node,
or over the predecessor's tunnel before that, on the bus's port; the guard and the openings
treat the bus as they do today.
Nothing is built of this before §10's bed passes; the MCP surface is a thin adapter over (2).
## 8. What a module sees, and what the wire does
**The contract a module is written against does not change.** `publish` on an envelope becomes a
publish on `mesh.events.<module>.<event>`; `subscribe` with a pattern becomes a durable JetStream
consumer on the matching subject filter; `request`/`handle` become a NATS request and a queue-group
subscription on `mesh.tools.<module>.<tool>`. The envelope's shape
([ADR 0042](../../02-DECISIONS/0042-the-shape-of-an-event-on-the-wire.md)) is unchanged; it is the
message body. A module built today runs on the new runtime without a rebuild — that is the test of
ADR 0039, and it is in §10.
**The wire underneath it changes completely, and that is a specification, not an implementation
detail.** Revision, second review ([ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md)):
an earlier draft of this section said "nothing new" and stopped there, which read as though the
change were contained inside the runtime. It is not.
[ADR 0074](../../02-DECISIONS/0074-the-wire-is-specified-not-the-types.md) settled that an SDK is an
implementation of a *specified wire*, checked by fixtures that must match byte for byte — because
two implementations that disagree about an envelope do not fail to compile, they ignore each other
while both keep running. That specification is
[design 19](19-the-module-protocol.md), and it is written in exchanges, a durable per-consumer queue
named `<node>.<module>.events`, and a shared `serve.<key>` queue. Every one of those is gone here.
So design 19 is rewritten from exchanges and queues to the subjects and streams of §2 and §3, its
fixtures are recaptured on NATS, and each SDK re-claims the capabilities it passes. That is step 3
of §9, and until it lands design 19 says in its own opening that its wire section describes the bus
being replaced. What is *not* rewritten: ADR 0074's model — protocol split per capability, an SDK
that implements the floor and events alone is legitimate, conformance executable per capability —
and ADR 0039's refusal. A new transport is when the pressure to grow the shared library is highest;
nothing is added to it here.
## 9. Moving from the bus the mesh has: five steps
Per ADR 0106 the bus moves once. Per
[ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md) the *build* is five steps,
each ending at something §10 proves, so that no part of this waits on the whole of it. Dividing the
build does not divide the bus: steps 1 to 4 leave every node on AMQP, and step 5 is still one
rollout.
**Step 1 — genesis raises the broker.** The `nats` module of §5, and a mesh raised on it from
nothing. This is built and proven although the mesh it is for will never travel this path: genesis
is where the foundation is *defined*, the one place the mesh comes from nothing, and the definition
every other path is measured against. A genesis path that exists only on paper is one nobody finds
wrong until there is a second mesh. *Ends at: the genesis bed.*
**Step 2 — adoption puts the broker in its seat.** A mesh already running does not get a foundation
module by being raised again; it adopts one in place
([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)). The
server is raised beside the AMQP broker on its own ports, carrying no mesh traffic yet, and the
`nats` module is adopted onto it. The seat it claims is **`mesh-broker`**, unchanged — the
foundation seats are named after the server's role rather than the product
([ADR 0079](../../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md)) for
exactly this case, and a seat named after the product would need renaming by every change the seat
exists to survive. *Ends at: the adoption bed.*
**Step 3 — the protocol gets its NATS binding.** §8's other half: design 19 rewritten from
exchanges and queues to subjects and streams, its conformance fixtures recaptured on NATS, each SDK
re-claiming the capabilities it passes. ADR 0074's model is unamended and ADR 0039's refusal holds —
no helper layer arrives with the new transport. A language may still arrive in pieces: connection
and events first, tools and provisioning when something needs them. *Ends at: the conformance suite,
per capability, per implementation.*
**Step 4 — the core speaks NATS.** The controller's link, the host's link, the tool runtime's
client — and with them the flows carried today by something other than the bus: a build source's
change reaching the builder, an installation, a module's own reports. Each is a conversion with a
named before and after. Observation — heartbeats, conditions, key-value state — belongs to
[research 017](../../01-RESEARCH/017-a-mesh-that-heals-itself/00-overview.md), which already
reserves it for after the move; this step does not pre-empt its design. Modules converted meanwhile
target the sdk contract and are untouched by any of it. *Ends at: each converted flow proved against
the behaviour it replaced.*
**Step 5 — the rollout.** Unchanged from ADR 0106 and previewed: the controller composes every
node's and module's account into the server standing since step 2, then rolls out the controller,
every host and every runtime built for NATS. Each node's host connects to the new bus as it comes up
and reports; the controller confirms every node heard before it stops listening on AMQP. The
predecessor's clients never notice — their broker is the compatibility module and stays. Then the
AMQP-side mesh accounts are removed from it, leaving only the predecessor's users, and it retires
when its retirement condition holds. The bus's port settings follow ADR 0100 like any port.
*Ends at: the cutover bed, then the rollout itself.*
What is not done, at any step: no dual-bus period for the mesh's own traffic, no bridge, no module
rebuilt.
**The work of these five, broken down and measured, is [design 28](28-building-the-bus.md)** —
including the two places their dependencies put a bed later than the step that names it.
## 10. How it is checked
**A bed per step, and each is green before the step after it starts** — the division in §9 is only
real if the proofs divide with it. A step that cannot name what its bed proves is not a step, and is
divided further before it is started.
**Step 1 — the genesis-broker bed**, a mesh raised from nothing, the server standing on it. The
mesh does not yet *live* on this bus — nothing speaks it until step 3's implementations exist — so
what this bed proves is the server, its configuration and the permissions, each of which the server
itself enforces and a plain client can therefore check:
- the four streams exist, asserted idempotently on a second start, with the retention of §3;
- every account and permission in the composed file is derived from the manifests' `emits` and
`consumes` and nothing else, with each user's own ack subject and its own inbox prefix;
- a user cannot publish outside its `emits` nor subscribe outside its `consumes` — refused by the
server, not by convention;
- a user cannot ack another user's delivery, and cannot subscribe another's inbox prefix — refused
by the server, not by the client's own good behaviour;
- the monitoring port is refused from anything but the private network;
- the `nats` container is not recreated when only its composed configuration file changes, and a
change to that file is live (a new user can connect, a revoked one cannot) within one
watcher-poll interval, without a restart.
**Step 2 — the adoption bed**, a mesh already running that has never had this server:
- the server is raised beside the AMQP broker on its own ports and the `nats` module is adopted
onto it in place, holding the data and the configuration it was raised with;
- the seat it claims is `mesh-broker`, and a second assignment of it anywhere in the mesh is
refused at resolution — *one per mesh*, as ADR 0079 requires;
- every node stays on AMQP throughout and nothing routes to the adopted server. This is the
check that makes steps 3 and 4 safe to run against a live mesh: adoption that quietly carried
traffic would be step 5 arriving early and unrehearsed.
**Step 3 — conformance, per capability, per implementation:**
- the fixtures of design 19, recaptured on NATS, produced and consumed byte for byte by every SDK
that claims the capability — the test ADR 0074 set, and the only one that catches two
implementations quietly ignoring each other;
- an SDK implementing connection and events alone passes those two and claims nothing more,
rather than failing as a whole;
- a module built before this design serves its tools unchanged on the new runtime — the test of
ADR 0039;
- the shared library gained nothing but the binding: its surface is the protocol and the
primitives, and a helper that arrived with the transport is a review failure, not a detail.
**Step 4 — the mesh living on it.** The implementations exist from step 3, so this is where a mesh
can first be raised on NATS and run:
- a node enrols over TLS with a claimed token, and the enrolment user cannot read a declaration;
- a push composes; the store is stopped; the push is held (nak with delay), the store returns, the
push applies, nothing was lost or duplicated;
- a node that was away gets exactly the newest declaration, and a replayed older one is refused
by sequence;
- an upgrade rolls out to two nodes;
- an enrolment request held by a `nak`-with-delay cycle still reaches the enrolling node's inbox
once the controller answers — proving the reply travels in the payload and not the transport
field a consumer's ack has already claimed;
- an event whose consumer keeps failing dead-letters after `max-deliver`;
- a module's tool is invoked from another node and from a person's client, each with an account
that can invoke it, and refused from one that cannot;
- a build source's change reaches the builder over the bus, and the build that follows is the one
the change asked for;
- an installation completes over the bus, with the same outcome the path it replaces produced;
- a node that was unreachable catches up on its reports rather than losing them.
**Step 5 — the cutover bed:** a mesh on AMQP with a predecessor stand-in on the compatibility
broker moves its bus in one rollout; every node reports on NATS afterwards; the stand-in's client
on AMQP is still connected throughout.
Unit tests hold the controller to composing accounts from `emits`/`consumes` and nothing else, to
creating the four streams and asserting them idempotently, and to spending a token exactly once;
the host to connecting as the enrolment user with nothing but enrolment permissions; the runtime
to mapping the sdk contract onto subjects exactly as §8 says.
## 11. Open, for the review
**Closed by the second revision** (2026-09-26, [ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md)):
the build was one undivided item (§9 is now five steps, each with its own bed in §10); there was no
adoption path for a mesh already running (§9 step 2); and §8 said a module sees "nothing new"
without distinguishing the sdk's contract, which does not change, from the specified wire, which
changes entirely and is design 19's (§8, §9 step 3). Two smaller corrections: the seat is
`mesh-broker` and not the product's name, and the shared library gains no conveniences with the new
transport.
**Closed by the first revision** (recorded in `MIGRATION-LOG.md`, 2026-09-24): the
`reload-on`/container mismatch (§5), the eaten reply subject on a CONTROL-stream message (§2, §6),
the missing ack permission (§4), and the un-scoped reply inbox under one account (§4). Each is
named where it was wrong, not silently fixed, so a reader comparing against the first version can
find what changed and why.
**Still open:**
- ~~Whether EVENTS should be one stream or one per emitting module.~~ **Closed**
([ADR 0117](../../02-DECISIONS/0117-the-bus-is-the-only-broker.md)): one stream, and not as a
preference — the streams the bus is made of are composed as configuration before any module
runs, because a provisioner is itself a module that needs a bus account to start. Bootstrapping
decides it.
- The heartbeat interval and the controller's "quiet" threshold on core NATS without persistence —
the same numbers as today are proposed.
- Whether the person's client is a catalogue module (runs on an enrolled workstation node) or a
standalone program (runs anywhere with credentials). Both, in that order, is proposed.
- Leaf nodes: a NATS leaf per machine would make every module's connection local and survive the
hub's restart. Deliberately out of scope; noted so it is not forgotten.
- **New, from this revision:** the `nats` image's own entrypoint now carries logic (watch a file,
signal a process) that no other module's container needed before. Is a one-file-watcher-and-
`SIGHUP` helper common enough across future modules with the same shape (a service that reloads
on `SIGHUP` but runs in a container) to belong in `mesh-sdk` rather than written once per module
that needs it? [ADR 0039](../../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md)'s test —
*does editing it recompile unrelated modules, and does it change often* — probably says no for
one instance; worth asking again if a second module needs the same shape.