Files
hq/03-DESIGN/01-to-be/25-the-bus-on-nats.md
jschoubben ce6ae943b7 Merge main: renumber this branch's records around the trunk's
Both lines of work numbered from the same point, so four decision records and one design
document existed twice with different content. The trunk keeps its numbers and this branch
yields — the only rule that scales, because the trunk's are already cited by what merged
before them.

  0117 the bus is the only broker        -> 0125
  0118 a module declares its own seats   -> 0126
  0119 amqp is a provision, not the bus  -> 0127
  0120 the mesh bus is required          -> 0128
  0123 a seat carries its role's protocol -> 0129
  0124 the predecessor is ending          -> 0130
  design 29, what a module declares       -> design 32

Applied to the code repositories too, because a stale reference is worse when numbers
collide than when they dangle: the reader lands on a real record that decided something
else.

Two reconciliations the merge forced, both real:

**0110 was marked wholly superseded and was not.** Its successor says in as many words that
everything 0110 decided about what a seat *is* stands untouched — and two records that
landed on the trunk rest on exactly that part. So it is accepted again, extended rather than
replaced, with a note saying which of its claims moved and where.

**A seat's protocol becomes columns, not fields.** The trunk moved the seat set out of
compiled code into a table the controller owns. This branch had added what a role accepts,
emits and serves to the Go slice. The decision is unaffected and the mechanism is better for
it: giving a role a protocol is now a write rather than a rebuild, which is the trunk's own
argument applied to what this branch added.

One check still fails and it fails on main too: a record resting on ADR 0112 while that is
still 'proposed'. Left alone — it is not this merge's to answer.
2026-09-27 18:23:41 +02:00

554 lines
39 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
layer: to-be
status: in-progress
code:
- mesh-controller internal/link (to be replaced)
- mesh-host internal/link (to be replaced)
- mesh-tools src/broker-amqp.ts (to be replaced)
- mesh-catalog modules/nats (to be written)
- mesh-sdk src (the protocol's NATS binding, step 3)
updated: 2026-09-27
decisions:
- 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md
- 02-DECISIONS/0116-the-bus-is-built-in-five-steps.md
- 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md
- 02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md
- 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
- 02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md
- 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
- 02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md
- 02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md
- 02-DECISIONS/0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md
---
# 25. The bus on NATS
**Status: proposed — a design to be reviewed before any code.** This is the architecture
[ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md) asks for. It says what rides the bus,
under which subject, with which guarantee, under whose account; how a node joins; how a person
reaches a tool; how the mesh moves from the bus it has to this one; and how each claim is checked.
Prose and diagrams only; no configuration is pasted.
## 1. What the bus is for
**The bus is where the mesh happens.** Not a transport the mesh sends things over — the place a
module is reachable at all, where a role is addressed without knowing who holds it, where the
mesh's own state lives, and where what a module may say is decided by what it declared.
| Traffic | Shape | Guarantee it needs |
|---|---|---|
| **control** — a node's report, a build's outcome, an enrolment | job | nothing lost while the store restarts; retried; in order per node |
| **heartbeat** — a node saying it is alive | fire and forget | none; a lost one is the next one |
| **declarations** — the controller tells a node what to be | state | the node gets the newest; a stale one is never applied |
| **builds** — work for the build machine | job | at least once, one worker at a time |
| **events** — a module says something happened | 1:many | delivered to every consumer that declared it; dead-lettered when it cannot be |
| **tools** — a module or a person asks another's tool | request/reply | one answer, from one server, or a timeout |
| **work to a role** — a module submits to a capability without knowing who provides it | job | exactly one holder does it; it queues while nobody does |
The last two rows are the ones worth dwelling on, because they are not messaging in the sense of
carrying bytes from A to B. **A role is addressable**, so a caller names the capability and never
the module or the node — and the implementation can be replaced under it without a caller
changing ([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md)). That is a
property of the mesh's architecture that happens to be expressed in subjects.
And more of the mesh lands here as it is built: conditions and observed state in key-value
buckets that anything may watch, the server's own advisories becoming observations like any other
([research 017](../../01-RESEARCH/017-a-mesh-that-heals-itself/00-overview.md)), and a person's
client speaking the bus directly rather than through a surface built over it (§7). None of that
is a message being moved; all of it is the bus being the mesh's centre.
**What a module sees of it is small and derived.** It declares what it emits, consumes, serves
and uses, and the subjects, streams, consumers and permissions all follow from that
([design 29](32-what-a-module-declares.md)). The sdk's contract — `request`, `handle`, `publish`,
`subscribe`, `close` — is the whole surface, and it does not change
([ADR 0039](../../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md)).
## 2. Subjects
NATS addresses everything by subject. The mesh's subject space is one tree, and every account's
permissions are expressed as which branches of it that account may publish to and subscribe from.
```
mesh.control.<node>.report a node's report (JetStream: CONTROL)
mesh.control.<node>.alive heartbeat (core, no persistence)
mesh.control.enrol an enrolment request (JetStream: CONTROL)
mesh.node.<node>.declare a declaration for a node (JetStream: NODES, last-per-subject)
mesh.mod.<module>.event.<event> an event (JetStream: EVENTS)
mesh.mod.<module>.tool.<tool> a tool invocation (core request/reply)
mesh.seat.<seat>.accept.<verb> work submitted to a role (JetStream: per-seat work queue)
mesh.seat.<seat>.event.<verb> a role's own event (JetStream: EVENTS)
mesh.seat.<seat>.tool.<verb> a role's tool (core request/reply)
mesh.ask.<node>.<command> the controller's command api (core request/reply)
```
**Revised 2026-09-27** ([ADR 0129](../../02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md)):
**`mesh.build.request`, `mesh.control.built` and the BUILDS stream are gone.** A build is work submitted to a role, and the
mesh already has a shape for that — a seat's `accept` subjects, on a work queue with a queue group of
holders, which is what a build queue shared by several machines *is*. Keeping a second mechanism for it
meant two things to reason about and two places for a permission to be wrong. A build's outcome is the
seat's own event — which is why the control branch loses its copy too: one publish reaches whoever
asked, the controller that records it and the catalogue that places it — the fan-out a shared exchange gave for free, as a derived subject rather than a
configured topology.
**Revised 2026-09-26** ([design 29](32-what-a-module-declares.md)): a module's events and tools
moved from `mesh.events.*` / `mesh.tools.*` into one namespace per module, `mesh.mod.<module>.>`,
so a module's authority over its own name is a single subject pattern the server enforces — and
each carries a **kind token**, without which an events stream's filter would capture tool calls.
Seats are the same shape, one namespace per role.
Two things this buys over the exchanges: **request/reply is native** — a tool call is one
`request` on `mesh.tools.<module>.<tool>` answered by whichever runtime serves it (a queue group per
tool, so several nodes may serve one tool); and **a declaration is last-per-subject** — the NODES
stream keeps only the newest message on `mesh.node.<node>.declare`, so a node that was away gets
exactly the current declaration and nothing older. That is the wire-level answer to
[issue 107](../../04-ISSUES/107-a-declaration-carries-no-order/00-report.md): the stream's sequence
*is* the order, and a node that sees sequence n refuses n−1 by construction.
**A reply-to travelling through a JetStream stream is carried in the payload, never in the
transport `Reply` field.** *Verified against a running server, 2026-09-27*: a caller published
asking for a reply to `_INBOX.LCr3M83q…`, and the consumer saw a `Reply` field of
`$JS.ACK.PROBE.probe_consumer.1.1.1…`. The address is replaced, not merely at risk — so the
payload-borne reply subject below is necessary rather than defensive, and the check is a test
rather than a note, because a future server that stopped doing this would leave enrolment
working and the reason for the field quietly becoming folklore.
Revision, first review: core NATS request/reply sets the requester's
ephemeral inbox as the message's `Reply` field, and a plain responder answers it directly — but a
message a JetStream consumer delivers has already had that field claimed for the consumer's own
ack address (`$JS.ACK.<stream>.<consumer>...`), so by the time the controller (§3's CONTROL
consumer) sees the message, `Reply` names where *it* must ack, not where the original caller is
waiting. `mesh.control.enrol` is the case that matters: a synchronous-feeling caller waiting on an
ephemeral inbox, over a subject the store-window guarantee may legitimately delay by several
`nak` cycles — exactly the combination that would otherwise deliver the answer to a caller who
has long since timed out and unsubscribed. So every CONTROL message that expects an answer states
its reply subject as an ordinary field of its own payload; the controller reads it from there and
publishes the answer to it explicitly, never via `Respond()`. Nothing else in this design routes
a reply through a stream — tools and heartbeats stay on core NATS, where `Reply` means what it has
always meant.
## 3. Streams, and the guarantees they carry
Core NATS is at-most-once. Everything the mesh must not lose lives in a JetStream stream:
| Stream | Subjects | Retention | Why |
|---|---|---|---|
| CONTROL | `mesh.control.>` except `alive` (a build's outcome moved to its seat, ADR 0121) | work queue, one consumer (the controller), explicit ack | the store-window guarantee ([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)): the controller `nak`s with a delay while its store is away and the message is redelivered; nothing is dropped |
| NODES | `mesh.node.>` | last per subject | one declaration per node, always the newest |
| EVENTS | `mesh.mod.*.event.>` | limits (age, size), durable consumer per subscribing module | a subscriber that was down catches up; after `max-deliver` attempts the advisory feeds `mesh.events.dead` (its own small stream) |
Tool calls and heartbeats stay on core NATS: a lost heartbeat is the next heartbeat; a lost tool
call is a timeout the caller already handles.
Streams and consumers are objects the controller creates at genesis and asserts on start; a module
declares nothing about them. The controller is the only writer of stream definitions.
### The store window, and what moving it into the server changes
The guarantee ([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)) is
that a push the controller cannot record because its store is restarting is **held and retried** —
never dropped, never falsely acknowledged. Here that is a `nak` with a delay: the server holds
the message and redelivers it, so the controller keeps no list of parked messages and one that
restarts mid-window loses nothing it was holding.
**That is a plain win, and it introduces one problem worth naming.** Holding a delivery in memory
let the controller drop an older report when a newer one for the same node arrived, because
acting on the older after the newer would undo the newer. A `nak`ed message belongs to the server
and comes back whatever happened meanwhile — so the older report is redelivered *after* the newer
was applied.
The answer was already in the message. A report carries the **digest of the declaration it is
about**, which exists because an earlier attempt to order reports by time lost the race it
invited: an apply that began under the previous declaration finishes after the next is sent, and
its report reads as newer than the send. Clocks cannot answer *which*.
So supersession stops being something the controller remembers and becomes something it checks —
a report whose digest is not the one outstanding for that node is acknowledged without being
acted on. The same shape as a node refusing a superseded declaration by sequence
([issue 107](../../04-ISSUES/107-a-declaration-carries-no-order/00-report.md)): **ordering settled
by what a message says, not by when it arrived.** And staleness is checked before the store is
waited on, so a redelivery that lost its race does not hold a slot in the window that a current
message needs.
## 4. Accounts
[ADR 0043](../../02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) says a
module's account may publish only what it `emits` and consume only what it `consumes`. NATS
expresses this exactly, per subject, and better than a vhost could:
- **One NATS account for the mesh.** Accounts in NATS isolate subject spaces entirely; the mesh
is one space, so it is one account. The predecessor's compatibility broker is not on this bus at
all.
**One account is a choice with a cost, stated plainly on revision:** none of NATS's own
isolation is free here, because there is only the one subject space, for everyone. Two
consequences that a single-account mesh must therefore grant on purpose, not by omission:
- **A durable consumer needs permission to ack, or it never really consumes.** Acking a
JetStream delivery is a publish to that consumer's own ack-reply address
(`$JS.ACK.<stream>.<consumer>.>`), a different subject from anything the consumer subscribes.
A module's user is therefore granted publish on `$JS.ACK.EVENTS.<module>.>` as well as its
emits — scoped to the one consumer name the controller derives for that module, so a module
can ack only its own deliveries. Without this, first review found, every message it receives
would be redelivered forever: refused by the permission list it already has.
- **A reply inbox needs a subject nothing else can guess or enumerate.** With one account,
inbox privacy is the permission list or it is nothing — there is no second account backing
it up. So no user is ever granted a bare `_INBOX.>`. Each user's inbox subject is derived
from its own identity (`_INBOX.<module>.<node>.>`, or `_INBOX.person.<name>.>`), and its
permissions name only that one prefix, for the reply to any request it makes and nothing
wider. First review found the account note without this and read it as "any user may
subscribe any inbox" — which was accurate against the text as it stood.
**And a scoped inbox needs `allow_responses`, or nothing can answer.** Revision, found while
composing the first real configuration: the rule above scopes each user's inbox to itself,
which is right — and leaves a responder unable to reply, because the answer goes to the
*caller's* inbox, which the responder has no permission for. The two ways out are granting
every responder `_INBOX.>`, which is exactly the blanket grant this bullet refuses, or NATS's
own `allow_responses`: the server permits one reply to the reply-subject of a message the user
actually received, within a TTL, and nothing else. So authority to answer is bounded by having
been asked, and only principals that serve something are granted it — a pure consumer gets
nothing. Without this the scoping is not merely incomplete: every tool call in the mesh times
out, and the permission list looks correct while it happens.
- **One user per module per node**, as today, with publish permissions
`mesh.events.<module>.<event>` for each emit, `mesh.tools.<module>.>` to serve its tools, its
own ack-reply subject for each durable consumer it holds, and its own inbox prefix; subscribe
permissions for each consumed event's subject, its tool subjects, and that same inbox prefix.
Nothing else. A module that tries to publish outside its emits is refused by the server, not by
convention.
- **The controller's user** owns `mesh.control.>`, `mesh.node.>` and the streams, and may submit work
to the seats the mesh's own flows use — a build, for one (ADR 0121).
**A host's user** may publish its own `mesh.control.<node>.>` and subscribe its own
`mesh.node.<node>.declare` — and nothing of any other node's.
- **A person's user** (§7) is a module-shaped user with permissions on the tool subjects it may
invoke, issued and revoked by the controller like any account.
**The server does not verify client certificates, and TLS is still required.** *Revision,
2026-09-27, found by building the module's image and connecting to it as a host would.* The first
composed configuration said `verify: true`, which makes the server demand a **client** certificate —
and nothing in the mesh presents one. A host pins this server's exact certificate and authenticates
with the password the mesh minted ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)),
and a module's runtime does the same. With it on, every connection in the mesh dies at the TLS
handshake before any password is looked at, and the error — "client didn't provide a certificate" —
reads as a fault in the client rather than in the bus's configuration. The `tls` block is what makes
TLS required; `verify` only decides whether client certificates are checked. What is given up is a
second factor the mesh has no machinery to issue or rotate — a certificate per module per node — and
what is kept is stronger than a name check in both directions: an exact pin outward, a password
scoped per user inward. **Mutual TLS is a later question and would need that machinery first.**
**The mesh composes the accounts; the module composes its server.** *Revision, 2026-09-27, while
building the composition.* An earlier reading of the paragraph below had the controller writing the
whole file. It writes only the user list. A server's ports, its TLS paths and its store directory
are properties of the container the module raises — they live in its image and its mounts and change
when it does — so the module declares its own configuration and `include`s the mesh's half. A
controller that wrote the whole file would have to be kept in step with a Dockerfile it never sees,
and a module could not change its own image without the mesh agreeing. Asking for the user list is
not enough to receive it: the file holds every user's password hash, so the claim on `mesh-broker` is
what authorises it. And the two files share one directory of necessity — an absolute include path is
resolved relative to the including file's own directory, so a server given one from elsewhere looks
for it underneath that directory and refuses to start.
**Accounts are configuration, not API calls.** The controller composes the mesh's user list and its
permissions into a file the host declares. **How that file reaches the running server is §5's,
not this one's** — revision, first review: an earlier draft said "reloads" and cited a precedent
that does not apply to a container (see §5). No management API, no credential travelling through a
management call, and the [issue 102](../../04-ISSUES/102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md)
discipline from the first day: an address or a permission is read where it is used, never stored
with a port. Passwords are minted and sealed exactly as today; the file holds bcrypt hashes.
Alternative considered and not taken: the operator/JWT model (`nsc`), where accounts are signed
tokens resolved by the server. It is the right model for a multi-tenant NATS; the mesh is one
tenant, already has a sealing key and a controller that writes files, and would gain a second
signing hierarchy for nothing.
## 5. The broker as a module
`nats` is a catalogue module claiming the seat `mesh-broker`
([ADR 0079](../../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md): the seat
is the server, and the server changes). It declares one container (a single binary; **JetStream on a host
directory bind, not a named volume** — revision, second review:
[issue 115](../../04-ISSUES/115-a-named-docker-volume-is-invisible-and-one-flag-from-gone/00-report.md)
is resolved, and converted the store, the broker and two others away from named volumes for the
reason it names; the bus's own data is not the place to reintroduce one), its listening ports —
**the client port, which carries TLS itself** rather than standing beside a plaintext one as the
AMQP broker's 5671/5672 pair did, and the monitoring endpoint on loopback — and a configuration
file the controller composes (accounts, permissions, TLS, JetStream).
**How that file's changes reach the running server, corrected on revision.** First review: the
earlier draft named `reload-on` as the mechanism, citing the container runtime's own trust file as
precedent. `reload-on` is real, but it is a **service** field
([mesh-host declaration.go](https://git.novox.be/novox/mesh-host), `Service.ReloadOn` —
`docker.service` is reloaded via systemd, which is what the cited precedent actually does). A
**container** resource has no reload field at all — only `restart-on`, and a container's
`restart-on` is documented, exactly, to mean *recreate*. Declared as the earlier draft had it,
either the field is silently meaningless on a container resource or — if read as the nearest real
equivalent — every account, permission, or key change recreates the bus's own server: every
connection dropped, every in-flight JetStream ack lost, mid-flight the moment a module is added,
reassigned, or a person's access changes. For the one resource everything else depends on, that
is not an edge case; it is the common case.
**The fix asks nothing new of the host.** `nats-server` already reloads its own configuration
live on `SIGHUP` — accounts, permissions, everything in §4 — without dropping a connection; this
is the server's own documented capability, not something built for the mesh. So the composed
configuration file is mounted into a **directory** resource, not directly — a directory's contents
are not compared for change the way [issue 103](../../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md)'s
fix made a directly-mounted file's content, so a rewritten file inside it is not, on its own, a
reason to recreate the container. The image's own entrypoint watches that one file and sends
`nats-server` its own process `SIGHUP` when it changes — self-contained, inside the module, the
same place `modules/gitea/token.ts` keeps its own state rather than asking the host to model it.
The host's only job is what it already does for any directory resource: keep the file's content
current. Nothing is declared as `reload-on` or `restart-on` for this resource at all.
Its guard is the same rule as the deprecated broker's: the monitoring port is refused from anything but
the private network. It is raised at genesis like the store, adopted as a module in the same
phase.
**The deprecated broker is an ordinary module, not a compatibility layer.** Revision, second review
([ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md)): earlier text here, and
ADR 0106 before it, called it `lavinmq-compat` — one purpose, the predecessor's clients, and a
retirement condition of no client connected for a period the operator sets. It is none of those.
A module may legitimately need an AMQP broker as a **backing service**, the way it needs a
database, and the provider that answers that is an ordinary module like any other: no seat, not
foundation, never raised at genesis, installed when something wants it and absent from a mesh
that does not. There is no retirement condition, because the day its last client disappears is
not a day anything is waiting for.
**Revised 2026-09-27** ([ADR 0130](../../02-DECISIONS/0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md)):
**that day is coming.** The predecessor is deprecated — some of it still running, none of it being
migrated, left to stop rather than moved — so the broker retires once nothing requires `amqp`. Still no
retirement *condition* and no end-date machinery: a provision with no consumers has its provider
unassigned, which is the ordinary mechanism and is ADR 0127 being paid off rather than revised. What
also goes with it is the tooling that reaches this installation's machines remotely, because the
predecessor's own mesh talks over that broker — so the rollout is driven from the node, or before the
broker stops.
What is deprecated is AMQP as **the mesh's transport**, which is this whole document. The rule
that remains is about direction rather than software: *inter-module communication goes over the
bus.* A module may hold a broker, a database or a cache for itself; it may not use one as a
channel to another module.
## 6. Joining: the enrolment handshake
Unchanged in shape, changed in transport. A node that has a token connects to the bus over TLS
with the **enrolment user** — a user that may publish `mesh.control.enrol` and subscribe its own
`_INBOX.enrol.<token-id>.>` and nothing else — publishes its request (the claim of the token, its
keys, its proof, its own reply subject as §2 now requires, and the found tunnel from
[ADR 0105](../../02-DECISIONS/0105-the-mesh-adopts-the-predecessors-tunnel-in-place.md)), and waits
on that inbox. The controller spends the token, records the node, composes the node's own user
into the server's configuration, and — reading the reply subject from the request's payload, never
from the transport `Reply` field the CONTROL consumer has already claimed for its own ack —
answers with the credentials sealed to the node's sealing key. The node reconnects as itself. The
enrolment user's permissions are what make a leaked token useless for anything but enrolling: it
cannot read a declaration or hear an event, and it cannot subscribe any inbox but the one its own
token derives.
## 7. A person's client
The operator asked for the mesh's tools from a workstation, and for it designed here rather than
bridged. It is three things:
1. **A person's account**: `operator issue <name>` on the controller creates a user whose
permissions are the tool subjects it may invoke — `mesh.tools.>` for an administrator, a list
for anyone else — and nothing on control, nodes or builds. It is issued, sealed to the person's
own key, and revoked, like a module's.
2. **A client that speaks the bus**: a small program on the workstation that connects as that user
over TLS, lists tools by asking the catalogue (`mesh.tools.mesh-catalog.catalog_tools`), and
turns each tool into a call — as an MCP server for an agent, and as a command line for a person.
It uses the sdk's `Broker` contract on the NATS runtime, so it is the same code path a module's
tools use, not a second protocol.
3. **Reachability**: the workstation reaches the bus over the private network once it is a node,
or over the predecessor's tunnel before that, on the bus's port; the guard and the openings
treat the bus as they do today.
Nothing is built of this before §10's bed passes; the MCP surface is a thin adapter over (2).
## 8. What a module sees, and what the wire does
**The contract a module is written against does not change.** `publish` on an envelope becomes a
publish on `mesh.events.<module>.<event>`; `subscribe` with a pattern becomes a durable JetStream
consumer on the matching subject filter; `request`/`handle` become a NATS request and a queue-group
subscription on `mesh.tools.<module>.<tool>`. The envelope's shape
([ADR 0042](../../02-DECISIONS/0042-the-shape-of-an-event-on-the-wire.md)) is unchanged; it is the
message body. A module built today runs on the new runtime without a rebuild — that is the test of
ADR 0039, and it is in §10.
**The wire underneath it changes completely, and that is a specification, not an implementation
detail.** Revision, second review ([ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md)):
an earlier draft of this section said "nothing new" and stopped there, which read as though the
change were contained inside the runtime. It is not.
[ADR 0074](../../02-DECISIONS/0074-the-wire-is-specified-not-the-types.md) settled that an SDK is an
implementation of a *specified wire*, checked by fixtures that must match byte for byte — because
two implementations that disagree about an envelope do not fail to compile, they ignore each other
while both keep running. That specification is
[design 19](19-the-module-protocol.md), and it is written in exchanges, a durable per-consumer queue
named `<node>.<module>.events`, and a shared `serve.<key>` queue. Every one of those is gone here.
So design 19 is rewritten from exchanges and queues to the subjects and streams of §2 and §3, its
fixtures are recaptured on NATS, and each SDK re-claims the capabilities it passes. That is step 3
of §9, and until it lands design 19 says in its own opening that its wire section describes the bus
being replaced. What is *not* rewritten: ADR 0074's model — protocol split per capability, an SDK
that implements the floor and events alone is legitimate, conformance executable per capability —
and ADR 0039's refusal. A new transport is when the pressure to grow the shared library is highest;
nothing is added to it here.
## 9. Moving from the bus the mesh has: five steps
Per ADR 0106 the bus moves once. Per
[ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md) the *build* is five steps,
each ending at something §10 proves, so that no part of this waits on the whole of it. Dividing the
build does not divide the bus: steps 1 to 4 leave every node on AMQP, and step 5 is still one
rollout.
**Step 1 — genesis raises the broker.** The `nats` module of §5, and a mesh raised on it from
nothing. This is built and proven although the mesh it is for will never travel this path: genesis
is where the foundation is *defined*, the one place the mesh comes from nothing, and the definition
every other path is measured against. A genesis path that exists only on paper is one nobody finds
wrong until there is a second mesh. *Ends at: the genesis bed.*
**Step 2 — adoption puts the broker in its seat.** A mesh already running does not get a foundation
module by being raised again; it adopts one in place
([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)). The
server is raised beside the deprecated broker on its own ports, carrying no mesh traffic yet, and the
`nats` module is adopted onto it. The seat it claims is **`mesh-broker`**, unchanged — the
foundation seats are named after the server's role rather than the product
([ADR 0079](../../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md)) for
exactly this case, and a seat named after the product would need renaming by every change the seat
exists to survive. *Ends at: the adoption bed.*
**Step 3 — the protocol gets its NATS binding.** §8's other half: design 19 rewritten from
exchanges and queues to subjects and streams, its conformance fixtures recaptured on NATS, each SDK
re-claiming the capabilities it passes. ADR 0074's model is unamended and ADR 0039's refusal holds —
no helper layer arrives with the new transport. A language may still arrive in pieces: connection
and events first, tools and provisioning when something needs them. *Ends at: the conformance suite,
per capability, per implementation.*
**Step 4 — the core speaks NATS.** The controller's link, the host's link, the tool runtime's
client — and with them the flows carried today by something other than the bus: a build source's
change reaching the builder, an installation, a module's own reports. Each is a conversion with a
named before and after. Observation — heartbeats, conditions, key-value state — belongs to
[research 017](../../01-RESEARCH/017-a-mesh-that-heals-itself/00-overview.md), which already
reserves it for after the move; this step does not pre-empt its design. Modules converted meanwhile
target the sdk contract and are untouched by any of it. *Ends at: each converted flow proved against
the behaviour it replaced.*
**Step 5 — the rollout.** Unchanged from ADR 0106 and previewed: the controller composes every
node's and module's account into the server standing since step 2, then rolls out the controller,
every host and every runtime built for NATS. Each node's host connects to the new bus as it comes up
and reports; the controller confirms every node heard before it stops listening on AMQP. The
predecessor's clients never notice — their broker is the compatibility module and stays. Then the
AMQP-side mesh accounts are removed from it, leaving only the predecessor's users, and it retires
when its retirement condition holds. The bus's port settings follow ADR 0100 like any port.
*Ends at: the cutover bed, then the rollout itself.*
What is not done, at any step: no dual-bus period for the mesh's own traffic, no bridge, no module
rebuilt.
**The work of these five, broken down and measured, is [design 28](28-building-the-bus.md)** —
including the two places their dependencies put a bed later than the step that names it.
## 10. How it is checked
**A bed per step, and each is green before the step after it starts** — the division in §9 is only
real if the proofs divide with it. A step that cannot name what its bed proves is not a step, and is
divided further before it is started.
**Step 1 — the genesis-broker bed**, a mesh raised from nothing, the server standing on it. The
mesh does not yet *live* on this bus — nothing speaks it until step 3's implementations exist — so
what this bed proves is the server, its configuration and the permissions, each of which the server
itself enforces and a plain client can therefore check:
- the four streams exist, asserted idempotently on a second start, with the retention of §3;
- every account and permission in the composed file is derived from the manifests' `emits` and
`consumes` and nothing else, with each user's own ack subject and its own inbox prefix;
- a user cannot publish outside its `emits` nor subscribe outside its `consumes` — refused by the
server, not by convention;
- a user cannot ack another user's delivery, and cannot subscribe another's inbox prefix — refused
by the server, not by the client's own good behaviour;
- the monitoring port is refused from anything but the private network;
- the `nats` container is not recreated when only its composed configuration file changes, and a
change to that file is live (a new user can connect, a revoked one cannot) within one
watcher-poll interval, without a restart.
**Step 2 — the adoption bed**, a mesh already running that has never had this server:
- the server is raised beside the deprecated broker on its own ports and the `nats` module is adopted
onto it in place, holding the data and the configuration it was raised with;
- the seat it claims is `mesh-broker`, and a second assignment of it anywhere in the mesh is
refused at resolution — *one per mesh*, as ADR 0079 requires;
- every node stays on AMQP throughout and nothing routes to the adopted server. This is the
check that makes steps 3 and 4 safe to run against a live mesh: adoption that quietly carried
traffic would be step 5 arriving early and unrehearsed.
**Step 3 — conformance, per capability, per implementation:**
- the fixtures of design 19, recaptured on NATS, produced and consumed byte for byte by every SDK
that claims the capability — the test ADR 0074 set, and the only one that catches two
implementations quietly ignoring each other;
- an SDK implementing connection and events alone passes those two and claims nothing more,
rather than failing as a whole;
- a module built before this design serves its tools unchanged on the new runtime — the test of
ADR 0039;
- the shared library gained nothing but the binding: its surface is the protocol and the
primitives, and a helper that arrived with the transport is a review failure, not a detail.
**Step 4 — the mesh living on it.** The implementations exist from step 3, so this is where a mesh
can first be raised on NATS and run:
- a node enrols over TLS with a claimed token, and the enrolment user cannot read a declaration;
- a push composes; the store is stopped; the push is held (nak with delay), the store returns, the
push applies, nothing was lost or duplicated;
- a node that was away gets exactly the newest declaration, and a replayed older one is refused
by sequence;
- an upgrade rolls out to two nodes;
- an enrolment request held by a `nak`-with-delay cycle still reaches the enrolling node's inbox
once the controller answers — proving the reply travels in the payload and not the transport
field a consumer's ack has already claimed;
- an event whose consumer keeps failing dead-letters after `max-deliver`;
- a module's tool is invoked from another node and from a person's client, each with an account
that can invoke it, and refused from one that cannot;
- a build source's change reaches the builder over the bus, and the build that follows is the one
the change asked for;
- an installation completes over the bus, with the same outcome the path it replaces produced;
- a node that was unreachable catches up on its reports rather than losing them.
**Step 5 — the cutover bed:** a mesh on AMQP with a predecessor stand-in on the compatibility
broker moves its bus in one rollout; every node reports on NATS afterwards; the stand-in's client
on AMQP is still connected throughout.
Unit tests hold the controller to composing accounts from `emits`/`consumes` and nothing else, to
creating the four streams and asserting them idempotently, and to spending a token exactly once;
the host to connecting as the enrolment user with nothing but enrolment permissions; the runtime
to mapping the sdk contract onto subjects exactly as §8 says.
## 11. Open, for the review
**Closed by the second revision** (2026-09-26, [ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md)):
the build was one undivided item (§9 is now five steps, each with its own bed in §10); there was no
adoption path for a mesh already running (§9 step 2); and §8 said a module sees "nothing new"
without distinguishing the sdk's contract, which does not change, from the specified wire, which
changes entirely and is design 19's (§8, §9 step 3). Two smaller corrections: the seat is
`mesh-broker` and not the product's name, and the shared library gains no conveniences with the new
transport.
**Closed by the first revision** (recorded in `MIGRATION-LOG.md`, 2026-09-24): the
`reload-on`/container mismatch (§5), the eaten reply subject on a CONTROL-stream message (§2, §6),
the missing ack permission (§4), and the un-scoped reply inbox under one account (§4). Each is
named where it was wrong, not silently fixed, so a reader comparing against the first version can
find what changed and why.
**Still open:**
- ~~Whether EVENTS should be one stream or one per emitting module.~~ **Closed**
([ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md) (superseding [ADR 0125](../../02-DECISIONS/0125-the-bus-is-the-only-broker.md))): one stream, and not as a
preference — the streams the bus is made of are composed as configuration before any module
runs, because a provisioner is itself a module that needs a bus account to start. Bootstrapping
decides it.
- The heartbeat interval and the controller's "quiet" threshold on core NATS without persistence —
the same numbers as today are proposed.
- Whether the person's client is a catalogue module (runs on an enrolled workstation node) or a
standalone program (runs anywhere with credentials). Both, in that order, is proposed.
- Leaf nodes: a NATS leaf per machine would make every module's connection local and survive the
hub's restart. Deliberately out of scope; noted so it is not forgotten.
- **New, from this revision:** the `nats` image's own entrypoint now carries logic (watch a file,
signal a process) that no other module's container needed before. Is a one-file-watcher-and-
`SIGHUP` helper common enough across future modules with the same shape (a service that reloads
on `SIGHUP` but runs in a container) to belong in `mesh-sdk` rather than written once per module
that needs it? [ADR 0039](../../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md)'s test —
*does editing it recompile unrelated modules, and does it change often* — probably says no for
one instance; worth asking again if a second module needs the same shape.