Taken during the outage of 2026-09-27, when the protocol leaked into the seat's contract: to hold mesh-broker a module had to provide amqp, so the module that will carry the bus could not hold the seat that names the bus, while the module being retired could. Supersedes 0127. Modules depend on the seat and reach the bus through the sdk; no manifest provides or requires amqp; the old broker's module and the two modules that required it leave the catalogue; the AMQP transport is deleted once every node reports on the new bus. Design 28 step 5 rewritten under it: the seat handover becomes its own task and is built first, because the seat the control plane dereferences cannot be empty in between — that emptiness was the outage. The cost note now carries what was measured rather than what was assumed. 0128 and 0130 extended 0127; each now rests on 0131 with a dated note and changes nothing it decided. Every other citation of 0127 names its replacement. records.py still fails on 0120/0112, which predates this branch.
54 KiB
layer, status, code, updated, decisions
| layer | status | code | updated | decisions | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| to-be | in-progress |
|
2026-09-27 |
|
28. Building the bus
The work of ADR 0116's five steps, in the order its dependencies allow, with what each ends at. Design 25 is the architecture and stays the authority on what is built; this document holds only the order, the sizes and the proofs, and it is wrong the moment it disagrees with design 25 rather than the other way round.
Each step ends at something runnable. A step that cannot name what its bed proves is not a step, and is divided further before it is started.
How this is built, and when it is run
Written as code with unit tests, committed per change, and taken to the lab once the pieces that would change the outcome are in place. The mistake this avoids is the one design 22 records: running a long bed against a mesh mid-transformation and debugging paths the next step deletes. Where a fault can be reasoned out of the code path, it is — reading, not running.
So the beds below are acceptance tests at the end of assembled work, not the tool for finding each bug, and a step's bed is run when that step is finished rather than while it is being written.
What the work is, measured
Counted 2026-09-26, non-test source only. The point of counting is that none of this is unknown territory: every piece has a shape already standing beside it.
| Piece | Today | Size | Becomes |
|---|---|---|---|
| the controller's link | Go, one package | ~1 800 lines | the same package on NATS |
| the host's link | Go, one package, mirroring the contracts rather than importing them | ~1 000 lines | the same, on NATS |
| the tool runtime's client | TypeScript, one file | ~390 lines | the same, on NATS |
| the sdk's messaging surface | TypeScript: messaging, events, tools, contracts, primitives | ~360 lines across five | unchanged, see below |
| the broker module | the adopted AMQP broker: client, tools, provisioner, bootstrap, image, manifest | ~340 lines of module code | the nats module, same shape |
| the beds | 39 lab scenarios, including a broker bed, an adoption bed, a genesis bed and a store-window bed | — | four analogues and one new |
Two measurements are worth stating on their own, because they change what the steps are.
The sdk speaks no AMQP, and never did. The word appears in its source three times, in three comments; its messaging module says in as many words that it "carries the contract, not a specific AMQP client build." ADR 0039 put the client in the runtime, and the payoff is collected here: no module is rebuilt for this change, and the sdk's own diff is three comments. That is the whole reason a bus can be replaced under a live mesh at all.
The wire therefore has three implementations, not two, and no suite pins any of them. ADR 0074 spoke of "the existing two implementations" — Go and TypeScript. Measured, the Go side is two separate packages that mirror rather than share (the host imports nothing, by ADR 0005), so the count is the controller's link, the host's link, and the runtime's client. And a search for conformance fixtures finds none anywhere in the four repositories: design 22's Phase 1.2 — the suite — has not been built.
This corrected the record. ADR 0116 said step 3's fixtures were recaptured on NATS. There is nothing to recapture, so step 3 builds the suite, and its first job is to pin the wire the mesh has before changing it — a suite written only against the new bus certifies whatever the new bus happens to do. A fact went stale while the decision stood, which is a progressive insight (
02-DECISIONS/README.md): it is marked and dated in ADR 0116 itself rather than left to be discovered here.
The order the work actually allows
The five steps are chunks of capability; the build order is not simply 1 to 5, and pretending otherwise would put two beds where they cannot run. Three edges decide it:
- A specification precedes the implementations it governs. ADR 0074's whole argument is that agreement is specified and checked, not hoped for. So the wire's NATS binding is written before the three implementations are, even though it is step 3 — and its conformance half can only finish once two implementations exist to disagree.
- A mesh cannot be raised on a bus nothing speaks. A bed that raises a mesh on NATS from genesis — enrolling a node, holding a push while the store restarts, rolling out an upgrade — needs the controller and the host to speak NATS already. That is the implementations, and they arrive with step 3.
- Adoption needs the module and nothing else. Step 2 puts a correctly configured server into a running mesh that continues to ignore it, which depends on no link at all.
So step 1's bed proves the server, from genesis, configured — not a mesh living on it. The full genesis bed is step 4's, where it can first run.
This corrected the record too. ADR 0116 first attributed "a mesh raised on NATS from genesis" to step 1. That bed cannot run until the links exist, and a step whose proof cannot run is the exact failure the record was written to prevent — so step 1 now ends at the server standing, correctly configured and carrying nothing, and the full bed is named under step 4. The five steps, their names, their order and the single rollout are unchanged; only where two beds run has moved. Marked and dated in ADR 0116 as a progressive insight, with what the record said before.
step 1 module, genesis places it ──┐
step 2 adoption into a running mesh ──┤ neither needs a link
│
step 3 the wire specified ──► three implementations ──► the suite
│
step 4 the flows, and the full genesis bed
│
step 5 the rollout
Step 1 — the module, and genesis raises it
Revised 2026-09-26 (ADR 0126, design 29). Tasks 1.3 and 1.4 said the controller composes every account and creates the four streams at genesis, from a fixed set. That is only the mesh's own half. A module declares seats with their protocols, so streams are created at registration and durable consumers at assignment — neither of which has happened at genesis. The fixed foundation set stays here; the derived machinery moves to step 3, where the declaration model it reads from is specified. Tasks 1.1 and 1.2, already done, are untouched by this: the module and its reload mechanism do not care what the configuration says.
Why here. Everything else needs a server to talk to, and genesis is where the foundation is defined. The mesh this is for will never travel this path — it is already running, and takes step 2 — but genesis is the definition every other path is measured against, and one that exists only on paper is wrong until there is a second mesh to find out.
-
1.1 the
natsmodule: manifest, image, one container, its client, TLS and monitoring ports, JetStream on a named volume — the shape of design 25 §5, and the same shape the broker module beside it already has -
1.2 the composed configuration as a directory resource, and the entrypoint that watches the one file and signals the server itself — design 25 §5's correction, kept inside the module because a container has no reload and a recreate would drop every connection the mesh has
-
1.3 the controller composes that file: accounts, permissions, TLS, JetStream — a user's permissions derived from its declaration and nothing else, over the three namespaces of design 29 §2, plus its own ack subject and its own inbox prefix (design 25 §4)
-
1.4 the mesh's own streams, created at genesis and asserted idempotently on start, by the controller as their only writer — the mesh's own, not all of them: a seat's streams are created when the module declaring it is registered, and a module's durable consumers when it is assigned, so this task is the fixed foundation set and 3.x carries the derived rest
-
1.5 genesis raises it as foundation, claiming the seat
mesh-broker— the seat is the server's role, not the product. Already true of the controller and needed no change: it resolves the broker by seat ("that is where the broker is, whatever else the topology says") and names no broker module anywhere in its source. What remains is namingnatsinstead of the deprecated broker where a genesis module set is declared, which is scenario and installer configuration — carried with 1.6 rather than before it. -
1.7 the composition, delivered — the controller gathering its principals, composing the file, and asserting the streams and consumers on start.
**In**: the user list is derived from the mesh's records and the credentials are kept. A bus user's bcrypt hash is now recorded, keyed by the username the file needs, and the plaintext is returned exactly once. That state is new and the reason is worth stating: on the bus the mesh runs on today an account is a management call — mint, hand over, seal to the holder, keep nothing — and that works because the broker remembers. Here the users are one file rewritten whenever any of it changes, so keeping nothing would mean **the first person's access change silently blanking every module's password**. **Permissions are not kept, only credentials.** Authority is derived from what each module declares every time the file is written ([ADR 0043](../../02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md)); a stored permission list would be a second account of a user's authority, able to disagree with the records it came from while both looked internally consistent. The derivation refuses two things where they can still be named: two users with one name — the server reads the file as one of them and which one depends on the order — and a module assigned but absent from the catalogue, which would compose a user with no authority and fail on its first publish with an authorisation error that says nothing about a missing manifest. A user the mesh has minted no password for is *named* rather than dropped or written as a user anybody is: an ordinary situation with an obvious remedy, and the caller decides whether a partial file is worth writing. A seat's protocol is gathered across the whole catalogue, not from one manifest, because a seat is declared by one module and held by another. **Out, and what each needs.** **Delivery is in, and it settled what a module declares.** The mesh writes the *accounts* and the module owns its *server*. The alternative was a manifest field enumerating ports, TLS paths and a store directory so the controller could write a whole configuration — wrong, because those are properties of the container the module raises and the controller would have to be kept in step with a Dockerfile it never sees. So a module declares its own configuration as a file resource and `bus-users` names where the mesh's half goes beside it; **asking is not enough to receive it**, because that file holds every user's password hash, so the claim on `mesh-broker` is what authorises it. Two things a running server changed. **An absolute include path is resolved relative to the including file's directory** — `include /etc/nats/accounts.conf` from another directory makes the server look for it *under* that directory and refuse to start — so both files share one. And **`verify: true` was refusing every connection in the mesh**: it makes the server demand a *client* certificate, and nothing in the mesh presents one — a host pins this server's exact certificate and authenticates with the password the mesh minted ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md), design 25 §4). Every connection would have died at the TLS handshake before any password was looked at, with an error that reads as a fault in the client. Removed; TLS is still required, because the block is what requires it and `verify` only decides whether client certificates are checked. **Design 25 §4 should say this**, and says nothing about it today. Also collected here: task 1.2's payoff, end to end against the module's own image — the user list rewritten, the module noticing and reloading the server itself with no signal from outside, and the connection the mesh already had still working afterwards. **Minting is in, on both halves.** A node at enrolment, a module when its credential is issued. Three things differ from a management call and each is the point of the move: the credential is minted into the mesh's records and becomes usable at the next composition, so no server need be reachable for it; the password travels beside the address rather than inside it, because a credential embedded in a URL leaks into every log line that prints a connection; and a module's durable consumer is derived from what it declared rather than named, so it cannot ask for delivery of something it did not say it consumes. A node reconnecting may be refused until the composition reaches the machine running the bus — which is what the host's reconnect backoff is for, where waiting for the push would hold an enrolment open for as long as a declaration takes to apply. **Which bus is one fact, and being told about both is refused at start.** Not warned about: a mesh half on each is one where a declaration goes out on one bus and the report comes back on the other, and every component logs success while it happens — ADR 0074's failure arriving through configuration instead of through code. A node that came away holding a credential for each could be half-moved, and nothing would say which half. **The objects are asserted on every start**, not created once at genesis: a stream somebody deleted, a mesh raised from a restored backup, or a bus whose data directory was replaced all have records and no objects, and a node whose consumer is missing hears nothing while everything else about it looks correct. Against a real server: every object accepted, asserting twice changes nothing (a start that failed the second time is a controller that cannot restart), a machine joining an already-raised bus accepted, each node's consumer bound to its own declaration subject and no other's, and CONTROL not dead-lettering — because the store window's bound is the controller's, and a server that gave up first would discard the push the stream exists to protect. **People are not in the list**, deliberately: the account model is built and `operator issue` is not (4.4), so there is nobody to derive. Left empty rather than guessed at. > **This corrects a tick, not a decision.** Tasks 1.3 and 1.4 are ticked and they are honest > about what they built — the composer, the derivation, the permission model, the stream and > consumer definitions, the asserter, all pure and held by unit tests and a golden > composition. What nobody wrote is the *caller*. Measured on the feature branch: outside the > package that defines them, there is **not one** use of the composer, the permission > derivation, the stream set, the stream asserter or the principal type. Step 1's "done when" > claims "every account and permission composed from the manifests", and a mesh raised today > would stand up a server with no user list at all. > > It also needs state the mesh does not keep. Design 25 §4 says the file holds bcrypt > hashes, and passwords are "minted and sealed exactly as today" — but today the mesh mints a > password, hands it to the broker through a management call, seals the plaintext to the > holder and **keeps nothing**. There is no management call here, so the hash has to survive > for every later recomposition: the first thing a person's access change or a new module > touches is a file that must still contain every other user's password. No bcrypt hash is > stored anywhere in the controller today. > > Named as its own task rather than folded into 1.3 so the gap is visible: the parts of > step 1 exist and the mesh does not yet do any of it. -
1.6 the genesis-broker bed — deferred: beds are run once, at the end, rather than per step (novox/hq design 22's rule, and the operator's instruction). Every claim step 1 makes is covered by a unit test or was demonstrated against the real server; what the bed adds is the claims that need a mesh.
Not done here, deliberately. The controller builds a module's broker credential as an
amqps://URL and defaults a portless genesis address to 5671. Those are correct until the rollout and must not move: steps 1 to 4 leave every node on AMQP (ADR 0116), so changing the credential's shape now would break the running bus to serve a bus nothing speaks yet. They change with the links, in step 3.
Done when. A mesh raised from nothing has the server standing with the streams asserted and
every account and permission composed from the manifests; a user cannot publish outside its
emits, subscribe outside its consumes, ack another user's delivery or subscribe another's inbox
prefix; the monitoring port is refused from anything but the private network; a change to the
composed file is live within one watcher interval without a restart, and the container is not
recreated by it.
The permission checks belong here rather than later because the server enforces them itself — a plain client proves them, no link required — and they are the whole of what ADR 0043 asks for. No mesh traffic is on the bus yet; that is step 4's bed, not this one's.
Step 2 — adoption puts it in the seat
Why here. It needs only step 1's module, it is the path the mesh that exists will actually take, and it is what makes steps 3 and 4 safe to develop against a live mesh. A running mesh does not get a foundation module by being raised again; it adopts one in place (ADR 0100).
- 2.1 the server raised beside the existing broker on its own ports, carrying nothing — installer-side, from the upstream image
- 2.2 the
natsmodule assigned, which recreates the container once, deliberately (see below), keeping its JetStream directory - 2.3 the seat claim, and the resolver's refusal of a second holder mesh-wide — already true and now proved: the refusal is generic to any mesh-scoped seat, and three tests pin what matters for this one — a second bus anywhere is refused naming the seat, a different bus implementation is refused for the same reason (which is what lets the bus be replaced at all), and the deprecated broker no longer contends for it, so both run on one mesh
- 2.4 the adoption bed — deferred with the other beds
Adoption here is not a no-op, and pretending it would be is the trap. The host keeps an existing container only when its spec matches the declaration exactly (
apply.go: existed && before.Spec == want && running → unchanged; anything else isrm -fand recreate). Genesis raises the server from the upstream image, because nothing has been built yet; the module declares the mesh-built artifact, which carries the entrypoint that reloads configuration in place. Those two specs differ, so assigning the module recreates the container.That is correct, and it is ADR 0067's pivot exactly: raise a temporary thing, then reinstall it as an ordinary module. It is safe only because it happens while the bus carries nothing — which is what 2.1 means by "carrying nothing", and why step 2 comes before anything speaks NATS rather than after. One recreate, at the one moment it costs nothing.
After that, never again. The configuration is a directory mount rather than a file, so rewriting accounts does not change the container's spec and the entrypoint reloads the server in place. That is the whole point of task 1.2, and this is the moment it pays: every later account, permission or person's access change touches a running bus with connections on it.
Done when. A mesh already running has the server adopted, holding mesh-broker; a second
assignment anywhere is refused at resolution — one per mesh; and every node is still on the old
bus with nothing routed to the new one. That last check is the point of the step: adoption that
quietly carried traffic would be step 5 arriving early and unrehearsed.
Step 3 — the protocol on NATS
Why here. The implementations cannot be written against an unwritten wire, and this is the step that decides what "agreeing" means for everything after it. It is the largest step and the one that pays for itself furthest away.
-
3.1/3.3 the fixtures — one directory in the sdk, read by each implementation's own runner rather than copied into either, because a fixture copied twice is two fixtures. The Go emitter and the runtime's NATS client both pass the first: every required header set, each value in the pinned shape, the subject derived the same way, and the payload the body alone.
**The suite also had to settle what "byte-for-byte" can mean**, which ADR 0074 stated and nothing had yet had to implement. The envelope is exact — subject, required headers, names and formats — because that is what two implementations get wrong invisibly. The body is not: Go sorts a map's keys and JavaScript keeps insertion order, so identical bytes would commit every implementation to a canonical JSON encoder, to buy a property the mesh never uses. Read strictly it would have sent somebody writing one. Still to capture: a served tool call, a grant and its answer, and the contributions file — the other three ADR 0074 names. -
3.2 design 19 rewritten from exchanges, queues and routing keys to the subjects and streams of design 25 §2–§3, per capability, with ADR 0074's model untouched: floor plus capabilities, an implementation legitimate when it claims less, identity from the sealed credential, dedup on
x-event-id. Claims checked against a running server are marked verified in the text, so a reader can tell what was measured from what was reasoned. One limitation lifts with the transport: a module may now call another's tool, which issue 049 recorded it could not. -
3.4 the controller's link on NATS — both halves are through the seam, and the store window is the server's.
Busstates the outbound in the mesh's words (publish an event, declare to a node) andControlstates the inbound (took it, dropped it, held it for the store); each has an AMQP and a NATS implementation, and both ship, because steps 1 to 4 leave every node on AMQP and both shipping is what holds them to one envelope.The outbound seam turned out to be eight call sites; the inbound was the larger half, and the reason: every handler took the transport's own delivery type, so the loop could not move without moving enrolment, reports, builds, upgrades and catch-up with it in one breath. **The window (ADR 0083) is now what decides, once, for both.** On the bus the mesh has, holding a message means an unacknowledged delivery kept in the controller, bounded by the prefetch and lost if it stops. On the bus being built it is a `nak` with a delay: the message stays the server's and the controller keeps only the moment it first could not take it, so one that restarts mid-window has nothing to lose. Seven claims about that were asked of a running server rather than reasoned — a report heard and gone from the work queue, one held through a store outage and recorded when it returned, one let go once the bound passed, a superseded one settled without being acted on, a heartbeat heard and nothing persisted, both followed events acknowledged on a stream the controller had no ack subject for, and the enrolment answer arriving at the address the request carried in its payload. **Three things the wiring forced into the open.** *Supersession is asked before the store, not after.* A report about a declaration the mesh has moved past would otherwise wait out a restarting store to be written, and then overwrite what the node is doing now. *Half of a report is not about a declaration, and that half is never stale.* What the machine **is** — the tunnel it took over, the ports its own bundle holds, what an adopted node found, a node moving its overlay key — reaches the mesh on a report and nowhere else. A rekey set aside as stale is a node whose overlay key never moves, and no retry is coming, because the node said it once. So staleness is asked only of a report that is purely an apply's account. *The controller could not have consumed a module event at all.* Its account granted no event subject to subscribe and no ack subject on the events stream, so every announcement would have been redelivered for ever, refused by the permission list it already had. Both are now granted, each subject named rather than by pattern — a controller subscribing every event in the mesh is a permission list that has stopped saying what it is for. Its consumers are **named beside the mesh's own streams rather than derived**, because the controller files no manifest and authority cannot come from a declaration that does not exist. Still outstanding: a build's own shape, which travels with the builder in step 4. -
3.5 the host's link on NATS — all three halves are through seams, mirroring the controller's and still importing nothing of the mesh's own (ADR 0005): the host's own interfaces over its own libraries, agreeing with the controller only because a fixture holds both to one envelope. A report goes through JetStream because it is the message the store-window guarantee is about; a heartbeat stays on core, because a heartbeat in a stream is the mesh's least valuable message competing for retention with its most valuable.
`Link` is dialling, hearing and saying in one interface, because **dialling is where the transport is chosen** and choosing it twice is how one half of a node ends up on a different bus from the other. `Asking` is the enrolment conversation, and it is separate for the opposite reason: almost nothing about it is the same, and a node that fails there is not in the mesh at all. **What the new bus took away, and what it would not give.** A host declares nothing here: on the old bus it declares its own queue, because a queue that is not there means a node that hears nothing, but the object it reads through now is a durable consumer and a host's account reaches no part of the JetStream API. So it **binds** to one the mesh made, and a missing one is said as the mesh's to answer rather than quietly created with whatever the client defaults to. Two things that had to be built for that: a node's declaration consumer (named after the node, because its ack grant is derived from the node's name, so any other name is a delivery it cannot acknowledge), and the enrolment user's **inbox** — design 25 §6 names it and the composer granted none, so an enrolling node would have published its request and waited out its timeout against a mesh that answered. **The reply address travels in the payload, and that is now proved from both ends.** The controller reads it from there (3.4) and the host writes it there and waits on it, and the test asserts the transport's own reply field held the *consumer's ack address* by the time the request arrived — so a future server that stopped claiming that field fails a test rather than letting the reason quietly become folklore. The host's **"newest wins" window narrows at the rollout rather than disappearing**, and that is now measured rather than predicted: three declarations pushed to an absent node leave one on the stream and it is the newest, so the catch-up half is the stream's — but three pushes to a connected node are still three deliveries, which is the half that stays. **The pin turned out easier here than in the tool runtime, not harder.** The Go client takes a `*tls.Config`, so the same pinned configuration with the same verify callback does the work; the subject-alternative-name constraint recorded under 3.6 is that client's, because it takes PEM strings with no verify hook. A host checks the fingerprint and nothing else. Nothing here composes an enrolment user per live token, and that is **1.7's**, not this task's: it is one input to a composition that does not happen at all yet. -
3.6 the tool runtime's client on NATS, behind the unchanged sdk contract — round-tripped against a real server: a tool answered across two connections, a throwing handler reaching the caller as an error rather than a timeout, an event delivered once with its key, body, node and event id intact. Ships beside the AMQP client and is selected at the rollout, because steps 1 to 4 leave every node on AMQP.
**A constraint it surfaced, recorded where somebody issuing a certificate will look.** The AMQP client pinned the exact certificate and switched hostname verification off, which is sound because a fingerprint is stronger than a name. The NATS client exposes no equivalent hook — its TLS options are PEM strings with no verify callback — so the pin still happens before dialling and the library's own name check happens beside it. **The bus's certificate must carry a subject-alternative name matching the address nodes dial it by**, or the connection is refused by a library error rather than by anything the mesh says. -
3.7 the sdk's three stale comments, and nothing else in it — three lines, which is the whole of the sdk's diff for the bus change, and the measurement that predicted it
-
3.8 the declaration model of design 29: local names derived to subjects, the three namespaces, permissions computed from a declaration, and a manifest that contains no subject. Done in the controller's composer (permissions, streams, consumers), in the runtime's client (subjects derived from the credential, never named by a module), and as a catalogue test asserting all 72 manifests hold no subject — because the rule held by construction, and a rule held by construction is one a later field breaks quietly.
-
3.9 seats declared by modules — the manifest now carries
seats(name, scope, accepts/emits/serves, retention) anduses, and registration refuses amesh-*name, a duplicate declarer, an undeclaredusesor claim, a seat with no protocol, a scope mismatch, and a holder that does not answer what its seat promises. Still to do: creating a seat's streams at registration and its holder's work-queue consumer at assignment, which need the JetStream client wired in.The refusal for an unknown claim *moved* rather than disappeared — the parser cannot judge it from one manifest any more, because another module may legitimately declare that seat, so it is registration's. The test that encoded the old rule was rewritten rather than deleted, and a second one pins the case the parser could not distinguish. **Done**: a seat's work queue is derived and created, and a holder's worker with it. The JetStream client behind them is wired and verified against a running server, which also completes 1.4's missing half — the pure `Asserter` had no implementation until now. -
3.10 the ten seat renames — done in the controller's table, the ten manifests that claim them, the controller's own shipped manifests, and every test. Not a migration after all: a holding is derived at resolution, never stored, so nothing recorded points at an old name (recorded as a progressive insight on ADR 0126). A kept rename table tells a manifest written against an old name what it became, because a module lives in its own repository and may be registered long after the catalogue stopped using one.
**A seat and the interface it delivers are different names.** The `git` seat became `mesh-git` while the `git` *provision* it delivers did not change, and the same for the package registry. A blanket replace got this wrong first and the failure read "the package registry is served on `<nil>`", which does not say "you renamed an interface" — so a test now pins every seat against the interface it delivers.
Done when. The fixtures are produced and consumed byte for byte by every implementation that claims the capability, and a module built before any of this serves its tools unchanged on the new runtime. The step is not done when the code runs — two implementations that disagree about an envelope do not fail to compile, they ignore each other while both keep running, which is the failure ADR 0074 exists to catch.
And the shared library gained nothing but the binding. A new transport is when the pressure to add conveniences is highest, and ADR 0039's rule does not bend for it: a helper that arrives with the bus is a review failure, not a detail. Code shared among a module's own features stays in that module.
Step 4 — the core speaks it
Why here. The links exist from step 3, so the flows that are not on the bus at all can move onto it, and the beds that need a mesh living on NATS can finally run.
A blocker surfaced here that is not this step's to fix. Every event name in the catalogue is still written the way a routing key on the bus the mesh has is written, so the derivation design 29 §1 specifies turns a consumer's declaration into a subject no emitter publishes — thirty-seven manifests, and one that cannot be composed at all. Nothing fails on the bus the mesh runs on today, where a routing key is matched literally; it fails on the first mesh raised on the new bus and not before, which is why wiring the controller's own subscription is what found it. Opened as issue 127. It holds 4.2, 4.3 and the catch-up half of 4.5; the node-facing flows — enrolment, reports, heartbeats, a build's outcome — are unaffected, because those subjects are the mesh's own and derive from nothing a module declares.
-
4.1 the full genesis bed — a mesh raised on NATS from nothing and living on it: a node enrols over TLS with a claimed token and the enrolment user cannot read a declaration; a push is held while the store restarts and applies after, nothing lost or duplicated; a node that was away gets exactly the newest declaration and refuses a replayed older one by sequence; an upgrade rolls out to two nodes; an event dead-letters after
max-deliver; and an enrolment held by anak-with-delay cycle still reaches the enrolling node, proving the reply travels in the payload and not the transport field the consumer's ack has claimed. The server-enforced permissions were proved at step 1 and are not re-proved here — nothing is outstanding but the bed itself. Both links speak NATS, the composition happens, and every claim above has a unit test or a check against a running server behind it. What none of them can stand in for is a mesh raising itself, which is what this bed is — so this is where the code stops and the lab starts -
4.2 a build source's change reaches the builder over the bus, and the build that follows is the one the change asked for — a build is work submitted to a role now (ADR 0129). Both sides are behind a seam with an implementation per bus, and on the bus being built one publish does what two did: the outcome is the role's own event, so the asker matches it by the id its request carried, the controller records it and the catalogue places it in the graph. A build machine needs a reply queue for nothing and a grant over nobody's inbox.
Checked against a running server: the round trip; a third party on the role's event hearing the same outcome the asker did, which is what the decision rests on; work leaving the queue once settled, so no second machine repeats it; work submitted with no machine holding the role **waiting rather than failing**, and being done when one arrives; and work a machine handed back coming round again. The outcome carries the module name, because only the manifest says what was built and one message now has three readers. A failed build names none: it produced no module version, and the catalogue would otherwise place something that was never made. -
[~] 4.3 an installation completes over the bus, with the same outcome as the path it replaces — the installer can raise it: a foundation template that stands up the server, writes the server's own settings and the mesh's first user list beside them, and starts a controller reaching the new bus. What remains is running it, which is 4.1's bed.
**The mesh composes its own user list, and at genesis there is no mesh to compose one.** So the installer carries the first — the controller's account at a well-known bootstrap password, exactly as the store is reached at `postgres:bootstrap` and the old bus at `guest:guest`, and rotated with them. From the controller's first composition onward the file is the controller's. That surfaced a gap reading would not have found: the controller's own account exists before there is a controller to mint one, so nothing recorded a hash for it and its first composition would have left the writer out of the file it was writing — a bus nothing can connect to, produced by the thing connected to it. It records a hash of the credential it is using, and only when none is recorded, so a restart cannot put the bootstrap password back over a rotated one. **The carried list and the derived one are checked against each other**, because they are two statements of one fact and a mesh cannot be raised twice to find out they disagreed. A template granting less than the controller derives produces a mesh that comes up, connects, and is refused on its first act, with an authorisation error naming a subject rather than the template that forgot it. The check earned itself at once: the composer was granting a role's whole event branch *and* the one event it follows, and the wider grant wins — so only the submitting half of a role is granted now, and what comes back is named exactly. -
4.4 a person's client — the account and the program are both in.
**The account**: a person is not a module and holds no seat, so their authority is a list of tools (or `*` for an administrator) and nothing else. Held to four properties, each a way of being wrong that would not announce itself: nothing but tools, so a person cannot claim a module said something; no ack subject, because authority over a consumer that does not exist is authority nobody audits; no ability to answer, because a person who can answer a request is impersonating a module on a bus where anyone may serve a tool; and two people do not share an inbox. Issued, listed and revoked by command; stating what somebody may call replaces what was there, because a list that could only grow is a permission nobody can take back; and forgetting somebody takes their credential with them, or it is not a revocation. **The program**: two surfaces over one thing — a command line and an MCP server — both adapters over the same three calls, because a second way of reaching a tool is a second thing to keep correct. It uses the client a module's runtime uses, so what a person may do is answered by the same permission list that answers it for a module and an audit has nothing separate to read. Three decisions in it worth keeping. It lists what the **catalogue** has rather than what this credential may call: somebody seeing only their own tools cannot tell "not installed" from "not yours", and those need different people to fix them. A failed call says which of three things happened — nobody serves it, this credential may not, or the tool was slow — because the remedies are in three different places and without that they are one timeout and a stack trace. And the MCP surface decides nothing: the names are the ones a person types, the schemas are the modules' own, an answer is passed through unshaped, and a tool that fails comes back as a tool error rather than a protocol error, because the request was well-formed and the mesh answered it. Both surfaces are driven against a running bus, including a host's notification being answered with nothing and an unknown method refused. > **Design 25 §7 says "nothing is built of this before §10's bed passes", and this was built > before.** Recorded rather than quietly ignored: the operator asked for it, it is on the > critical path for nothing and blocked by nothing, and the bed it waits for is 4.1's. If the > bed changes what a person's client should be, this is what gets changed. -
4.5 reports and catch-up: a node that was unreachable catches up rather than losing them.
**The reports half is in and proved against a server** (3.4): held through the store's absence by the server rather than by the controller, superseded ones settled by the digest they carry. **The catch-up half needed nothing built, and that was the answer.** It existed because a queue on the bus the mesh runs on today receives only what is published after it is bound, so everything built before the catalogue existed was announced to nobody — and on a fresh mesh that is always the foundation, because those are the things the catalogue needed in order to exist ([issue 050](../../04-ISSUES/050-the-catalogue-knows-nothing-built-before-it/00-report.md)). A whole mechanism followed: the catalogue asks, the controller re-publishes. A stream is a log and a consumer is a position in it. A consumer created later starts at the beginning, so the builds are simply there — asked of a running server rather than assumed, since the decision rested on it: three builds published with nothing listening, then a consumer created, and all three waiting for it. So the question of *who replays* has no answer because nothing replays. > **This is the shape of the whole change, in one task.** Three ways to do the replay were weighed > — a namespace for the mesh's own voice, the controller answering a question, a consumer reading > from the start — and the right answer was that the bus being moved to already does it. The > mechanism was never about builds; it was about a queue that could not remember. **A conversion > that carried it across would have carried a workaround for a limitation that no longer exists**, > and nothing would have looked wrong. Retiring it is step 5's, with the rest of what only the old bus needs: the request, the re-publishing, and the `replay` flag that told a consumer to register history without acting on it.
Done when. Each converted flow is proved against the behaviour it replaced, and the full genesis bed is green. Observation is not in this step — heartbeats, conditions and key-value state are research 017's, that effort already reserves them for after the move, and a flow built ahead of its design would be rebuilt.
Step 5 — the rollout
Why here. It is the only step that moves a node's bus, and it moves every node's at once.
The order the repositories land in is part of the rollout, not paperwork. Derived 2026-09-27 while merging, and not obvious from any one repository, which is why it is written here rather than left to be re-derived under time pressure:
| order | repository | why it cannot be later |
|---|---|---|
| 1 | this one | prose; nothing deploys |
| 2 | the sdk | comments only, and no module rebuilds for it |
| 3 | the client library | it is what a module calls to emit, and it is where the subject is derived. Until it lands, a locally-named event is published under the local name itself |
| 4 | the catalogue | every manifest and every module's code, renamed together. Safe only once the runtime derives |
| 5 | the controller | it refuses an old-style event name outright, so landing it before the catalogue makes every unconverted module unregisterable |
| 6 | the hosts | last, because nothing else waits on them |
Two properties make the sequence safe rather than merely ordered, and both are pinned by tests. A name already in the old form passes through the derivation untouched, so a module nobody has converted keeps working at every step. And a converted name derives to exactly the key the old bus published, so steps 3 and 4 change nothing on the wire — the move to the new bus is step 5.2 and one environment variable, not a side effect of deploying.
The failure this ordering avoids is issue 127's own: a publisher and a subscriber that disagree about a subject produce no error anywhere. Nothing logs, nothing retries, and the mesh reports itself healthy while reacting to nothing.
-
5.1 the cutover bed: a mesh on AMQP with a predecessor stand-in on the deprecated broker moves its bus in one rollout, every node reporting on NATS afterwards, the stand-in's own client still connected throughout
-
[~] 5.2 the rollout: accounts composed, then the controller, every host and every runtime together; every node confirmed heard before AMQP stops.
**The readiness half is in and is the half worth having.** The move takes every node at once, so there is nothing to inspect afterwards and no half to roll back — either the mesh was ready or it was not. `rollout check` answers that from records, with one dial: is a bus answering, does a machine hold the seat, has it been sent the composed user list, does every machine and every module that speaks have a credential. Each missing thing names its own next step, because "not ready" that cannot be acted on is not an answer at the point where the next step is irreversible. **A machine with no credential is what must stop it.** It keeps running, cannot come back, and afterwards there is no bus to tell it anything over. The move itself is deliberately not written yet, and the command says so rather than pretending: it waits on the check having been run against a real mesh. Writing the irreversible half before the question it depends on has ever been asked of something real is how the plan's own rule about beds gets broken by another route. > **What this costs if it goes wrong, measured rather than assumed.** Nothing in a served > request's path goes over the mesh's own bus: modules serve from their own containers. What a > failed move costs is the mesh's ability to *change* anything — pushes, tool calls, new > provisioning — until it is finished or undone. That is worth knowing before rather than > after, and it is why the operator's "as long as my services keep running" is a reasonable > position rather than a gamble. **Measured on 2026-09-27**, when a seat emptied itself > mid-change: 52 containers stayed up and the broker never stopped; the control plane > crash-looped for two hours and nothing could be deployed until it was repaired by hand. > An earlier version of this note said the old broker stays as an ordinary provider of > `amqp` ([ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md)); that is withdrawn by [ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md) — see 5.4. -
5.3 the seat changes hands as one act. A command takes a seat and the assignment taking it over, and the seat is never empty in between — the emptiness is the outage of 2026-09-27, when the control plane, which finds its own bus through this seat, lost the address and looped. Today only
seat renameexists. This is what 5.2 uses to movemesh-brokerfrom the old broker's assignment to the new one's, and it is built first (ADR 0131). -
5.4 the old broker and everything that named AMQP leave the mesh (ADR 0131, superseding ADR 0127): the two modules that required
amqpare removed, the broker's module is unassigned and removed, registration refuses a manifest that provides or requiresamqp, and a whole-catalogue check asserts none does. Not a retirement condition — a decision, taken, with the operator's "I don't care if the predecessor breaks" on record (ADR 0130). Retiring with it: the build outcome's second announcement under the module's own name, which existed only so a catalogue deployed before the rename and one after both heard it.> **The remote tooling goes with it too.** The predecessor's own mesh talks over that broker, so > shutting it down ends the path that reaches this installation's machines from a workstation. > The rollout is driven from the node, or before the broker stops — a sequencing constraint on > 5.2, not an afterthought. -
5.5 the AMQP transport is deleted from the control plane and the hosts, and the variable that selected a transport is refused at start as unknown. One bus, nothing to select (ADR 0131).
The old 5.4 note is history. It recorded that a retirement condition was wrong from ADR 0127 onward, which framed the old broker as an ordinary provider with no end. ADR 0131 ends that framing in turn: the broker is not kept as a provider either, because AMQP is not a provision. Both readings are kept here so the two reversals can be read in order.
Done when. Every node reports on NATS, and nothing of the mesh's own is left connected to the deprecated broker.
The through-line
The order is dependency, not preference. Steps 1 to 4 leave every node on AMQP, so the cost of being wrong is bounded until the last step: a step may be abandoned, or reordered after step 2, without a rollback. The server stands before anything speaks to it; the wire is specified before it is implemented three times; the flows move once there is something to move them onto; and the bus itself moves once, at the end, on one day.
What is deliberately not here
- Observation — research 017's, after the move, by its own design.
- Leaf nodes — design 25 §11 keeps this out of scope and says so; a leaf per machine is a later question, noted so it is not forgotten.
- The predecessor's world. It is AMQP and it is not moving — ADR 0130: it is deprecated, some of it is still running, and it is being left to stop rather than migrated. Its broker goes with it, unassigned like any provider whose provision nothing requires.
How this list is kept true
A task is ticked when its change is committed, not when it is written. A step is done when its bed is green, not when its tasks are ticked. If a step's tasks are all ticked and its bed has not run, the step is in progress and this document says so — that gap is the thing the whole shape is built to make visible.