Files
hq/03-DESIGN/01-to-be/28-building-the-bus.md
jschoubben d7d6f2eda0 Design 28: the move needs a credential and a membership for machines already enrolled
The first live attempt at 5.2 found the half nobody had built. With the seat
handed over on record and the new bus's module assigned beside the old one, the
push was refused: not one user has a credential for the new bus. The check is
right. A credential is minted only at enrolment, at `module issue` and for a
person; nothing mints one for a machine already enrolled or for the control plane
itself, and on the host the membership is written once at enrolment and never
rewritten. So "move each machine" had no mechanism under it on either side.

Written into 5.2 as the mechanism to build before anything moves: the control
plane mints what is missing and delivers each plaintext where its owner reads it —
a machine's as a sealed membership in its declaration, a module's as its broker
secret, its own as its module secret — and the host saves a delivered membership
and re-dials on it through the reconnect path it already has. 5.3 is ticked as
built; 5.4's catalogue half is done and its live half waits on 5.2.
2026-09-27 23:47:45 +02:00

56 KiB

layer, status, code, updated, decisions
layer status code updated decisions
to-be in-progress
mesh-catalog modules/nats
mesh-controller internal/catalogue
mesh-lab scenarios
2026-09-27
02-DECISIONS/0116-the-bus-is-built-in-five-steps.md
02-DECISIONS/0106-the-bus-is-nats.md
02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
02-DECISIONS/0126-a-module-declares-its-own-seats.md
02-DECISIONS/0074-the-wire-is-specified-not-the-types.md
02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md
02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md
02-DECISIONS/0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md

28. Building the bus

The work of ADR 0116's five steps, in the order its dependencies allow, with what each ends at. Design 25 is the architecture and stays the authority on what is built; this document holds only the order, the sizes and the proofs, and it is wrong the moment it disagrees with design 25 rather than the other way round.

Each step ends at something runnable. A step that cannot name what its bed proves is not a step, and is divided further before it is started.

How this is built, and when it is run

Written as code with unit tests, committed per change, and taken to the lab once the pieces that would change the outcome are in place. The mistake this avoids is the one design 22 records: running a long bed against a mesh mid-transformation and debugging paths the next step deletes. Where a fault can be reasoned out of the code path, it is — reading, not running.

So the beds below are acceptance tests at the end of assembled work, not the tool for finding each bug, and a step's bed is run when that step is finished rather than while it is being written.

What the work is, measured

Counted 2026-09-26, non-test source only. The point of counting is that none of this is unknown territory: every piece has a shape already standing beside it.

Piece Today Size Becomes
the controller's link Go, one package ~1 800 lines the same package on NATS
the host's link Go, one package, mirroring the contracts rather than importing them ~1 000 lines the same, on NATS
the tool runtime's client TypeScript, one file ~390 lines the same, on NATS
the sdk's messaging surface TypeScript: messaging, events, tools, contracts, primitives ~360 lines across five unchanged, see below
the broker module the adopted AMQP broker: client, tools, provisioner, bootstrap, image, manifest ~340 lines of module code the nats module, same shape
the beds 39 lab scenarios, including a broker bed, an adoption bed, a genesis bed and a store-window bed — four analogues and one new

Two measurements are worth stating on their own, because they change what the steps are.

The sdk speaks no AMQP, and never did. The word appears in its source three times, in three comments; its messaging module says in as many words that it "carries the contract, not a specific AMQP client build." ADR 0039 put the client in the runtime, and the payoff is collected here: no module is rebuilt for this change, and the sdk's own diff is three comments. That is the whole reason a bus can be replaced under a live mesh at all.

The wire therefore has three implementations, not two, and no suite pins any of them. ADR 0074 spoke of "the existing two implementations" — Go and TypeScript. Measured, the Go side is two separate packages that mirror rather than share (the host imports nothing, by ADR 0005), so the count is the controller's link, the host's link, and the runtime's client. And a search for conformance fixtures finds none anywhere in the four repositories: design 22's Phase 1.2 — the suite — has not been built.

This corrected the record. ADR 0116 said step 3's fixtures were recaptured on NATS. There is nothing to recapture, so step 3 builds the suite, and its first job is to pin the wire the mesh has before changing it — a suite written only against the new bus certifies whatever the new bus happens to do. A fact went stale while the decision stood, which is a progressive insight (02-DECISIONS/README.md): it is marked and dated in ADR 0116 itself rather than left to be discovered here.

The order the work actually allows

The five steps are chunks of capability; the build order is not simply 1 to 5, and pretending otherwise would put two beds where they cannot run. Three edges decide it:

  • A specification precedes the implementations it governs. ADR 0074's whole argument is that agreement is specified and checked, not hoped for. So the wire's NATS binding is written before the three implementations are, even though it is step 3 — and its conformance half can only finish once two implementations exist to disagree.
  • A mesh cannot be raised on a bus nothing speaks. A bed that raises a mesh on NATS from genesis — enrolling a node, holding a push while the store restarts, rolling out an upgrade — needs the controller and the host to speak NATS already. That is the implementations, and they arrive with step 3.
  • Adoption needs the module and nothing else. Step 2 puts a correctly configured server into a running mesh that continues to ignore it, which depends on no link at all.

So step 1's bed proves the server, from genesis, configured — not a mesh living on it. The full genesis bed is step 4's, where it can first run.

This corrected the record too. ADR 0116 first attributed "a mesh raised on NATS from genesis" to step 1. That bed cannot run until the links exist, and a step whose proof cannot run is the exact failure the record was written to prevent — so step 1 now ends at the server standing, correctly configured and carrying nothing, and the full bed is named under step 4. The five steps, their names, their order and the single rollout are unchanged; only where two beds run has moved. Marked and dated in ADR 0116 as a progressive insight, with what the record said before.

step 1  module, genesis places it        ──┐
step 2  adoption into a running mesh     ──┤  neither needs a link
                                           │
step 3  the wire specified ──► three implementations ──► the suite
                                           │
step 4  the flows, and the full genesis bed
                                           │
step 5  the rollout

Step 1 — the module, and genesis raises it

Revised 2026-09-26 (ADR 0126, design 29). Tasks 1.3 and 1.4 said the controller composes every account and creates the four streams at genesis, from a fixed set. That is only the mesh's own half. A module declares seats with their protocols, so streams are created at registration and durable consumers at assignment — neither of which has happened at genesis. The fixed foundation set stays here; the derived machinery moves to step 3, where the declaration model it reads from is specified. Tasks 1.1 and 1.2, already done, are untouched by this: the module and its reload mechanism do not care what the configuration says.

Why here. Everything else needs a server to talk to, and genesis is where the foundation is defined. The mesh this is for will never travel this path — it is already running, and takes step 2 — but genesis is the definition every other path is measured against, and one that exists only on paper is wrong until there is a second mesh to find out.

  • 1.1 the nats module: manifest, image, one container, its client, TLS and monitoring ports, JetStream on a named volume — the shape of design 25 §5, and the same shape the broker module beside it already has

  • 1.2 the composed configuration as a directory resource, and the entrypoint that watches the one file and signals the server itself — design 25 §5's correction, kept inside the module because a container has no reload and a recreate would drop every connection the mesh has

  • 1.3 the controller composes that file: accounts, permissions, TLS, JetStream — a user's permissions derived from its declaration and nothing else, over the three namespaces of design 29 §2, plus its own ack subject and its own inbox prefix (design 25 §4)

  • 1.4 the mesh's own streams, created at genesis and asserted idempotently on start, by the controller as their only writer — the mesh's own, not all of them: a seat's streams are created when the module declaring it is registered, and a module's durable consumers when it is assigned, so this task is the fixed foundation set and 3.x carries the derived rest

  • 1.5 genesis raises it as foundation, claiming the seat mesh-broker — the seat is the server's role, not the product. Already true of the controller and needed no change: it resolves the broker by seat ("that is where the broker is, whatever else the topology says") and names no broker module anywhere in its source. What remains is naming nats instead of the deprecated broker where a genesis module set is declared, which is scenario and installer configuration — carried with 1.6 rather than before it.

  • 1.7 the composition, delivered — the controller gathering its principals, composing the file, and asserting the streams and consumers on start.

    **In**: the user list is derived from the mesh's records and the credentials are kept.
    
    A bus user's bcrypt hash is now recorded, keyed by the username the file needs, and the
    plaintext is returned exactly once. That state is new and the reason is worth stating: on the
    bus the mesh runs on today an account is a management call — mint, hand over, seal to the
    holder, keep nothing — and that works because the broker remembers. Here the users are one
    file rewritten whenever any of it changes, so keeping nothing would mean **the first person's
    access change silently blanking every module's password**.
    
    **Permissions are not kept, only credentials.** Authority is derived from what each module
    declares every time the file is written ([ADR 0043](../../02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md));
    a stored permission list would be a second account of a user's authority, able to disagree
    with the records it came from while both looked internally consistent.
    
    The derivation refuses two things where they can still be named: two users with one name — the
    server reads the file as one of them and which one depends on the order — and a module
    assigned but absent from the catalogue, which would compose a user with no authority and fail
    on its first publish with an authorisation error that says nothing about a missing manifest.
    A user the mesh has minted no password for is *named* rather than dropped or written as a user
    anybody is: an ordinary situation with an obvious remedy, and the caller decides whether a
    partial file is worth writing. A seat's protocol is gathered across the whole catalogue, not
    from one manifest, because a seat is declared by one module and held by another.
    
    **Out, and what each needs.**
    
    **Delivery is in, and it settled what a module declares.** The mesh writes the *accounts* and
    the module owns its *server*. The alternative was a manifest field enumerating ports, TLS
    paths and a store directory so the controller could write a whole configuration — wrong,
    because those are properties of the container the module raises and the controller would have
    to be kept in step with a Dockerfile it never sees. So a module declares its own configuration
    as a file resource and `bus-users` names where the mesh's half goes beside it; **asking is not
    enough to receive it**, because that file holds every user's password hash, so the claim on
    `mesh-broker` is what authorises it.
    
    Two things a running server changed. **An absolute include path is resolved relative to the
    including file's directory** — `include /etc/nats/accounts.conf` from another directory makes
    the server look for it *under* that directory and refuse to start — so both files share one.
    And **`verify: true` was refusing every connection in the mesh**: it makes the server demand a
    *client* certificate, and nothing in the mesh presents one — a host pins this server's exact
    certificate and authenticates with the password the mesh minted ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md),
    design 25 §4). Every connection would have died at the TLS handshake before any password was
    looked at, with an error that reads as a fault in the client. Removed; TLS is still required,
    because the block is what requires it and `verify` only decides whether client certificates
    are checked. **Design 25 §4 should say this**, and says nothing about it today.
    
    Also collected here: task 1.2's payoff, end to end against the module's own image — the user
    list rewritten, the module noticing and reloading the server itself with no signal from
    outside, and the connection the mesh already had still working afterwards.
    
    **Minting is in, on both halves.** A node at enrolment, a module when its credential is
    issued. Three things differ from a management call and each is the point of the move: the
    credential is minted into the mesh's records and becomes usable at the next composition, so no
    server need be reachable for it; the password travels beside the address rather than inside it,
    because a credential embedded in a URL leaks into every log line that prints a connection; and
    a module's durable consumer is derived from what it declared rather than named, so it cannot ask
    for delivery of something it did not say it consumes. A node reconnecting may be refused until
    the composition reaches the machine running the bus — which is what the host's reconnect backoff
    is for, where waiting for the push would hold an enrolment open for as long as a declaration
    takes to apply.
    
    **Which bus is one fact, and being told about both is refused at start.** Not warned about: a
    mesh half on each is one where a declaration goes out on one bus and the report comes back on
    the other, and every component logs success while it happens — ADR 0074's failure arriving
    through configuration instead of through code. A node that came away holding a credential for
    each could be half-moved, and nothing would say which half.
    
    **The objects are asserted on every start**, not created once at genesis: a stream somebody
    deleted, a mesh raised from a restored backup, or a bus whose data directory was replaced all
    have records and no objects, and a node whose consumer is missing hears nothing while everything
    else about it looks correct. Against a real server: every object accepted, asserting twice
    changes nothing (a start that failed the second time is a controller that cannot restart), a
    machine joining an already-raised bus accepted, each node's consumer bound to its own
    declaration subject and no other's, and CONTROL not dead-lettering — because the store window's
    bound is the controller's, and a server that gave up first would discard the push the stream
    exists to protect.
    
    **People are not in the list**, deliberately: the account model is built and `operator issue`
    is not (4.4), so there is nobody to derive. Left empty rather than guessed at.
    
    > **This corrects a tick, not a decision.** Tasks 1.3 and 1.4 are ticked and they are honest
    > about what they built — the composer, the derivation, the permission model, the stream and
    > consumer definitions, the asserter, all pure and held by unit tests and a golden
    > composition. What nobody wrote is the *caller*. Measured on the feature branch: outside the
    > package that defines them, there is **not one** use of the composer, the permission
    > derivation, the stream set, the stream asserter or the principal type. Step 1's "done when"
    > claims "every account and permission composed from the manifests", and a mesh raised today
    > would stand up a server with no user list at all.
    >
    > It also needs state the mesh does not keep. Design 25 §4 says the file holds bcrypt
    > hashes, and passwords are "minted and sealed exactly as today" — but today the mesh mints a
    > password, hands it to the broker through a management call, seals the plaintext to the
    > holder and **keeps nothing**. There is no management call here, so the hash has to survive
    > for every later recomposition: the first thing a person's access change or a new module
    > touches is a file that must still contain every other user's password. No bcrypt hash is
    > stored anywhere in the controller today.
    >
    > Named as its own task rather than folded into 1.3 so the gap is visible: the parts of
    > step 1 exist and the mesh does not yet do any of it.
    
  • 1.6 the genesis-broker bed — deferred: beds are run once, at the end, rather than per step (novox/hq design 22's rule, and the operator's instruction). Every claim step 1 makes is covered by a unit test or was demonstrated against the real server; what the bed adds is the claims that need a mesh.

Not done here, deliberately. The controller builds a module's broker credential as an amqps:// URL and defaults a portless genesis address to 5671. Those are correct until the rollout and must not move: steps 1 to 4 leave every node on AMQP (ADR 0116), so changing the credential's shape now would break the running bus to serve a bus nothing speaks yet. They change with the links, in step 3.

Done when. A mesh raised from nothing has the server standing with the streams asserted and every account and permission composed from the manifests; a user cannot publish outside its emits, subscribe outside its consumes, ack another user's delivery or subscribe another's inbox prefix; the monitoring port is refused from anything but the private network; a change to the composed file is live within one watcher interval without a restart, and the container is not recreated by it.

The permission checks belong here rather than later because the server enforces them itself — a plain client proves them, no link required — and they are the whole of what ADR 0043 asks for. No mesh traffic is on the bus yet; that is step 4's bed, not this one's.

Step 2 — adoption puts it in the seat

Why here. It needs only step 1's module, it is the path the mesh that exists will actually take, and it is what makes steps 3 and 4 safe to develop against a live mesh. A running mesh does not get a foundation module by being raised again; it adopts one in place (ADR 0100).

  • 2.1 the server raised beside the existing broker on its own ports, carrying nothing — installer-side, from the upstream image
  • 2.2 the nats module assigned, which recreates the container once, deliberately (see below), keeping its JetStream directory
  • 2.3 the seat claim, and the resolver's refusal of a second holder mesh-wide — already true and now proved: the refusal is generic to any mesh-scoped seat, and three tests pin what matters for this one — a second bus anywhere is refused naming the seat, a different bus implementation is refused for the same reason (which is what lets the bus be replaced at all), and the deprecated broker no longer contends for it, so both run on one mesh
  • 2.4 the adoption bed — deferred with the other beds

Adoption here is not a no-op, and pretending it would be is the trap. The host keeps an existing container only when its spec matches the declaration exactly (apply.go: existed && before.Spec == want && running → unchanged; anything else is rm -f and recreate). Genesis raises the server from the upstream image, because nothing has been built yet; the module declares the mesh-built artifact, which carries the entrypoint that reloads configuration in place. Those two specs differ, so assigning the module recreates the container.

That is correct, and it is ADR 0067's pivot exactly: raise a temporary thing, then reinstall it as an ordinary module. It is safe only because it happens while the bus carries nothing — which is what 2.1 means by "carrying nothing", and why step 2 comes before anything speaks NATS rather than after. One recreate, at the one moment it costs nothing.

After that, never again. The configuration is a directory mount rather than a file, so rewriting accounts does not change the container's spec and the entrypoint reloads the server in place. That is the whole point of task 1.2, and this is the moment it pays: every later account, permission or person's access change touches a running bus with connections on it.

Done when. A mesh already running has the server adopted, holding mesh-broker; a second assignment anywhere is refused at resolution — one per mesh; and every node is still on the old bus with nothing routed to the new one. That last check is the point of the step: adoption that quietly carried traffic would be step 5 arriving early and unrehearsed.

Step 3 — the protocol on NATS

Why here. The implementations cannot be written against an unwritten wire, and this is the step that decides what "agreeing" means for everything after it. It is the largest step and the one that pays for itself furthest away.

  • 3.1/3.3 the fixtures — one directory in the sdk, read by each implementation's own runner rather than copied into either, because a fixture copied twice is two fixtures. The Go emitter and the runtime's NATS client both pass the first: every required header set, each value in the pinned shape, the subject derived the same way, and the payload the body alone.

    **The suite also had to settle what "byte-for-byte" can mean**, which ADR 0074 stated and
    nothing had yet had to implement. The envelope is exact — subject, required headers, names
    and formats — because that is what two implementations get wrong invisibly. The body is
    not: Go sorts a map's keys and JavaScript keeps insertion order, so identical bytes would
    commit every implementation to a canonical JSON encoder, to buy a property the mesh never
    uses. Read strictly it would have sent somebody writing one.
    
    Still to capture: a served tool call, a grant and its answer, and the contributions file —
    the other three ADR 0074 names.
    
  • 3.2 design 19 rewritten from exchanges, queues and routing keys to the subjects and streams of design 25 §2–§3, per capability, with ADR 0074's model untouched: floor plus capabilities, an implementation legitimate when it claims less, identity from the sealed credential, dedup on x-event-id. Claims checked against a running server are marked verified in the text, so a reader can tell what was measured from what was reasoned. One limitation lifts with the transport: a module may now call another's tool, which issue 049 recorded it could not.

  • 3.4 the controller's link on NATS — both halves are through the seam, and the store window is the server's. Bus states the outbound in the mesh's words (publish an event, declare to a node) and Control states the inbound (took it, dropped it, held it for the store); each has an AMQP and a NATS implementation, and both ship, because steps 1 to 4 leave every node on AMQP and both shipping is what holds them to one envelope.

    The outbound seam turned out to be eight call sites; the inbound was the larger half, and
    the reason: every handler took the transport's own delivery type, so the loop could not move
    without moving enrolment, reports, builds, upgrades and catch-up with it in one breath.
    
    **The window (ADR 0083) is now what decides, once, for both.** On the bus the mesh has,
    holding a message means an unacknowledged delivery kept in the controller, bounded by the
    prefetch and lost if it stops. On the bus being built it is a `nak` with a delay: the
    message stays the server's and the controller keeps only the moment it first could not take
    it, so one that restarts mid-window has nothing to lose. Seven claims about that were asked
    of a running server rather than reasoned — a report heard and gone from the work queue, one
    held through a store outage and recorded when it returned, one let go once the bound passed,
    a superseded one settled without being acted on, a heartbeat heard and nothing persisted,
    both followed events acknowledged on a stream the controller had no ack subject for, and the
    enrolment answer arriving at the address the request carried in its payload.
    
    **Three things the wiring forced into the open.**
    
    *Supersession is asked before the store, not after.* A report about a declaration the mesh
    has moved past would otherwise wait out a restarting store to be written, and then overwrite
    what the node is doing now.
    
    *Half of a report is not about a declaration, and that half is never stale.* What the machine
    **is** — the tunnel it took over, the ports its own bundle holds, what an adopted node found,
    a node moving its overlay key — reaches the mesh on a report and nowhere else. A rekey set
    aside as stale is a node whose overlay key never moves, and no retry is coming, because the
    node said it once. So staleness is asked only of a report that is purely an apply's account.
    
    *The controller could not have consumed a module event at all.* Its account granted no event
    subject to subscribe and no ack subject on the events stream, so every announcement would
    have been redelivered for ever, refused by the permission list it already had. Both are now
    granted, each subject named rather than by pattern — a controller subscribing every event in
    the mesh is a permission list that has stopped saying what it is for. Its consumers are
    **named beside the mesh's own streams rather than derived**, because the controller files no
    manifest and authority cannot come from a declaration that does not exist.
    
    Still outstanding: a build's own shape, which travels with the builder in step 4.
    
  • 3.5 the host's link on NATS — all three halves are through seams, mirroring the controller's and still importing nothing of the mesh's own (ADR 0005): the host's own interfaces over its own libraries, agreeing with the controller only because a fixture holds both to one envelope. A report goes through JetStream because it is the message the store-window guarantee is about; a heartbeat stays on core, because a heartbeat in a stream is the mesh's least valuable message competing for retention with its most valuable.

    `Link` is dialling, hearing and saying in one interface, because **dialling is where the
    transport is chosen** and choosing it twice is how one half of a node ends up on a different
    bus from the other. `Asking` is the enrolment conversation, and it is separate for the
    opposite reason: almost nothing about it is the same, and a node that fails there is not in
    the mesh at all.
    
    **What the new bus took away, and what it would not give.** A host declares nothing here: on
    the old bus it declares its own queue, because a queue that is not there means a node that
    hears nothing, but the object it reads through now is a durable consumer and a host's account
    reaches no part of the JetStream API. So it **binds** to one the mesh made, and a missing one
    is said as the mesh's to answer rather than quietly created with whatever the client defaults
    to. Two things that had to be built for that: a node's declaration consumer (named after the
    node, because its ack grant is derived from the node's name, so any other name is a delivery
    it cannot acknowledge), and the enrolment user's **inbox** — design 25 §6 names it and the
    composer granted none, so an enrolling node would have published its request and waited out
    its timeout against a mesh that answered.
    
    **The reply address travels in the payload, and that is now proved from both ends.** The
    controller reads it from there (3.4) and the host writes it there and waits on it, and the
    test asserts the transport's own reply field held the *consumer's ack address* by the time the
    request arrived — so a future server that stopped claiming that field fails a test rather than
    letting the reason quietly become folklore.
    
    The host's **"newest wins" window narrows at the rollout rather than disappearing**, and that
    is now measured rather than predicted: three declarations pushed to an absent node leave one
    on the stream and it is the newest, so the catch-up half is the stream's — but three pushes to
    a connected node are still three deliveries, which is the half that stays.
    
    **The pin turned out easier here than in the tool runtime, not harder.** The Go client takes a
    `*tls.Config`, so the same pinned configuration with the same verify callback does the work;
    the subject-alternative-name constraint recorded under 3.6 is that client's, because it takes
    PEM strings with no verify hook. A host checks the fingerprint and nothing else.
    
    Nothing here composes an enrolment user per live token, and that is **1.7's**, not this
    task's: it is one input to a composition that does not happen at all yet.
    
  • 3.6 the tool runtime's client on NATS, behind the unchanged sdk contract — round-tripped against a real server: a tool answered across two connections, a throwing handler reaching the caller as an error rather than a timeout, an event delivered once with its key, body, node and event id intact. Ships beside the AMQP client and is selected at the rollout, because steps 1 to 4 leave every node on AMQP.

    **A constraint it surfaced, recorded where somebody issuing a certificate will look.** The
    AMQP client pinned the exact certificate and switched hostname verification off, which is
    sound because a fingerprint is stronger than a name. The NATS client exposes no equivalent
    hook — its TLS options are PEM strings with no verify callback — so the pin still happens
    before dialling and the library's own name check happens beside it. **The bus's certificate
    must carry a subject-alternative name matching the address nodes dial it by**, or the
    connection is refused by a library error rather than by anything the mesh says.
    
  • 3.7 the sdk's three stale comments, and nothing else in it — three lines, which is the whole of the sdk's diff for the bus change, and the measurement that predicted it

  • 3.8 the declaration model of design 29: local names derived to subjects, the three namespaces, permissions computed from a declaration, and a manifest that contains no subject. Done in the controller's composer (permissions, streams, consumers), in the runtime's client (subjects derived from the credential, never named by a module), and as a catalogue test asserting all 72 manifests hold no subject — because the rule held by construction, and a rule held by construction is one a later field breaks quietly.

  • 3.9 seats declared by modules — the manifest now carries seats (name, scope, accepts/emits/serves, retention) and uses, and registration refuses a mesh-* name, a duplicate declarer, an undeclared uses or claim, a seat with no protocol, a scope mismatch, and a holder that does not answer what its seat promises. Still to do: creating a seat's streams at registration and its holder's work-queue consumer at assignment, which need the JetStream client wired in.

    The refusal for an unknown claim *moved* rather than disappeared — the parser cannot judge
    it from one manifest any more, because another module may legitimately declare that seat,
    so it is registration's. The test that encoded the old rule was rewritten rather than
    deleted, and a second one pins the case the parser could not distinguish.
    
    **Done**: a seat's work queue is derived and created, and a holder's worker with it. The
    JetStream client behind them is wired and verified against a running server, which also
    completes 1.4's missing half — the pure `Asserter` had no implementation until now.
    
  • 3.10 the ten seat renames — done in the controller's table, the ten manifests that claim them, the controller's own shipped manifests, and every test. Not a migration after all: a holding is derived at resolution, never stored, so nothing recorded points at an old name (recorded as a progressive insight on ADR 0126). A kept rename table tells a manifest written against an old name what it became, because a module lives in its own repository and may be registered long after the catalogue stopped using one.

    **A seat and the interface it delivers are different names.** The `git` seat became
    `mesh-git` while the `git` *provision* it delivers did not change, and the same for the
    package registry. A blanket replace got this wrong first and the failure read "the package
    registry is served on `<nil>`", which does not say "you renamed an interface" — so a test
    now pins every seat against the interface it delivers.
    

Done when. The fixtures are produced and consumed byte for byte by every implementation that claims the capability, and a module built before any of this serves its tools unchanged on the new runtime. The step is not done when the code runs — two implementations that disagree about an envelope do not fail to compile, they ignore each other while both keep running, which is the failure ADR 0074 exists to catch.

And the shared library gained nothing but the binding. A new transport is when the pressure to add conveniences is highest, and ADR 0039's rule does not bend for it: a helper that arrives with the bus is a review failure, not a detail. Code shared among a module's own features stays in that module.

Step 4 — the core speaks it

Why here. The links exist from step 3, so the flows that are not on the bus at all can move onto it, and the beds that need a mesh living on NATS can finally run.

A blocker surfaced here that is not this step's to fix. Every event name in the catalogue is still written the way a routing key on the bus the mesh has is written, so the derivation design 29 §1 specifies turns a consumer's declaration into a subject no emitter publishes — thirty-seven manifests, and one that cannot be composed at all. Nothing fails on the bus the mesh runs on today, where a routing key is matched literally; it fails on the first mesh raised on the new bus and not before, which is why wiring the controller's own subscription is what found it. Opened as issue 127. It holds 4.2, 4.3 and the catch-up half of 4.5; the node-facing flows — enrolment, reports, heartbeats, a build's outcome — are unaffected, because those subjects are the mesh's own and derive from nothing a module declares.

  • 4.1 the full genesis bed — a mesh raised on NATS from nothing and living on it: a node enrols over TLS with a claimed token and the enrolment user cannot read a declaration; a push is held while the store restarts and applies after, nothing lost or duplicated; a node that was away gets exactly the newest declaration and refuses a replayed older one by sequence; an upgrade rolls out to two nodes; an event dead-letters after max-deliver; and an enrolment held by a nak-with-delay cycle still reaches the enrolling node, proving the reply travels in the payload and not the transport field the consumer's ack has claimed. The server-enforced permissions were proved at step 1 and are not re-proved here — nothing is outstanding but the bed itself. Both links speak NATS, the composition happens, and every claim above has a unit test or a check against a running server behind it. What none of them can stand in for is a mesh raising itself, which is what this bed is — so this is where the code stops and the lab starts

  • 4.2 a build source's change reaches the builder over the bus, and the build that follows is the one the change asked for — a build is work submitted to a role now (ADR 0129). Both sides are behind a seam with an implementation per bus, and on the bus being built one publish does what two did: the outcome is the role's own event, so the asker matches it by the id its request carried, the controller records it and the catalogue places it in the graph. A build machine needs a reply queue for nothing and a grant over nobody's inbox.

    Checked against a running server: the round trip; a third party on the role's event hearing
    the same outcome the asker did, which is what the decision rests on; work leaving the queue
    once settled, so no second machine repeats it; work submitted with no machine holding the role
    **waiting rather than failing**, and being done when one arrives; and work a machine handed
    back coming round again.
    
    The outcome carries the module name, because only the manifest says what was built and one
    message now has three readers. A failed build names none: it produced no module version, and
    the catalogue would otherwise place something that was never made.
    
  • [~] 4.3 an installation completes over the bus, with the same outcome as the path it replaces — the installer can raise it: a foundation template that stands up the server, writes the server's own settings and the mesh's first user list beside them, and starts a controller reaching the new bus. What remains is running it, which is 4.1's bed.

    **The mesh composes its own user list, and at genesis there is no mesh to compose one.** So the
    installer carries the first — the controller's account at a well-known bootstrap password,
    exactly as the store is reached at `postgres:bootstrap` and the old bus at `guest:guest`, and
    rotated with them. From the controller's first composition onward the file is the controller's.
    
    That surfaced a gap reading would not have found: the controller's own account exists before
    there is a controller to mint one, so nothing recorded a hash for it and its first composition
    would have left the writer out of the file it was writing — a bus nothing can connect to,
    produced by the thing connected to it. It records a hash of the credential it is using, and only
    when none is recorded, so a restart cannot put the bootstrap password back over a rotated one.
    
    **The carried list and the derived one are checked against each other**, because they are two
    statements of one fact and a mesh cannot be raised twice to find out they disagreed. A template
    granting less than the controller derives produces a mesh that comes up, connects, and is
    refused on its first act, with an authorisation error naming a subject rather than the template
    that forgot it. The check earned itself at once: the composer was granting a role's whole event
    branch *and* the one event it follows, and the wider grant wins — so only the submitting half of
    a role is granted now, and what comes back is named exactly.
    
  • 4.4 a person's client — the account and the program are both in.

    **The account**: a person is not a module and holds no seat, so their authority is a list of
    tools (or `*` for an administrator) and nothing else. Held to four properties, each a way of
    being wrong that would not announce itself: nothing but tools, so a person cannot claim a
    module said something; no ack subject, because authority over a consumer that does not exist
    is authority nobody audits; no ability to answer, because a person who can answer a request is
    impersonating a module on a bus where anyone may serve a tool; and two people do not share an
    inbox. Issued, listed and revoked by command; stating what somebody may call replaces what was
    there, because a list that could only grow is a permission nobody can take back; and forgetting
    somebody takes their credential with them, or it is not a revocation.
    
    **The program**: two surfaces over one thing — a command line and an MCP server — both adapters
    over the same three calls, because a second way of reaching a tool is a second thing to keep
    correct. It uses the client a module's runtime uses, so what a person may do is answered by the
    same permission list that answers it for a module and an audit has nothing separate to read.
    
    Three decisions in it worth keeping. It lists what the **catalogue** has rather than what this
    credential may call: somebody seeing only their own tools cannot tell "not installed" from "not
    yours", and those need different people to fix them. A failed call says which of three things
    happened — nobody serves it, this credential may not, or the tool was slow — because the
    remedies are in three different places and without that they are one timeout and a stack trace.
    And the MCP surface decides nothing: the names are the ones a person types, the schemas are the
    modules' own, an answer is passed through unshaped, and a tool that fails comes back as a tool
    error rather than a protocol error, because the request was well-formed and the mesh answered it.
    
    Both surfaces are driven against a running bus, including a host's notification being answered
    with nothing and an unknown method refused.
    
    > **Design 25 §7 says "nothing is built of this before §10's bed passes", and this was built
    > before.** Recorded rather than quietly ignored: the operator asked for it, it is on the
    > critical path for nothing and blocked by nothing, and the bed it waits for is 4.1's. If the
    > bed changes what a person's client should be, this is what gets changed.
    
  • 4.5 reports and catch-up: a node that was unreachable catches up rather than losing them.

    **The reports half is in and proved against a server** (3.4): held through the store's absence by
    the server rather than by the controller, superseded ones settled by the digest they carry.
    
    **The catch-up half needed nothing built, and that was the answer.** It existed because a queue on
    the bus the mesh runs on today receives only what is published after it is bound, so everything
    built before the catalogue existed was announced to nobody — and on a fresh mesh that is always
    the foundation, because those are the things the catalogue needed in order to exist
    ([issue 050](../../04-ISSUES/050-the-catalogue-knows-nothing-built-before-it/00-report.md)). A
    whole mechanism followed: the catalogue asks, the controller re-publishes.
    
    A stream is a log and a consumer is a position in it. A consumer created later starts at the
    beginning, so the builds are simply there — asked of a running server rather than assumed, since
    the decision rested on it: three builds published with nothing listening, then a consumer created,
    and all three waiting for it. So the question of *who replays* has no answer because nothing
    replays.
    
    > **This is the shape of the whole change, in one task.** Three ways to do the replay were weighed
    > — a namespace for the mesh's own voice, the controller answering a question, a consumer reading
    > from the start — and the right answer was that the bus being moved to already does it. The
    > mechanism was never about builds; it was about a queue that could not remember. **A conversion
    > that carried it across would have carried a workaround for a limitation that no longer exists**,
    > and nothing would have looked wrong.
    
    Retiring it is step 5's, with the rest of what only the old bus needs: the request, the
    re-publishing, and the `replay` flag that told a consumer to register history without acting on it.
    

Done when. Each converted flow is proved against the behaviour it replaced, and the full genesis bed is green. Observation is not in this step — heartbeats, conditions and key-value state are research 017's, that effort already reserves them for after the move, and a flow built ahead of its design would be rebuilt.

Step 5 — the rollout

Why here. It is the only step that moves a node's bus, and it moves every node's at once.

The order the repositories land in is part of the rollout, not paperwork. Derived 2026-09-27 while merging, and not obvious from any one repository, which is why it is written here rather than left to be re-derived under time pressure:

order repository why it cannot be later
1 this one prose; nothing deploys
2 the sdk comments only, and no module rebuilds for it
3 the client library it is what a module calls to emit, and it is where the subject is derived. Until it lands, a locally-named event is published under the local name itself
4 the catalogue every manifest and every module's code, renamed together. Safe only once the runtime derives
5 the controller it refuses an old-style event name outright, so landing it before the catalogue makes every unconverted module unregisterable
6 the hosts last, because nothing else waits on them

Two properties make the sequence safe rather than merely ordered, and both are pinned by tests. A name already in the old form passes through the derivation untouched, so a module nobody has converted keeps working at every step. And a converted name derives to exactly the key the old bus published, so steps 3 and 4 change nothing on the wire — the move to the new bus is step 5.2 and one environment variable, not a side effect of deploying.

The failure this ordering avoids is issue 127's own: a publisher and a subscriber that disagree about a subject produce no error anywhere. Nothing logs, nothing retries, and the mesh reports itself healthy while reacting to nothing.

  • 5.1 the cutover bed: a mesh on AMQP with a predecessor stand-in on the deprecated broker moves its bus in one rollout, every node reporting on NATS afterwards, the stand-in's own client still connected throughout

  • [~] 5.2 the rollout: accounts composed, then the controller, every host and every runtime together; every node confirmed heard before AMQP stops.

    **The readiness half is in and is the half worth having.** The move takes every node at once, so
    there is nothing to inspect afterwards and no half to roll back — either the mesh was ready or it
    was not. `rollout check` answers that from records, with one dial: is a bus answering, does a
    machine hold the seat, has it been sent the composed user list, does every machine and every
    module that speaks have a credential. Each missing thing names its own next step, because "not
    ready" that cannot be acted on is not an answer at the point where the next step is irreversible.
    
    **A machine with no credential is what must stop it.** It keeps running, cannot come back, and
    afterwards there is no bus to tell it anything over.
    
    The move itself is deliberately not written yet, and the command says so rather than pretending:
    it waits on the check having been run against a real mesh. Writing the irreversible half before
    the question it depends on has ever been asked of something real is how the plan's own rule about
    beds gets broken by another route.
    
    > **What this costs if it goes wrong, measured rather than assumed.** Nothing in a served
    > request's path goes over the mesh's own bus: modules serve from their own containers. What a
    > failed move costs is the mesh's ability to *change* anything — pushes, tool calls, new
    > provisioning — until it is finished or undone. That is worth knowing before rather than
    > after, and it is why the operator's "as long as my services keep running" is a reasonable
    > position rather than a gamble. **Measured on 2026-09-27**, when a seat emptied itself
    > mid-change: 52 containers stayed up and the broker never stopped; the control plane
    > crash-looped for two hours and nothing could be deployed until it was repaired by hand.
    > An earlier version of this note said the old broker stays as an ordinary provider of
    > `amqp` ([ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md)); that is withdrawn by [ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md) — see 5.4.
    
    **What the first live attempt found, 2026-09-27.** With the seat handed over on record and the
    new bus's module registered and assigned beside the old one, `push` refused the control node:
    *not one user has a credential for the new bus*. The check is right — a bus whose user list is
    empty refuses every connection in the mesh — and it exposed the half of this task nobody had
    built. A credential is minted at three moments only: a machine's at enrolment, a module's at
    `module issue`, a person's at `operator`. **Nothing mints one for a machine already enrolled, or
    for the control plane itself.** And on the host, the membership — bus address, fingerprint,
    password, transport — is written once, at enrolment, and nothing ever rewrites it. So "move
    each machine and confirm it reports" had no mechanism under it on either side.
    
    The mechanism, to build before anything moves:
    - **the control plane mints what is missing** — every user the records derive with no hash —
      and delivers each plaintext where its owner reads it: a machine's inside its declaration, as a
      sealed *membership* for the new bus (address, fingerprint, password, transport); a module's as
      its broker secret, the path `module issue` already uses; the control plane's own as its module
      secret, so it reads it the way any module does;
    - **the host saves a delivered membership and re-dials on it** — the same file enrolment wrote,
      the same reconnect path a lost connection takes, so a machine moved this way is a machine
      that came back, and nothing new has to be right for it to work;
    - **the switch is then two acts in one push**: `MESH_BUS_NATS` on the control plane, and
      `seat mesh-broker --to <node>/<the new bus's module>` — the seat never empty, every machine
      already holding a credential that works on the other side.
    
    Until the first bullet exists the check keeps refusing, and it should: a machine moved without
    a credential cannot come back, and afterwards there is no bus to tell it anything over.
    
  • 5.3 the seat changes hands as one act. A command takes a seat and the assignment taking it over, and the seat is never empty in between — the emptiness is the outage of 2026-09-27, when the control plane, which finds its own bus through this seat, lost the address and looped. Built 2026-09-27 (mesh-controller seat_holding, migration 0039; design 26 says how it is checked). Its first live use recorded the standing holder — which the row moving under it had made unable to satisfy what the seat delivers, so the first handover on a mesh that predates the record writes down who holds without re-judging them. This is what 5.2 uses to move mesh-broker from the old broker's assignment to the new one's (ADR 0131).

  • [~] 5.4 the old broker and everything that named AMQP leave the mesh — the catalogue half done 2026-09-27 (three modules removed; registration refuses the word; the seat's row delivers mesh-bus, migration 0040); the live half — unassigning the old broker — waits on 5.2 (ADR 0131, superseding ADR 0127): the two modules that required amqp are removed, the broker's module is unassigned and removed, registration refuses a manifest that provides or requires amqp, and a whole-catalogue check asserts none does. Not a retirement condition — a decision, taken, with the operator's "I don't care if the predecessor breaks" on record (ADR 0130). Retiring with it: the build outcome's second announcement under the module's own name, which existed only so a catalogue deployed before the rename and one after both heard it.

    > **The remote tooling goes with it too.** The predecessor's own mesh talks over that broker, so
    > shutting it down ends the path that reaches this installation's machines from a workstation.
    > The rollout is driven from the node, or before the broker stops — a sequencing constraint on
    > 5.2, not an afterthought.
    
  • 5.5 the AMQP transport is deleted from the control plane and the hosts, and the variable that selected a transport is refused at start as unknown. One bus, nothing to select (ADR 0131).

The old 5.4 note is history. It recorded that a retirement condition was wrong from ADR 0127 onward, which framed the old broker as an ordinary provider with no end. ADR 0131 ends that framing in turn: the broker is not kept as a provider either, because AMQP is not a provision. Both readings are kept here so the two reversals can be read in order.

Done when. Every node reports on NATS, and nothing of the mesh's own is left connected to the deprecated broker.

The through-line

The order is dependency, not preference. Steps 1 to 4 leave every node on AMQP, so the cost of being wrong is bounded until the last step: a step may be abandoned, or reordered after step 2, without a rollback. The server stands before anything speaks to it; the wire is specified before it is implemented three times; the flows move once there is something to move them onto; and the bus itself moves once, at the end, on one day.

What is deliberately not here

  • Observation — research 017's, after the move, by its own design.
  • Leaf nodes — design 25 §11 keeps this out of scope and says so; a leaf per machine is a later question, noted so it is not forgotten.
  • The predecessor's world. It is AMQP and it is not moving — ADR 0130: it is deprecated, some of it is still running, and it is being left to stop rather than migrated. Its broker goes with it, unassigned like any provider whose provision nothing requires.

How this list is kept true

A task is ticked when its change is committed, not when it is written. A step is done when its bed is green, not when its tasks are ticked. If a step's tasks are all ticked and its bed has not run, the step is in progress and this document says so — that gap is the thing the whole shape is built to make visible.