Files
hq/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md
T
jschoubben eef54917ec Issue 146: the double enrolment was two consumers sharing a delivery subject
Not about enrolment. A push consumer delivers onto an ordinary subject and
everything subscribed to it gets a copy; the controller's two consumers were
both named after it, so both were given the same subject and the one process
acted on every message twice. Enrolment is where it drew blood because a second
enrolment mints a second credential.
2026-09-29 21:32:52 +02:00

11 KiB

Diagnosis

2026-09-29, by raising a first node in the lab over and over and writing down each thing it hit.

Not one fault. Four, stacked, each hidden behind the one before it, and every one of them the same shape: a step that was right while the mesh ran on the previous broker and was never asked a question again after the bus changed. Nothing had raised a foundation since, so nothing said so.

1 — the bundle's bus image is named for a registry that is gone (fixed)

The newer bundle pins <a lab registry>/nats@…, which resolves nowhere outside the lab that raised that registry. The lab already rewrites the store's and the previous broker's references to upstream ones for a machine with an uplink; the bus had no such rule because no bed had ever tried to raise this bundle. Added (mesh-lab test/integration/harness.ts). The digest is the bundle's own — what the registry served was a copy, so the same digest resolves upstream, and this is a prefix being removed rather than a reference being replaced.

2 — the bus's certificate was made by a tool the bus does not have (fixed)

failed bus-certificate … docker exited 127

The step ran openssl inside the broker's image. The previous broker's image carried it; the bus's does not — it is Alpine with a shell and no openssl — and neither does any other image the bundle names, so there was nothing to substitute. The program that needs the certificate now makes it: mesh-controller broker certificate --into <dir>, with --check as the step's verify. The controller is already on the machine at that point (the schema step ran it) and needs nothing from the image it writes into. Self-signed, as before and on purpose — a host pins this server's exact certificate (ADR 0004) and at that moment there is no authority to ask. Idempotent, because a second certificate is one every host that pinned the first no longer believes. It runs --user 0:0: the volume is root's, and the control plane's image runs as nobody, which is right for the long-lived server and wrong for a one-shot writing into a fresh volume.

3 — enrolment dialled TLS at a bus that speaks first (fixed)

mesh-host: cannot reach the broker at …:5671: tls: first record does not look like a TLS handshake

Enrolment opened a raw TLS connection to check the pinned certificate before saying anything. NATS speaks its own protocol and upgrades afterwards, so the handshake met a plaintext greeting. The pin was never the problem: the same pinned configuration is handed to the client that presents the token, and the verification runs inside that handshake — so what ADR 0004 requires still holds, and holds better, because the one-time secret is sent only after the certificate has been checked. The raw dial is gone from the enrolment path and kept only as what its tests always proved: that a wrong certificate is refused before a byte of application data is sent.

Then, immediately behind it:

mesh-host: this token is for the "" bus, and the mesh's bus is nats

The enrolment left the transport empty and meant whatever the mesh runs today, which was true while two buses existed and became a refusal the moment one did. The host knows which bus the mesh runs; it says so now.

4 — a first node cannot be let onto its own bus (open, and this is the real one)

mesh-host: cannot reach the bus at …:5671 as anchor: nats: Authorization Violation

The bus's user list is a file beside its configuration. The installer carries the first one — the controller's own account at a bootstrap password — and the controller composes every user after that (design 25 §6; the controller's own test asserts the carried list matches what it would derive). On the running mesh that composition reaches the bus because the bus is a module, with the list delivered to it the way anything is delivered to a module.

At genesis there is no module. The foundation's bus is raised by the installer, the control plane is given no way to write beside its configuration — it mounts the certificate and nothing else — and so the account a joining node needs cannot come into existence. The first node cannot join the mesh it just raised.

That is not a line to fix in a bundle. It is the open half of the mesh delivering its own components (ADR 0142) and of the bus becoming a module: either the installer's bus is raised as the module the mesh will go on managing, or genesis carries a user list that includes the first node's enrolment and the controller takes over from there. Both are decisions, not patches, and both belong to the genesis step that was deliberately left until last.

5 — the composed user list has to be placed by hand at genesis (fixed)

The account a token is the password of is not recorded at all: the composer names an enrolment user for every machine with a live token, nothing minted a credential for it, and the composition left it out as a user with no password. The comment above the issuing code already claimed otherwise — "the account is created before the token is handed over" — which is how it went unnoticed. Issuing a token now records that account, with the token's own secret as its password, because that is the string the machine will present.

Placing it is the other half. The list reaches the machine running the bus in that machine's declaration, which a machine that has not enrolled does not get, so at genesis it cannot arrive that way. The control plane composes and says what it composed — broker accounts, to standard output — and whoever is raising the machine writes it beside the bus's configuration and makes the server re-read it. Twice, because two accounts come into existence at different moments: the enrolment when the token is issued, and the machine's own when it enrols. A control plane that wrote the file itself would have to know where the bus keeps its configuration and how to make it reload, which is the module's knowledge and is what the module takes over on the first push.

With that, a first node enrols against the bus it just raised — measured, from bare, in the lab.

6 — and is enrolled twice, keeping a credential the mesh has replaced (fixed)

mesh-controller: enrolled anchor
mesh-controller: enrolled anchor      (the same second)

One enrol on the machine, two enrolments in the control plane. Each mints the node a fresh bus password and returns it; the machine keeps the answer to the first, and the mesh keeps the hash of the second. The machine then reconnects for ever as a user whose password the mesh rotated out from under it — authentication error - User "node.anchor" on the bus, Authorization Violation in the host's log, and a node that never reports.

What is ruled out: the host asking twice — it asks again only when the mesh says try again, and a refused attempt is not logged as an enrolment. Redelivery by the consumer — there is one consumer, its acknowledgement window is thirty seconds, and the handler is quick.

What is left: the client re-publishing when an acknowledgement is slow, which is what its defaults do. That was addressed by giving the publish a message id derived from its own bytes, so the stream discards the copy — and the duplicate survived it, so either the id is not reaching the stream or the second copy is not a copy. This is where the trail stops.

Worth saying plainly: the mint is the fragile part, not the delivery. An enrolment answered twice is survivable if the answer is the same both times, and it cannot be — the mesh keeps only the hash, so a second answer is necessarily a different credential.

Found, and it is not about enrolment at all (fixed). The bus's own counters settled it: one message published, one held in the stream, one delivery, nothing redelivered — and the controller enrolled the machine twice. So the handler ran twice on one delivery.

A push consumer delivers onto an ordinary subject, and everything subscribed to that subject gets a copy. The controller holds a consumer called controller on CONTROL and another called controller on EVENTS, and the delivery subject was derived from the consumer's name alone — so both were _DELIVER.controller, the one process held both subscriptions, and every message from either stream was acted on twice.

Enrolment is where it drew blood, because enrolling twice mints twice and the second credential replaces the first. But it applied to every report and every event the controller follows, and it is the kind of fault that leaves no trace: nothing is redelivered, no counter is wrong, the work simply happens twice. The comment in the receiving code about a merge that ran the whole catalogue five times over on 2026-09-28 is the same shape seen from the other end.

The stream is in the delivery subject now, because the pair is what identifies a consumer — the server scopes a durable's name to its stream, and this subject was the one place that scoping was dropped. A subscriber's permission gains the same shape, keeping the bare name so an existing consumer keeps working until the controller's next assertion moves it.

How it is checked: the consumers the mesh asks for are asserted to deliver onto distinct subjects, in the controller's own suite. Against a server it would be invisible, which is the point.

Where it belongs

mesh-host (the bundle and the enrolment path) and mesh-controller (the certificate command, and the composition that cannot reach the bus at genesis). Three of the four are fixed on branches; the fourth is the genesis work.

What made it slow, and what was changed so it is not

Six faults behind one another, each found by raising a machine and reading what it said. What cost the most was not the faults:

  • Every bed's own instructions named the bundle that cannot work, so the first three attempts ended in a control plane crash-looping on a missing bus. They name the working one now.
  • A host binary built without its system refuses everything it is given with this host was built for "", which reads like a broken bundle. The lab's README says so.
  • make image in the control plane had been broken for as long as its base was pinned: the Dockerfile's fallback is a Go older than the module asks for, and the pipeline never saw it because the pipeline passes the declared base in. It reads the base from the manifest now.
  • Leaving the machine standing is what answers the question. Every finding above came from shelling in afterwards — the host's log, the bus's log, the file the bus was actually handed — and none from the test's own output, which says only that nothing converged. The bed takes MESH_LAB_KEEP, and the README says to reach for it first.

What it cost, for the next person

Every lab bed still names foundation-first-node.lock in its own instructions, and that bundle raises the previous broker with a control plane that refuses to start without MESH_BUS_NATS. Until the fourth fault is answered and the two bundles become one, a bed runs with MESH_LAB_BUNDLE pointing at the NATS bundle by hand, and stops at the enrolment.