From f6ed3545b765ba6e6bb36148456a12b728b43b33 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 17:36:40 +0200 Subject: [PATCH 1/5] Issue 146: a first node now enrols, and is enrolled twice MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two more faults behind the three already fixed. The account a token is the password of was never recorded, and the comment above the issuing code said it was; issuing now records it. Placing the composed list at genesis is the other half — the control plane says what it composed and whoever raises the machine writes it beside the bus, because no declaration can reach a machine that has not enrolled. With that a first node enrols. It is then enrolled twice from one attempt, each minting a credential, and it keeps the answer to the first while the mesh keeps the second. The trail for that one stops at a duplicate that survived message-id deduplication. --- .../01-diagnosis.md | 48 +++++++++++++++++++ 1 file changed, 48 insertions(+) diff --git a/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md b/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md index bcbd17b..34c2817 100644 --- a/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md +++ b/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md @@ -82,6 +82,54 @@ managing, or genesis carries a user list that includes the first node's enrolmen takes over from there. Both are decisions, not patches, and both belong to the genesis step that was deliberately left until last. +## 5 — the composed user list has to be placed by hand at genesis *(fixed)* + +The account a token is the password of is **not recorded at all**: the composer names an enrolment +user for every machine with a live token, nothing minted a credential for it, and the composition +left it out as a user with no password. The comment above the issuing code already claimed +otherwise — *"the account is created before the token is handed over"* — which is how it went +unnoticed. Issuing a token now records that account, with the token's own secret as its password, +because that is the string the machine will present. + +Placing it is the other half. The list reaches the machine running the bus in that machine's +declaration, which a machine that has not enrolled does not get, so at genesis it cannot arrive +that way. **The control plane composes and says what it composed** — `broker accounts`, to standard +output — and whoever is raising the machine writes it beside the bus's configuration and makes the +server re-read it. Twice, because two accounts come into existence at different moments: the +enrolment when the token is issued, and the machine's own when it enrols. A control plane that +wrote the file itself would have to know where the bus keeps its configuration and how to make it +reload, which is the module's knowledge and is what the module takes over on the first push. + +With that, **a first node enrols against the bus it just raised** — measured, from bare, in the +lab. + +## 6 — and is enrolled twice, keeping a credential the mesh has replaced *(open)* + +``` +mesh-controller: enrolled anchor +mesh-controller: enrolled anchor (the same second) +``` + +One `enrol` on the machine, two enrolments in the control plane. Each mints the node a fresh bus +password and returns it; the machine keeps the answer to the first, and the mesh keeps the hash of +the second. The machine then reconnects for ever as a user whose password the mesh rotated out from +under it — *authentication error - User "node.anchor"* on the bus, `Authorization Violation` in the +host's log, and a node that never reports. + +What is ruled out: the host asking twice — it asks again only when the mesh says *try again*, and +a refused attempt is not logged as an enrolment. Redelivery by the consumer — there is one +consumer, its acknowledgement window is thirty seconds, and the handler is quick. + +What is left: the client re-publishing when an acknowledgement is slow, which is what its defaults +do. That was addressed by giving the publish a message id derived from its own bytes, so the stream +discards the copy — **and the duplicate survived it**, so either the id is not reaching the stream +or the second copy is not a copy. This is where the trail stops. + +Worth saying plainly: **the mint is the fragile part, not the delivery.** An enrolment answered +twice is survivable if the answer is the same both times, and it cannot be — the mesh keeps only +the hash, so a second answer is necessarily a different credential. Whatever closes this either +makes the enrolment arrive once, or stops the second arrival from rotating anything. + ## Where it belongs `mesh-host` (the bundle and the enrolment path) and `mesh-controller` (the certificate command, and From e9b1010bc0c0fa58dc885f10a8ecb46f3f9137b5 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 17:45:25 +0200 Subject: [PATCH 2/5] Issue 146: what made it slow, and what was changed so it is not --- .../01-diagnosis.md | 17 +++++++++++++++++ 1 file changed, 17 insertions(+) diff --git a/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md b/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md index 34c2817..6bd82ee 100644 --- a/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md +++ b/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md @@ -136,6 +136,23 @@ makes the enrolment arrive once, or stops the second arrival from rotating anyth the composition that cannot reach the bus at genesis). Three of the four are fixed on branches; the fourth is the genesis work. +## What made it slow, and what was changed so it is not + +Six faults behind one another, each found by raising a machine and reading what it said. What cost +the most was not the faults: + +- **Every bed's own instructions named the bundle that cannot work**, so the first three attempts + ended in a control plane crash-looping on a missing bus. They name the working one now. +- **A host binary built without its system** refuses everything it is given with *this host was + built for ""*, which reads like a broken bundle. The lab's README says so. +- **`make image` in the control plane had been broken for as long as its base was pinned**: the + Dockerfile's fallback is a Go older than the module asks for, and the pipeline never saw it + because the pipeline passes the declared base in. It reads the base from the manifest now. +- **Leaving the machine standing is what answers the question.** Every finding above came from + shelling in afterwards — the host's log, the bus's log, the file the bus was actually handed — + and none from the test's own output, which says only that nothing converged. The bed takes + `MESH_LAB_KEEP`, and the README says to reach for it first. + ## What it cost, for the next person Every lab bed still names `foundation-first-node.lock` in its own instructions, and that bundle From eef54917ec4a08479b8eb425aa3137391e670b52 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 21:32:52 +0200 Subject: [PATCH 3/5] Issue 146: the double enrolment was two consumers sharing a delivery subject Not about enrolment. A push consumer delivers onto an ordinary subject and everything subscribed to it gets a copy; the controller's two consumers were both named after it, so both were given the same subject and the one process acted on every message twice. Enrolment is where it drew blood because a second enrolment mints a second credential. --- .../01-diagnosis.md | 29 +++++++++++++++++-- 1 file changed, 26 insertions(+), 3 deletions(-) diff --git a/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md b/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md index 6bd82ee..317bb01 100644 --- a/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md +++ b/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md @@ -103,7 +103,7 @@ reload, which is the module's knowledge and is what the module takes over on the With that, **a first node enrols against the bus it just raised** — measured, from bare, in the lab. -## 6 — and is enrolled twice, keeping a credential the mesh has replaced *(open)* +## 6 — and is enrolled twice, keeping a credential the mesh has replaced *(fixed)* ``` mesh-controller: enrolled anchor @@ -127,8 +127,31 @@ or the second copy is not a copy. This is where the trail stops. Worth saying plainly: **the mint is the fragile part, not the delivery.** An enrolment answered twice is survivable if the answer is the same both times, and it cannot be — the mesh keeps only -the hash, so a second answer is necessarily a different credential. Whatever closes this either -makes the enrolment arrive once, or stops the second arrival from rotating anything. +the hash, so a second answer is necessarily a different credential. + +**Found, and it is not about enrolment at all** *(fixed)*. The bus's own counters settled it: one +message published, one held in the stream, one delivery, nothing redelivered — and the controller +enrolled the machine twice. So the handler ran twice on one delivery. + +A push consumer delivers onto an ordinary subject, and **everything subscribed to that subject gets +a copy**. The controller holds a consumer called `controller` on CONTROL and another called +`controller` on EVENTS, and the delivery subject was derived from the consumer's name alone — so +both were `_DELIVER.controller`, the one process held both subscriptions, and every message from +either stream was acted on twice. + +Enrolment is where it drew blood, because enrolling twice mints twice and the second credential +replaces the first. But it applied to **every report and every event the controller follows**, and +it is the kind of fault that leaves no trace: nothing is redelivered, no counter is wrong, the work +simply happens twice. The comment in the receiving code about a merge that ran the whole catalogue +five times over on 2026-09-28 is the same shape seen from the other end. + +The stream is in the delivery subject now, because the pair is what identifies a consumer — the +server scopes a durable's name to its stream, and this subject was the one place that scoping was +dropped. A subscriber's permission gains the same shape, keeping the bare name so an existing +consumer keeps working until the controller's next assertion moves it. + +*How it is checked:* the consumers the mesh asks for are asserted to deliver onto distinct subjects, +in the controller's own suite. Against a server it would be invisible, which is the point. ## Where it belongs From 0dd00e88b61a92d54a85a30866723d9ec0d1f1df Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 21:39:55 +0200 Subject: [PATCH 4/5] Issue 147: the operator's tools still dial the bus that was removed MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Every tool call fails with 'AMQP not connected', on every node including the local one, because the tool surface still opens an AMQP connection and that transport was deleted at the cut-over. The mesh reports healthy throughout — what broke is the thing standing outside asking it questions, so nothing the mesh checks is about it. --- .../00-report.md | 58 +++++++++++++++++++ 1 file changed, 58 insertions(+) create mode 100644 04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md diff --git a/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md b/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md new file mode 100644 index 0000000..f4e2d57 --- /dev/null +++ b/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md @@ -0,0 +1,58 @@ +--- +status: open +opened: 2026-09-29 +located-in: [] +--- + +# 147 — the operator's tools still dial the bus that was removed + +## What was observed + +Every tool call an operator makes against the mesh fails, on every machine, with the same answer: + +``` +AMQP not connected — cannot reach hal/mesh@novox +AMQP not connected — cannot reach hal/mesh@shanks +``` + +The mesh moved to one bus and the previous transport was deleted +([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md), cut over +2026-09-28). The surface an operator drives the mesh through — the tool bridge their assistant +speaks to — still opens an AMQP connection, so it cannot reach anything. Not one node answers, +including the machine the operator is sitting at. + +**What that leaves.** The mesh itself is healthy: nodes are current, modules run, the bus carries +the mesh's own traffic. What is gone is the way a person asks it anything. Every question — what a +node runs, what is assigned, what a module's settings are — has to be asked by opening a shell on +the machine and running the control plane's binary inside its container, which is precisely the +path the tool surface exists to remove, and which nothing checks, records or permits. + +It also silently changes how work gets done: an assistant told to use the mesh's tools finds them +dead, falls back to `ssh` and `docker exec`, and the operator discovers the fallback rather than +the fault. + +## Why this is here and not a note in the knowledge base + +The rule is that everything happens over the mesh's bus. A surface that cannot reach that bus is +that rule enforced by nothing — and the mesh reports itself healthy throughout, because what broke +is not a module, a node or a provision but the thing standing outside asking them questions. + +The same cut-over removed the assistant's long-term memory (`recall`, `memorize`) for the same +reason, which is why lessons from the last two days were written into this repository by hand. + +## What would have prevented it + +- **The tool surface as a consumer of the bus like any other**, so moving the bus moves it — rather + than a separate bridge with its own connection settings that nothing resolves. +- **A check that a tool call reaches a node**, run where the mesh's other checks run. Every check + the mesh makes today is about the relationship between the mesh and a machine; none asks whether + a person can ask it anything ([issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md), + the same shape one level out). + +## Evidence to carry into diagnosis + +- The failure text names `hal/mesh@`, so the bridge is resolving a node and then dialling the + old transport. +- It fails identically for the local machine, which rules out reachability and points at the + transport alone. +- The mesh's own traffic over the new bus is unaffected: nodes report, declarations apply. From 3c0f7082e64b7f52c4e482d908cd2f3686d5ac27 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 21:50:50 +0200 Subject: [PATCH 5/5] Issue 147: the tool surface is not the mesh's, it is the predecessor's MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Diagnosed from the configuration: the tool server is HAL's brain, a local process on the workstation with the predecessor's broker URL in the assistant's own config. No manifest, no assignment, no seat, no account. Nothing regressed — the mesh removed a transport this program still dials, and the program was never part of the mesh. The mesh has a tool model and nothing publishes an operator-facing surface onto it. --- .../00-report.md | 21 +++++++++++++++++-- 1 file changed, 19 insertions(+), 2 deletions(-) diff --git a/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md b/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md index f4e2d57..b056a21 100644 --- a/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md +++ b/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md @@ -1,7 +1,7 @@ --- -status: open +status: located opened: 2026-09-29 -located-in: [] +located-in: [nothing in the mesh — the surface in use is the predecessor's, installed on the workstation] --- # 147 — the operator's tools still dial the bus that was removed @@ -49,6 +49,23 @@ reason, which is why lessons from the last two days were written into this repos a person can ask it anything ([issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md), the same shape one level out). +## Diagnosed at once, because the answer was in the configuration + +**The surface is not the mesh's.** The tool server the operator's assistant speaks to is the +predecessor's brain, installed on the workstation and started as a local process, with the +predecessor's broker URL — `amqp://…@` — written into the assistant's own +configuration. Nothing about it is a module: it has no manifest, no assignment, no seat, no account +on the mesh's bus, and the mesh has never known it exists. + +So nothing regressed. The mesh removed a transport that this program still dials, and the program +was never part of the mesh to be moved. **The mesh has a tool model** — `tools` in a manifest, the +calls a seat accepts, the subjects a module answers on, and a control-plane command that asks one — +and **nothing publishes an operator-facing surface onto it.** What an operator uses is the thing +that came before, kept alive by a URL in a file. + +That is the issue, and it is larger than a broken connection: the way a person drives this mesh is +outside the mesh. + ## Evidence to carry into diagnosis - The failure text names `hal/mesh@`, so the bridge is resolving a node and then dialling the