From f6ed3545b765ba6e6bb36148456a12b728b43b33 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 17:36:40 +0200 Subject: [PATCH 01/13] Issue 146: a first node now enrols, and is enrolled twice MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two more faults behind the three already fixed. The account a token is the password of was never recorded, and the comment above the issuing code said it was; issuing now records it. Placing the composed list at genesis is the other half — the control plane says what it composed and whoever raises the machine writes it beside the bus, because no declaration can reach a machine that has not enrolled. With that a first node enrols. It is then enrolled twice from one attempt, each minting a credential, and it keeps the answer to the first while the mesh keeps the second. The trail for that one stops at a duplicate that survived message-id deduplication. --- .../01-diagnosis.md | 48 +++++++++++++++++++ 1 file changed, 48 insertions(+) diff --git a/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md b/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md index bcbd17b..34c2817 100644 --- a/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md +++ b/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md @@ -82,6 +82,54 @@ managing, or genesis carries a user list that includes the first node's enrolmen takes over from there. Both are decisions, not patches, and both belong to the genesis step that was deliberately left until last. +## 5 — the composed user list has to be placed by hand at genesis *(fixed)* + +The account a token is the password of is **not recorded at all**: the composer names an enrolment +user for every machine with a live token, nothing minted a credential for it, and the composition +left it out as a user with no password. The comment above the issuing code already claimed +otherwise — *"the account is created before the token is handed over"* — which is how it went +unnoticed. Issuing a token now records that account, with the token's own secret as its password, +because that is the string the machine will present. + +Placing it is the other half. The list reaches the machine running the bus in that machine's +declaration, which a machine that has not enrolled does not get, so at genesis it cannot arrive +that way. **The control plane composes and says what it composed** — `broker accounts`, to standard +output — and whoever is raising the machine writes it beside the bus's configuration and makes the +server re-read it. Twice, because two accounts come into existence at different moments: the +enrolment when the token is issued, and the machine's own when it enrols. A control plane that +wrote the file itself would have to know where the bus keeps its configuration and how to make it +reload, which is the module's knowledge and is what the module takes over on the first push. + +With that, **a first node enrols against the bus it just raised** — measured, from bare, in the +lab. + +## 6 — and is enrolled twice, keeping a credential the mesh has replaced *(open)* + +``` +mesh-controller: enrolled anchor +mesh-controller: enrolled anchor (the same second) +``` + +One `enrol` on the machine, two enrolments in the control plane. Each mints the node a fresh bus +password and returns it; the machine keeps the answer to the first, and the mesh keeps the hash of +the second. The machine then reconnects for ever as a user whose password the mesh rotated out from +under it — *authentication error - User "node.anchor"* on the bus, `Authorization Violation` in the +host's log, and a node that never reports. + +What is ruled out: the host asking twice — it asks again only when the mesh says *try again*, and +a refused attempt is not logged as an enrolment. Redelivery by the consumer — there is one +consumer, its acknowledgement window is thirty seconds, and the handler is quick. + +What is left: the client re-publishing when an acknowledgement is slow, which is what its defaults +do. That was addressed by giving the publish a message id derived from its own bytes, so the stream +discards the copy — **and the duplicate survived it**, so either the id is not reaching the stream +or the second copy is not a copy. This is where the trail stops. + +Worth saying plainly: **the mint is the fragile part, not the delivery.** An enrolment answered +twice is survivable if the answer is the same both times, and it cannot be — the mesh keeps only +the hash, so a second answer is necessarily a different credential. Whatever closes this either +makes the enrolment arrive once, or stops the second arrival from rotating anything. + ## Where it belongs `mesh-host` (the bundle and the enrolment path) and `mesh-controller` (the certificate command, and -- 2.54.0 From e9b1010bc0c0fa58dc885f10a8ecb46f3f9137b5 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 17:45:25 +0200 Subject: [PATCH 02/13] Issue 146: what made it slow, and what was changed so it is not --- .../01-diagnosis.md | 17 +++++++++++++++++ 1 file changed, 17 insertions(+) diff --git a/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md b/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md index 34c2817..6bd82ee 100644 --- a/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md +++ b/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md @@ -136,6 +136,23 @@ makes the enrolment arrive once, or stops the second arrival from rotating anyth the composition that cannot reach the bus at genesis). Three of the four are fixed on branches; the fourth is the genesis work. +## What made it slow, and what was changed so it is not + +Six faults behind one another, each found by raising a machine and reading what it said. What cost +the most was not the faults: + +- **Every bed's own instructions named the bundle that cannot work**, so the first three attempts + ended in a control plane crash-looping on a missing bus. They name the working one now. +- **A host binary built without its system** refuses everything it is given with *this host was + built for ""*, which reads like a broken bundle. The lab's README says so. +- **`make image` in the control plane had been broken for as long as its base was pinned**: the + Dockerfile's fallback is a Go older than the module asks for, and the pipeline never saw it + because the pipeline passes the declared base in. It reads the base from the manifest now. +- **Leaving the machine standing is what answers the question.** Every finding above came from + shelling in afterwards — the host's log, the bus's log, the file the bus was actually handed — + and none from the test's own output, which says only that nothing converged. The bed takes + `MESH_LAB_KEEP`, and the README says to reach for it first. + ## What it cost, for the next person Every lab bed still names `foundation-first-node.lock` in its own instructions, and that bundle -- 2.54.0 From eef54917ec4a08479b8eb425aa3137391e670b52 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 21:32:52 +0200 Subject: [PATCH 03/13] Issue 146: the double enrolment was two consumers sharing a delivery subject Not about enrolment. A push consumer delivers onto an ordinary subject and everything subscribed to it gets a copy; the controller's two consumers were both named after it, so both were given the same subject and the one process acted on every message twice. Enrolment is where it drew blood because a second enrolment mints a second credential. --- .../01-diagnosis.md | 29 +++++++++++++++++-- 1 file changed, 26 insertions(+), 3 deletions(-) diff --git a/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md b/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md index 6bd82ee..317bb01 100644 --- a/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md +++ b/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md @@ -103,7 +103,7 @@ reload, which is the module's knowledge and is what the module takes over on the With that, **a first node enrols against the bus it just raised** — measured, from bare, in the lab. -## 6 — and is enrolled twice, keeping a credential the mesh has replaced *(open)* +## 6 — and is enrolled twice, keeping a credential the mesh has replaced *(fixed)* ``` mesh-controller: enrolled anchor @@ -127,8 +127,31 @@ or the second copy is not a copy. This is where the trail stops. Worth saying plainly: **the mint is the fragile part, not the delivery.** An enrolment answered twice is survivable if the answer is the same both times, and it cannot be — the mesh keeps only -the hash, so a second answer is necessarily a different credential. Whatever closes this either -makes the enrolment arrive once, or stops the second arrival from rotating anything. +the hash, so a second answer is necessarily a different credential. + +**Found, and it is not about enrolment at all** *(fixed)*. The bus's own counters settled it: one +message published, one held in the stream, one delivery, nothing redelivered — and the controller +enrolled the machine twice. So the handler ran twice on one delivery. + +A push consumer delivers onto an ordinary subject, and **everything subscribed to that subject gets +a copy**. The controller holds a consumer called `controller` on CONTROL and another called +`controller` on EVENTS, and the delivery subject was derived from the consumer's name alone — so +both were `_DELIVER.controller`, the one process held both subscriptions, and every message from +either stream was acted on twice. + +Enrolment is where it drew blood, because enrolling twice mints twice and the second credential +replaces the first. But it applied to **every report and every event the controller follows**, and +it is the kind of fault that leaves no trace: nothing is redelivered, no counter is wrong, the work +simply happens twice. The comment in the receiving code about a merge that ran the whole catalogue +five times over on 2026-09-28 is the same shape seen from the other end. + +The stream is in the delivery subject now, because the pair is what identifies a consumer — the +server scopes a durable's name to its stream, and this subject was the one place that scoping was +dropped. A subscriber's permission gains the same shape, keeping the bare name so an existing +consumer keeps working until the controller's next assertion moves it. + +*How it is checked:* the consumers the mesh asks for are asserted to deliver onto distinct subjects, +in the controller's own suite. Against a server it would be invisible, which is the point. ## Where it belongs -- 2.54.0 From 0dd00e88b61a92d54a85a30866723d9ec0d1f1df Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 21:39:55 +0200 Subject: [PATCH 04/13] Issue 147: the operator's tools still dial the bus that was removed MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Every tool call fails with 'AMQP not connected', on every node including the local one, because the tool surface still opens an AMQP connection and that transport was deleted at the cut-over. The mesh reports healthy throughout — what broke is the thing standing outside asking it questions, so nothing the mesh checks is about it. --- .../00-report.md | 58 +++++++++++++++++++ 1 file changed, 58 insertions(+) create mode 100644 04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md diff --git a/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md b/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md new file mode 100644 index 0000000..f4e2d57 --- /dev/null +++ b/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md @@ -0,0 +1,58 @@ +--- +status: open +opened: 2026-09-29 +located-in: [] +--- + +# 147 — the operator's tools still dial the bus that was removed + +## What was observed + +Every tool call an operator makes against the mesh fails, on every machine, with the same answer: + +``` +AMQP not connected — cannot reach hal/mesh@novox +AMQP not connected — cannot reach hal/mesh@shanks +``` + +The mesh moved to one bus and the previous transport was deleted +([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md), cut over +2026-09-28). The surface an operator drives the mesh through — the tool bridge their assistant +speaks to — still opens an AMQP connection, so it cannot reach anything. Not one node answers, +including the machine the operator is sitting at. + +**What that leaves.** The mesh itself is healthy: nodes are current, modules run, the bus carries +the mesh's own traffic. What is gone is the way a person asks it anything. Every question — what a +node runs, what is assigned, what a module's settings are — has to be asked by opening a shell on +the machine and running the control plane's binary inside its container, which is precisely the +path the tool surface exists to remove, and which nothing checks, records or permits. + +It also silently changes how work gets done: an assistant told to use the mesh's tools finds them +dead, falls back to `ssh` and `docker exec`, and the operator discovers the fallback rather than +the fault. + +## Why this is here and not a note in the knowledge base + +The rule is that everything happens over the mesh's bus. A surface that cannot reach that bus is +that rule enforced by nothing — and the mesh reports itself healthy throughout, because what broke +is not a module, a node or a provision but the thing standing outside asking them questions. + +The same cut-over removed the assistant's long-term memory (`recall`, `memorize`) for the same +reason, which is why lessons from the last two days were written into this repository by hand. + +## What would have prevented it + +- **The tool surface as a consumer of the bus like any other**, so moving the bus moves it — rather + than a separate bridge with its own connection settings that nothing resolves. +- **A check that a tool call reaches a node**, run where the mesh's other checks run. Every check + the mesh makes today is about the relationship between the mesh and a machine; none asks whether + a person can ask it anything ([issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md), + the same shape one level out). + +## Evidence to carry into diagnosis + +- The failure text names `hal/mesh@`, so the bridge is resolving a node and then dialling the + old transport. +- It fails identically for the local machine, which rules out reachability and points at the + transport alone. +- The mesh's own traffic over the new bus is unaffected: nodes report, declarations apply. -- 2.54.0 From 3c0f7082e64b7f52c4e482d908cd2f3686d5ac27 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 21:50:50 +0200 Subject: [PATCH 05/13] Issue 147: the tool surface is not the mesh's, it is the predecessor's MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Diagnosed from the configuration: the tool server is HAL's brain, a local process on the workstation with the predecessor's broker URL in the assistant's own config. No manifest, no assignment, no seat, no account. Nothing regressed — the mesh removed a transport this program still dials, and the program was never part of the mesh. The mesh has a tool model and nothing publishes an operator-facing surface onto it. --- .../00-report.md | 21 +++++++++++++++++-- 1 file changed, 19 insertions(+), 2 deletions(-) diff --git a/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md b/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md index f4e2d57..b056a21 100644 --- a/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md +++ b/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md @@ -1,7 +1,7 @@ --- -status: open +status: located opened: 2026-09-29 -located-in: [] +located-in: [nothing in the mesh — the surface in use is the predecessor's, installed on the workstation] --- # 147 — the operator's tools still dial the bus that was removed @@ -49,6 +49,23 @@ reason, which is why lessons from the last two days were written into this repos a person can ask it anything ([issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md), the same shape one level out). +## Diagnosed at once, because the answer was in the configuration + +**The surface is not the mesh's.** The tool server the operator's assistant speaks to is the +predecessor's brain, installed on the workstation and started as a local process, with the +predecessor's broker URL — `amqp://…@` — written into the assistant's own +configuration. Nothing about it is a module: it has no manifest, no assignment, no seat, no account +on the mesh's bus, and the mesh has never known it exists. + +So nothing regressed. The mesh removed a transport that this program still dials, and the program +was never part of the mesh to be moved. **The mesh has a tool model** — `tools` in a manifest, the +calls a seat accepts, the subjects a module answers on, and a control-plane command that asks one — +and **nothing publishes an operator-facing surface onto it.** What an operator uses is the thing +that came before, kept alive by a URL in a file. + +That is the issue, and it is larger than a broken connection: the way a person drives this mesh is +outside the mesh. + ## Evidence to carry into diagnosis - The failure text names `hal/mesh@`, so the bridge is resolving a node and then dialling the -- 2.54.0 From 14be8576f80a97b6aff9616e229799f8b67758a4 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 22:25:25 +0200 Subject: [PATCH 06/13] Grooming: five issues were fixed and never closed, and one is not MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 088, 089, 120, 128 and 130 each name a commit that is on main and cites them — the forge's address following a moved port, a route naming its endpoint, a provisioner asking the backend what is there, the hosts file written into a marked block, and undeclaring giving a unit back the state it was found in. Each says it was closed by reading commits rather than by a run, so nobody reads a green that was never measured. 129 stays located on purpose: ca-trust is merged and no machine holds it, so the symptom it opened on is still true everywhere. --- .../00-report.md | 12 +++++++++--- .../00-report.md | 12 +++++++++--- .../00-report.md | 10 ++++++++-- .../00-report.md | 9 ++++++++- .../01-diagnosis.md | 13 +++++++++++++ .../130-undeclaring-a-service-stops-it/00-report.md | 9 ++++++++- 6 files changed, 55 insertions(+), 10 deletions(-) diff --git a/04-ISSUES/088-the-forges-own-address-names-a-port-it-may-not-have/00-report.md b/04-ISSUES/088-the-forges-own-address-names-a-port-it-may-not-have/00-report.md index 4aa2eea..890b6c8 100644 --- a/04-ISSUES/088-the-forges-own-address-names-a-port-it-may-not-have/00-report.md +++ b/04-ISSUES/088-the-forges-own-address-names-a-port-it-may-not-have/00-report.md @@ -1,8 +1,8 @@ --- -status: open +status: resolved opened: 2026-09-22 -located-in: [] -fixed-by: +located-in: [mesh-catalog modules/gitea, mesh-controller internal/catalogue/declaration.go] +fixed-by: mesh-controller 7352c84, merged in #46 — a module is told its port in a container's environment too, as a file already was amended-design: --- @@ -39,3 +39,9 @@ the assignment happens to differ. - Should composition refuse an environment value that names a port the module does not fix, the way it refuses other claims a module cannot make? - Which other modules write their own address, with a port, into their environment? + +## Closed + +*2026-09-29, in a grooming pass rather than by whoever fixed it.* The forge's address follows a moved port the same way every other reader does. Found by +reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was +not re-verified on a machine, and this record says so rather than implying a run that did not happen. diff --git a/04-ISSUES/089-a-contributed-route-names-a-port-the-node-may-have-moved/00-report.md b/04-ISSUES/089-a-contributed-route-names-a-port-the-node-may-have-moved/00-report.md index 80df81c..038b028 100644 --- a/04-ISSUES/089-a-contributed-route-names-a-port-the-node-may-have-moved/00-report.md +++ b/04-ISSUES/089-a-contributed-route-names-a-port-the-node-may-have-moved/00-report.md @@ -1,8 +1,8 @@ --- -status: open +status: resolved opened: 2026-09-22 -located-in: [] -fixed-by: +located-in: [mesh-controller internal/catalogue/declaration.go, mesh-catalog (every routed module)] +fixed-by: mesh-controller bdf965d (a route names the endpoint it serves) with `portOfEndpoint` and `AtPublishedPort` — the contribution carries the endpoint's declared port and the machine-side redirection is applied to it like any other amended-design: --- @@ -45,3 +45,9 @@ precisely because the predecessor holds the usual one. keeping the mapping out of rendered configuration? - What should refuse a declaration whose contributed route names a port nothing on that node listens on? + +## Closed + +*2026-09-29, in a grooming pass rather than by whoever fixed it.* A route names an endpoint rather than a port, and the redirection that turns a declared port into the published one is applied to contributions too. Found by +reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was +not re-verified on a machine, and this record says so rather than implying a run that did not happen. diff --git a/04-ISSUES/120-a-provisioner-remembers-what-it-did-not-what-is/00-report.md b/04-ISSUES/120-a-provisioner-remembers-what-it-did-not-what-is/00-report.md index 9b697f3..078d27c 100644 --- a/04-ISSUES/120-a-provisioner-remembers-what-it-did-not-what-is/00-report.md +++ b/04-ISSUES/120-a-provisioner-remembers-what-it-did-not-what-is/00-report.md @@ -1,8 +1,8 @@ --- -status: located +status: resolved opened: 2026-09-26 located-in: [mesh-sdk src/provisioner, mesh-catalog modules/redis] -fixed-by: +fixed-by: mesh-catalog bbda88c, merged in #84 — every credential provider says whether it still holds a consumer amended-design: --- @@ -60,3 +60,9 @@ checks it after the first pass. instance and leaves the gap for the others. - Where does the record of what was applied live, if not in memory? ADR 0114, still proposed, puts rotation state with the vault. The same place may answer this. + +## Closed + +*2026-09-29, in a grooming pass rather than by whoever fixed it.* A provisioner asks the backend what is there rather than trusting what it remembers doing. Found by +reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was +not re-verified on a machine, and this record says so rather than implying a run that did not happen. diff --git a/04-ISSUES/128-the-hosts-file-is-written-whole/00-report.md b/04-ISSUES/128-the-hosts-file-is-written-whole/00-report.md index fa02a00..509ee3f 100644 --- a/04-ISSUES/128-the-hosts-file-is-written-whole/00-report.md +++ b/04-ISSUES/128-the-hosts-file-is-written-whole/00-report.md @@ -1,5 +1,6 @@ --- -status: located +status: resolved +fixed-by: mesh-host 1cb8953 and fdc768c — the mesh writes into a marked block of a text file instead of over it opened: 2026-09-26 located-in: [mesh-controller internal/catalogue/facts.go, mesh-host internal/apply] --- @@ -71,3 +72,9 @@ private network loses that name too. - The host's file resource supports `into: "json"` only; anything else is a whole write. - `node show ` on the adopted workstation: `holds file /etc/hosts mesh-wireguard.fact-node-names`, original kept. + +## Closed + +*2026-09-29, in a grooming pass rather than by whoever fixed it.* A shared hosts file keeps every line that is not the mesh's. Found by +reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was +not re-verified on a machine, and this record says so rather than implying a run that did not happen. diff --git a/04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/01-diagnosis.md b/04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/01-diagnosis.md index 0071b16..e262608 100644 --- a/04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/01-diagnosis.md +++ b/04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/01-diagnosis.md @@ -30,3 +30,16 @@ network, installs it as a trust anchor, refreshes the machine's bundles, and — unassigned stops its unit, and stopping the unit is what undoes it — takes both away again. The owner is therefore `mesh-catalog`, module `ca-trust`, and nothing in the control plane. + +## The module exists, and this stays open until a machine holds it + +*2026-09-29.* `ca-trust` is in the catalogue and merged +([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md)), and what it renders +is checked in the control plane's own suite: the script fetches from the authority it was bound to, +and the unit runs it both ways. + +**No machine has been assigned it, and nothing has verified a name because of it.** The bed written +for that cannot run ([issue 146](../146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md)), +and the live mesh has not been given the module. So the symptom this record opened on — every +internal name failing verification on every machine — is still true everywhere, and the record stays +`located` until it is not. Closing it on a module that exists would be closing it on an intention. diff --git a/04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md b/04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md index 57fb3de..b9ea3ec 100644 --- a/04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md +++ b/04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md @@ -1,5 +1,6 @@ --- -status: located +status: resolved +fixed-by: mesh-host 3112c88 — undeclaring gives a unit back the state it was found in, and removes only a process the mesh made opened: 2026-09-27 located-in: [mesh-host internal/apply/apply.go (remove)] amended-design: 02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md @@ -64,3 +65,9 @@ something to settle in passing. The unassign preview is partly answered — the host's plan names each unit it will stop — and the controller's side is left open. + +## Closed + +*2026-09-29, in a grooming pass rather than by whoever fixed it.* Undeclaring no longer stops a unit the mesh only reloaded or only kept running. Found by +reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was +not re-verified on a machine, and this record says so rather than implying a run that did not happen. -- 2.54.0 From 9a1dc4665cdf2bf7f259b92ec9ed0422421b9f44 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 22:26:13 +0200 Subject: [PATCH 07/13] Grooming: issue 006's knowledge base is the predecessor's, and is gone MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The record catches README claiming this repository is indexed into a knowledge base. Nothing indexed it, and since the cut-over there is nothing to index it into — the surface that answered is on the transport the mesh removed (issue 147). Noted where the record is, so the next reader does not go looking for a search that cannot exist. --- .../00-report.md | 14 ++++++++++++++ 1 file changed, 14 insertions(+) diff --git a/04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md b/04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md index 7c8280e..2c8c762 100644 --- a/04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md +++ b/04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md @@ -155,3 +155,17 @@ design document here, and get it back. That check fails today by design. **What stands until then** is the signpost, and the honest description of it: reachable, not surfacing. + +## Where this stands, 2026-09-29 + +*Added in a grooming pass.* The knowledge base this record is about is the **predecessor's**, and it +is no longer reachable from anything: the surface that answered `recall_search` speaks the transport +the mesh removed at the cut-over +([issue 147](../147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md)). + +So the sentence in `README.md` that this record catches — *these documents are still indexed into +the knowledge base* — is now wrong twice over: nothing indexed them, and there is nothing to index +them into. The record stays open, and its answer is no longer "index this repository somewhere"; it +is whatever the mesh grows as its own knowledge surface, if it grows one. Until then the honest fix +is the README, which should stop claiming a property nothing provides. + -- 2.54.0 From 4eb16f1028c01769df3341c249c4bb42aa4e065b Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 22:31:55 +0200 Subject: [PATCH 08/13] Three proposed records were already built; two are still yours to call MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 0037 (where a module lives), 0113 (the vault makes every secret) and 0115 (one assignment of a module per node) described arrangements the mesh has — a catalogue of other people's software plus a table filled by module add, a vault answering a secret requirement six modules make, and a rule the assignment table's primary key already enforces. Each is accepted against what was built, and says so in its own words. 0037's other half is not built: a manifest outside this catalogue has no check, which is issue 148. 0068 (the lab takes requests) and 0114 (a shared credential rotates over two credentials) stay proposed. Neither is built, and both are decisions rather than records of something that happened. --- 02-DECISIONS/0037-where-a-module-lives.md | 18 ++++++++- .../0113-the-vault-makes-every-secret.md | 10 ++++- ...115-one-assignment-of-a-module-per-node.md | 9 ++++- 02-DECISIONS/README.md | 6 +-- .../00-report.md | 37 +++++++++++++++++++ 5 files changed, 74 insertions(+), 6 deletions(-) create mode 100644 04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md diff --git a/02-DECISIONS/0037-where-a-module-lives.md b/02-DECISIONS/0037-where-a-module-lives.md index a697d12..c01e6c6 100644 --- a/02-DECISIONS/0037-where-a-module-lives.md +++ b/02-DECISIONS/0037-where-a-module-lives.md @@ -1,6 +1,6 @@ --- topic: building it -status: proposed +status: accepted date: 2026-09-01 deciders: jochen reconstructed: false @@ -101,3 +101,19 @@ the digest down after building. **Whether kind 4 deserves a module at all.** Thirty-five descriptions that say *install this and write these files* may be better as one module with settings than as thirty-five modules. Left open deliberately; it is a question about the shape of the catalogue, not about whether to have one. + +## Accepted, 2026-09-29, against what was built + +*Marked in a grooming pass: the mesh was built to this record and the record still said `proposed`.* + +The proposal is the arrangement that exists. `mesh-catalog` holds descriptions of software we did +not write and the programs that provision it, and holds neither the mesh's own components nor an +application's own module. The mesh's list of modules is a table in the control plane, filled by +`module add`, and every module records the source it came from with the commit it was read at. + +**One half is not built: `module check` as a command on the control plane's binary.** A manifest is +still validated by a test that reaches into the control plane's internals — which works for this +catalogue and gives nothing at all to somebody describing their own application in their own +repository, which this record says is the case that matters most. That is +[issue 148](../04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md). + diff --git a/02-DECISIONS/0113-the-vault-makes-every-secret.md b/02-DECISIONS/0113-the-vault-makes-every-secret.md index a8ae287..b4b61fa 100644 --- a/02-DECISIONS/0113-the-vault-makes-every-secret.md +++ b/02-DECISIONS/0113-the-vault-makes-every-secret.md @@ -1,6 +1,6 @@ --- topic: what runs on it -status: proposed +status: accepted date: 2026-09-25 deciders: jochen reconstructed: false @@ -275,3 +275,11 @@ On acceptance, each of these is superseded or amended by this record, not edited everything a module needs is a requirement - [Issue 095](../04-ISSUES/095-a-module-assigned-after-genesis-has-no-broker-account/00-report.md), [issue 103](../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md): what fails today + +## Accepted, 2026-09-29, against what was built + +*Marked in a grooming pass.* The vault is a module providing `secret` at mesh scope, and six +modules in the catalogue require it — so a shared secret is a requirement answered by the vault, +which is what this record asks for. Private keys are still made where they are used and never +travel, which is the other half and was never in question. + diff --git a/02-DECISIONS/0115-one-assignment-of-a-module-per-node.md b/02-DECISIONS/0115-one-assignment-of-a-module-per-node.md index 341a3ff..5936529 100644 --- a/02-DECISIONS/0115-one-assignment-of-a-module-per-node.md +++ b/02-DECISIONS/0115-one-assignment-of-a-module-per-node.md @@ -1,6 +1,6 @@ --- topic: what runs on it -status: proposed +status: accepted date: 2026-09-26 deciders: jochen extends: 0112-a-module-definition-names-no-node-mesh-or-path.md @@ -47,3 +47,10 @@ other boundary already is: the module name. node runs one of each (ADR 0115)" — instead of failing on whichever name collides first. - Multi-tenant asks are answered in the catalogue (a second module definition), not in the control plane. + +## Accepted, 2026-09-29, against what was built + +*Marked in a grooming pass.* The rule is enforced where it cannot be forgotten: `assignment`'s +primary key is `(node, module)`, so a second assignment of one module to one machine is not a thing +the mesh can hold. The record read `proposed` while the schema had already settled it. + diff --git a/02-DECISIONS/README.md b/02-DECISIONS/README.md index 3e1cf39..cbf086a 100644 --- a/02-DECISIONS/README.md +++ b/02-DECISIONS/README.md @@ -207,9 +207,9 @@ python3 00-META/checks/index.py fail if stale - **0099** — [A step that runs once names what it reads, and runs again when it changed](0099-a-step-that-runs-once-names-what-it-reads.md) - **0110** — [A seat is held by one assignment, from a closed set, and it may deliver a provision](0110-a-seat-is-a-module-assignment-from-a-closed-set.md) - **0112** — [A module definition names no node, no mesh and no path: everything it needs is a requirement the mesh resolves](0112-a-module-definition-names-no-node-mesh-or-path.md) -- **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md) *(proposed)* +- **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md) - **0114** — [A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached](0114-a-shared-credential-rotates-over-two-credentials.md) *(proposed)* -- **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md) *(proposed)* +- **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md) - **0117** — [A machine's uplink is a seat: the mesh configures the manager, never the link](0117-a-machines-uplink-is-a-seat.md) - **0118** — [Undeclaring removes what the mesh made, and gives a unit back the state it was found in](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md) - **0120** — [A roster fact carries its format as a template: the mesh owns the data, the module owns the format](0120-a-roster-fact-carries-its-format-as-a-template.md) @@ -237,7 +237,7 @@ python3 00-META/checks/index.py fail if stale - **0014** — [No workspace — each module is a standalone package consuming published dependencies](0014-no-npm-workspace.md) - **0015** — [Applications live in their own repository; the monorepo is for the mesh](0015-applications-live-in-their-own-repository.md) - **0016** — [The lab](0016-the-lab.md) -- **0037** — [Where a module lives](0037-where-a-module-lives.md) *(proposed)* +- **0037** — [Where a module lives](0037-where-a-module-lives.md) - **0039** — [What the SDK holds, and what it refuses](0039-what-the-sdk-holds-and-refuses.md) - **0068** — [The lab takes requests, one at a time, and runs each from its own copy](0068-the-lab-takes-requests.md) *(proposed)* - **0069** — [A module is a repository and a path within it](0069-a-module-is-a-repository-and-a-path.md) diff --git a/04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md b/04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md new file mode 100644 index 0000000..1bd3f00 --- /dev/null +++ b/04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md @@ -0,0 +1,37 @@ +--- +status: open +opened: 2026-09-29 +located-in: [mesh-controller cmd/mesh-controller] +--- + +# 148 — a manifest outside this catalogue has no check + +## What was observed + +A module's manifest is validated by a **test** — `internal/catalogue`'s suite parses every manifest +in the catalogue checkout beside it and fails on one it cannot resolve. That works, and it is how +several real faults were caught before a machine saw them. + +It is available to exactly one repository: this one. Somebody describing their own application in +their own repository — the case +[ADR 0037](../../02-DECISIONS/0037-where-a-module-lives.md) calls *the case that matters most* — +has no check at all. They write a manifest, register it with a running mesh, and find out whether +it is valid when the mesh refuses it, or later, when a machine applies something that resolved and +should not have. + +The same record asks for the answer: **a `module check` command on the control plane's binary**, so +a manifest is checked by the tool rather than by a test that imports the tool's internals. + +## What would have prevented it + +Nothing prevents this; it was noticed and left. ADR 0037 named it on 2026-09-01 and the record sat +`proposed` until 2026-09-29, so the missing half was never anybody's task. + +## Evidence to carry into diagnosis + +- `mesh-controller/internal/catalogue` — `ParseManifest` and `CatalogueProblems` are the check, and + both are internal. +- The catalogue-wide test is `TestEveryCatalogueManifestDeclaresWhatItMounts` and its siblings; they + take a path from `MESH_CATALOG`, so the mechanism is already path-driven and not repository-bound. +- `mesh-controller module add` refuses a bad manifest at registration, which is the same check far + too late: by then it is in a running mesh's records. -- 2.54.0 From 96bdffa9bc481daffd94bb9e7d1b0789a70dfda0 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 22:33:40 +0200 Subject: [PATCH 09/13] Two records were numbered 127; the second becomes 149, and is resolved A number identifies a record, and two were given 127 on 2026-09-27. The one three documents and three source files cite by number keeps it; the other becomes 149, says so in its own heading, and its two inbound references are repointed. It is also resolved: an empty declaration is sent carrying owns_nothing rather than skipped, and the host refuses an empty body that does not carry it, so emptiness cannot be read as a truncated declaration. --- ...s-a-unit-back-the-state-it-was-found-in.md | 2 +- .../00-report.md | 2 +- .../00-report.md | 25 ++++++++++++++++--- 3 files changed, 24 insertions(+), 5 deletions(-) rename 04-ISSUES/{127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent => 149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent}/00-report.md (57%) diff --git a/02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md b/02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md index 317ec90..c5958f1 100644 --- a/02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md +++ b/02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md @@ -12,7 +12,7 @@ extends: 0102-the-mesh-writes-into-a-shared-file-never-over-it.md ## Context When a resource stops being declared — its module unassigned, the node sent a -deliberately-empty declaration ([issue 127](../04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)), +deliberately-empty declaration ([issue 149](../04-ISSUES/149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)), or a new catalogue version renaming its id — the host undoes it. The host's own code states the rule it means to follow: **it removes what it made and leaves what it merely configured.** For almost every resource it does exactly that: diff --git a/04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md b/04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md index b9ea3ec..431a3f8 100644 --- a/04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md +++ b/04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md @@ -16,7 +16,7 @@ found that the host's `remove` path stops every `service` resource that is no lo delete". `store.Orphans` matches by id alone. So any of these stops the unit: - the module is unassigned — by mistake, or to switch it for another; -- the node is sent a deliberately-empty declaration ([issue 127](../127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)); +- the node is sent a deliberately-empty declaration ([issue 149](../149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)); - a later catalogue version renames the resource's `id`. That is right for a service the mesh brought into being. It is wrong for a unit the mesh diff --git a/04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md b/04-ISSUES/149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md similarity index 57% rename from 04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md rename to 04-ISSUES/149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md index 9a57cb6..b5d797b 100644 --- a/04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md +++ b/04-ISSUES/149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md @@ -1,10 +1,15 @@ --- -status: located +status: resolved opened: 2026-09-27 -located-in: [mesh-controller cmd/mesh-controller/push.go] +located-in: [mesh-controller cmd/mesh-controller/push.go, mesh-controller cmd/mesh-controller/sendable.go] +fixed-by: mesh-controller sendable.go and push.go — an empty declaration is sent carrying `owns_nothing`, and the host refuses an empty body that does not carry it --- -# A declaration that shrinks to empty is skipped, so the node keeps what it should drop +# 149 — a declaration that shrinks to empty is skipped, so the node keeps what it should drop + +*Opened as 127 and renumbered on 2026-09-29: two records were given that number on the same day.* +*The other kept it, because three documents and three source files cite it by number and nothing +cited this one but a decision and a sibling issue, both corrected with this move.* ## What was observed @@ -37,3 +42,17 @@ mean "own nothing", which the host already applies correctly when it receives on On ace, one command drops it permanently (the corrected controller never re-composes it): `sudo ufw delete allow 5671`. At ace's converge it would clear on its own. + +## Closed + +*2026-09-29, in a grooming pass.* Both halves are on `main` and both name this issue. + +- The control plane **sends** it: a declaration that composes to no resources goes out with + `owns_nothing`, and `push` says *sent, not skipped*. +- The host **refuses an empty body that does not carry it**, so a truncated or mis-composed + declaration can never be read as "own nothing" — which is the failure the fix had to avoid while + making the empty case expressible. + +Closed by reading the code rather than by watching a machine let go of a stray resource; the record +says so rather than implying a run. + -- 2.54.0 From 2b5119ecd2dc29070f1fa9c2f5211cdb09b05a34 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 22:38:11 +0200 Subject: [PATCH 10/13] Issue 114 is answered: the controller is a process, by ADR 0142 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit It asked a question rather than reporting a defect, and the question was taken five days later — the mesh's own components are binaries on the machine, and third-party software stays a container because an image is the right way to carry somebody else's build. The delivery of them is issue 142. --- .../00-report.md | 22 +++++++++++++++++-- 1 file changed, 20 insertions(+), 2 deletions(-) diff --git a/04-ISSUES/114-should-the-controller-be-a-container-or-a-process/00-report.md b/04-ISSUES/114-should-the-controller-be-a-container-or-a-process/00-report.md index 6b03b96..194fbd7 100644 --- a/04-ISSUES/114-should-the-controller-be-a-container-or-a-process/00-report.md +++ b/04-ISSUES/114-should-the-controller-be-a-container-or-a-process/00-report.md @@ -1,8 +1,8 @@ --- -status: open +status: resolved opened: 2026-09-24 located-in: [mesh-controller module.json, mesh-host internal/apply] -fixed-by: +fixed-by: ADR 0142 — the mesh's own components are delivered as binaries on the machine, so the controller is a process; the delivery itself is issue 142 amended-design: --- @@ -84,3 +84,21 @@ restart and run-to-completion semantics — so this would not need host-side wor The two do not collapse into one. The controller is not a code-carrying sidecar, and `network: host` is what makes the asymmetry visible here and nowhere else. + +## Answered + +*2026-09-29, in a grooming pass.* This asked a question rather than reporting a defect, and the +question was taken: [ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md) +decides that the mesh's own components — the host, the controller, the catalogue, the builder, the +vault — are **binaries on the machine**, delivered by the mechanism +[ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) built, and that +third-party software (the store, the registry, the broker) stays a container because an image is the +right way to carry somebody else's build. + +So the operating experience this record was written from — every mutating command reached through +`docker exec mesh-controller` — is answered, and answered against the container. + +**The delivery is a separate matter and is not this record's.** Step 1 of it is built and no +component travels yet; that is +[issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md). + -- 2.54.0 From 72eaf52867b34ebe100d83a48c0485bce6a4b792 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 23:13:54 +0200 Subject: [PATCH 11/13] Issue 152: a node whose plan will not compose withdraws its names from every machine ace numbered its two records 147 and 148, which this repository already uses; they become 150 and 151, as 127 became 149. The new record is why the control node could not stop applying: the roster alternates between two values because routeNamesInTheMesh swallows a per-node plan failure, and the roster is part of every container's identity. The loop closes through the control plane's own store, which each pass replaces. --- .../00-report.md | 2 +- .../00-report.md | 4 +- .../00-report.md | 100 ++++++++++++++++++ 3 files changed, 103 insertions(+), 3 deletions(-) rename 04-ISSUES/{147-a-route-is-contributed-before-its-module-is-taken => 150-a-route-is-contributed-before-its-module-is-taken}/00-report.md (97%) rename 04-ISSUES/{148-a-new-name-recreates-every-container-in-the-mesh => 151-a-new-name-recreates-every-container-in-the-mesh}/00-report.md (96%) create mode 100644 04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md diff --git a/04-ISSUES/147-a-route-is-contributed-before-its-module-is-taken/00-report.md b/04-ISSUES/150-a-route-is-contributed-before-its-module-is-taken/00-report.md similarity index 97% rename from 04-ISSUES/147-a-route-is-contributed-before-its-module-is-taken/00-report.md rename to 04-ISSUES/150-a-route-is-contributed-before-its-module-is-taken/00-report.md index f9afe62..77220e6 100644 --- a/04-ISSUES/147-a-route-is-contributed-before-its-module-is-taken/00-report.md +++ b/04-ISSUES/150-a-route-is-contributed-before-its-module-is-taken/00-report.md @@ -6,7 +6,7 @@ located-in: fixed-by: --- -# 147 — A route is contributed before its module is taken +# 150 — A route is contributed before its module is taken ## What was observed diff --git a/04-ISSUES/148-a-new-name-recreates-every-container-in-the-mesh/00-report.md b/04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md similarity index 96% rename from 04-ISSUES/148-a-new-name-recreates-every-container-in-the-mesh/00-report.md rename to 04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md index 8185c6a..bf44ae6 100644 --- a/04-ISSUES/148-a-new-name-recreates-every-container-in-the-mesh/00-report.md +++ b/04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md @@ -7,13 +7,13 @@ located-in: fixed-by: --- -# 148 — A new name recreates every container in the mesh +# 151 — A new name recreates every container in the mesh ## What was observed Migrating one small module on ace (searxng) took four routine controller actions: `node public-domain ace zurag.be`, `assign ace searxng` + push, `unassign ace searxng` + push (a rollback, see -[issue 147](../147-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take +[issue 150](../150-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take + push again. Each push to ace also pushed novox ("this push left g14, novox, shanks behind … sending it too"). novox's host then **replaced every container it runs, twice**: diff --git a/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md new file mode 100644 index 0000000..f65fa1a --- /dev/null +++ b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md @@ -0,0 +1,100 @@ +--- +status: located +opened: 2026-09-29 +located-in: + - mesh-controller cmd/mesh-controller/plan.go (routeNamesInTheMesh skips a node whose plan will not compose) +fixed-by: +amended-design: +--- + +# 152 — A node whose plan will not compose silently removes its names from every machine + +## What was observed + +For at least seventeen minutes after the last operator action, the control-node's host applied all +327 of its resources every ~6.5 minutes without pause, replacing every container on the machine each +time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail, +and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine +sat near 8. Nothing was converging: each pass ended and the next began four seconds later. + +The two machines carrying no containers were not churning. They were only knocked off the bus each +time the control node re-created it, reconnected, re-heard the same declaration and applied it again +as a no-op. + +## Why: the roster alternates between two values, and it is part of every container + +Two consecutive declarations were compared by reading the `--add-host` entries of four containers +the moment each pass created them: + +``` +23:02 ace.internal drive.novox.be g14.internal keycloak.novox.be novox.internal + office.novox.be portainer.novox.be shanks.internal umami.novox.be (9 names) +23:07 … the same nine, and searxng.zurag.be (10 names) +``` + +One routed name — belonging to a module on another machine entirely — leaves the roster and comes +back. Because the roster is part of every container's spec digest ([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)), +each flip is a different identity for every container on the machine, and a running container cannot +have its hosts changed. So every flip replaces all of them. + +**What makes it flip is a swallowed error.** `routeNamesInTheMesh` composes every node's plan to +find the names it serves, and when one will not compose it moves on: + +``` +plan, settings, err := planFor(ctx, open, n.Name) +if err != nil { + continue +} +``` + +A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does +not know", but "the mesh states these names do not exist", to every machine at once. + +## Why it cannot recover on its own + +The loop closes through the control plane's own database: + +1. An apply replaces `mesh-store` — the store the control plane reads — by removing the container, + so postgres comes back through crash recovery. +2. While it recovers it refuses connections: `FATAL: the database system is not yet accepting + connections / Consistent recovery state has not been yet reached` (observed, 21:08:39 UTC, every + pass). +3. `planFor` for the other machine fails against that store. `routeNamesInTheMesh` swallows it and + drops its routed name. +4. The roster changed, so all 327 resources differ, so all are replaced — including `mesh-store`, + and including the bus, which is why the host also cannot report: `applied, and could not tell the + mesh: reporting: nats: connection closed`. +5. Back to 1. + +It is stable in its instability: every pass destroys the evidence the next pass needs to decide it +has nothing to do. Nothing outside the machine has to be wrong for this to continue, and nothing +about it stops. + +## What it is not + +- Not the operator's four actions on the other machine. Those explain the first passes + ([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)); they were + finished seventeen minutes and three full passes before these measurements. +- Not a file that keeps changing. `/etc/hosts`, the bus's account list and the vault's export were + hashed across passes and are byte-identical, and the directory reported `mode 755 to 700` every + pass while already being `700`. Those resources are **misreported as changed** and are worth their + own question, but they are not what moves a container's identity. +- Not the lost report alone. A report that cannot be delivered explains a re-apply; it does not + explain a re-apply that finds 327 differences. + +## Why it matters beyond this outage + +The same `continue` makes every routed name in the mesh conditional on every node's plan composing at +the moment any machine is pushed to. One unreachable or half-migrated machine is enough to withdraw +its names from everywhere — and the withdrawal is indistinguishable, on the receiving machine, from +the operator having removed them. + +The codebase already states the rule this breaks, forty lines away, about the same kind of lookup: + +> A lookup failure is an error, never "not found": collapsing the two composed a declaration without +> the trust whenever the inventory hiccuped, delivered by a push that reported success. + +## How the fix is checked + +A test that composes the roster with one node's plan failing, and requires the compose to fail rather +than return a roster missing that node's names. -- 2.54.0 From 3c2b4fc6b6b95bf3491a6e0b129b0a4a98104833 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 23:26:20 +0200 Subject: [PATCH 12/13] Issue 152 is fixed: a lookup failure is no longer an absence The three gatherers pass over a node whose set does not compose, and raise anything else. 151 stays open: this removes the false reasons a roster changes, not the fact that a real change still replaces every container. --- .../00-report.md | 24 +++++++++++++++---- 1 file changed, 19 insertions(+), 5 deletions(-) diff --git a/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md index f65fa1a..ccef9d9 100644 --- a/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md +++ b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md @@ -1,9 +1,9 @@ --- -status: located +status: resolved opened: 2026-09-29 located-in: - mesh-controller cmd/mesh-controller/plan.go (routeNamesInTheMesh skips a node whose plan will not compose) -fixed-by: +fixed-by: mesh-control fix/152-a-lookup-failure-is-not-an-absence amended-design: --- @@ -94,7 +94,21 @@ The codebase already states the rule this breaks, forty lines away, about the sa > A lookup failure is an error, never "not found": collapsing the two composed a declaration without > the trust whenever the inventory hiccuped, delivered by a push that reported success. -## How the fix is checked +## How it was fixed, and how the fix is checked -A test that composes the roster with one node's plan failing, and requires the compose to fail rather -than return a roster missing that node's names. +`planFor` now marks the two failures that really are a statement about the node — its set not +composing, and a setting that reaches nothing — and the three gatherers pass over those and only +those. Every other failure is raised, naming the machine and the read. + +Five tests hold it: a set that cannot compose is marked as the node's own; a store that cannot be +read is *not*; one incoherent node still does not cost the rest their names; a roster is never +returned beside an error; and the raised failure names what could not be read. + +The two sibling gatherers were audited and fixed the same way — the grant composer, which would have +withheld a consumer's credential, and the private-network membership, which would have taken a +machine off the overlay. Three other `planFor` callers were audited and left alone: they refuse or +report rather than silently withdraw, which is the safe direction. + +**This does not close [issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md).** +A roster that changes for a real reason still replaces every container in the mesh. This removes the +false reasons; whether the roster belongs in a container's identity at all is that record's question. -- 2.54.0 From 18f37c25b2ef0cba6d56d51ebe697db6f9aef2d3 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 23:27:51 +0200 Subject: [PATCH 13/13] Issue 152: the loop is metastable, and it cleared at 23:11 It ran 22:46-23:11, five full replacements, and stopped when a pass happened to read the store during a window it was up. The record said nothing about it stops; that was wrong. Exiting by luck is the finding, not a mitigation. --- .../00-report.md | 24 +++++++++++++------ 1 file changed, 17 insertions(+), 7 deletions(-) diff --git a/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md index ccef9d9..6dcd31f 100644 --- a/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md +++ b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md @@ -11,9 +11,9 @@ amended-design: ## What was observed -For at least seventeen minutes after the last operator action, the control-node's host applied all -327 of its resources every ~6.5 minutes without pause, replacing every container on the machine each -time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail, +For twenty-five minutes, and for seventeen of them after the last operator action, the control-node's +host applied all 327 of its resources every ~6.5 minutes without pause, replacing every container on +the machine each time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail, and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine sat near 8. Nothing was converging: each pass ended and the next began four seconds later. @@ -50,7 +50,7 @@ if err != nil { A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does not know", but "the mesh states these names do not exist", to every machine at once. -## Why it cannot recover on its own +## Why it sustains itself The loop closes through the control plane's own database: @@ -66,9 +66,19 @@ The loop closes through the control plane's own database: mesh: reporting: nats: connection closed`. 5. Back to 1. -It is stable in its instability: every pass destroys the evidence the next pass needs to decide it -has nothing to do. Nothing outside the machine has to be wrong for this to continue, and nothing -about it stops. +Every pass destroys the evidence the next pass needs to decide it has nothing to do, and nothing +outside the machine has to be wrong for it to continue. + +**It is metastable, not permanent.** It ran from 22:46 to 23:11 — five full replacements of every +container on the machine — and then stopped on its own, when one pass happened to read the store +during a window it was up, composed the same roster twice running, and found nothing to do. Load fell +from 7.8 to 1.7 and the machine returned to its five-minute idle tick. + +That it ends by luck is the point, not a mitigation. The exit condition is a race the mesh does not +control, the operator cannot see, and nothing reports; the same four actions on a slower machine, or +a larger store, would not have found it. An outage that clears itself after twenty-five minutes and +five restarts of the forge, the directory and mail is not a smaller fault than one that does not — it +is the same fault, harder to catch. ## What it is not -- 2.54.0