From 4eb16f1028c01769df3341c249c4bb42aa4e065b Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 22:31:55 +0200 Subject: [PATCH 1/6] Three proposed records were already built; two are still yours to call MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 0037 (where a module lives), 0113 (the vault makes every secret) and 0115 (one assignment of a module per node) described arrangements the mesh has — a catalogue of other people's software plus a table filled by module add, a vault answering a secret requirement six modules make, and a rule the assignment table's primary key already enforces. Each is accepted against what was built, and says so in its own words. 0037's other half is not built: a manifest outside this catalogue has no check, which is issue 148. 0068 (the lab takes requests) and 0114 (a shared credential rotates over two credentials) stay proposed. Neither is built, and both are decisions rather than records of something that happened. --- 02-DECISIONS/0037-where-a-module-lives.md | 18 ++++++++- .../0113-the-vault-makes-every-secret.md | 10 ++++- ...115-one-assignment-of-a-module-per-node.md | 9 ++++- 02-DECISIONS/README.md | 6 +-- .../00-report.md | 37 +++++++++++++++++++ 5 files changed, 74 insertions(+), 6 deletions(-) create mode 100644 04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md diff --git a/02-DECISIONS/0037-where-a-module-lives.md b/02-DECISIONS/0037-where-a-module-lives.md index a697d12..c01e6c6 100644 --- a/02-DECISIONS/0037-where-a-module-lives.md +++ b/02-DECISIONS/0037-where-a-module-lives.md @@ -1,6 +1,6 @@ --- topic: building it -status: proposed +status: accepted date: 2026-09-01 deciders: jochen reconstructed: false @@ -101,3 +101,19 @@ the digest down after building. **Whether kind 4 deserves a module at all.** Thirty-five descriptions that say *install this and write these files* may be better as one module with settings than as thirty-five modules. Left open deliberately; it is a question about the shape of the catalogue, not about whether to have one. + +## Accepted, 2026-09-29, against what was built + +*Marked in a grooming pass: the mesh was built to this record and the record still said `proposed`.* + +The proposal is the arrangement that exists. `mesh-catalog` holds descriptions of software we did +not write and the programs that provision it, and holds neither the mesh's own components nor an +application's own module. The mesh's list of modules is a table in the control plane, filled by +`module add`, and every module records the source it came from with the commit it was read at. + +**One half is not built: `module check` as a command on the control plane's binary.** A manifest is +still validated by a test that reaches into the control plane's internals — which works for this +catalogue and gives nothing at all to somebody describing their own application in their own +repository, which this record says is the case that matters most. That is +[issue 148](../04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md). + diff --git a/02-DECISIONS/0113-the-vault-makes-every-secret.md b/02-DECISIONS/0113-the-vault-makes-every-secret.md index a8ae287..b4b61fa 100644 --- a/02-DECISIONS/0113-the-vault-makes-every-secret.md +++ b/02-DECISIONS/0113-the-vault-makes-every-secret.md @@ -1,6 +1,6 @@ --- topic: what runs on it -status: proposed +status: accepted date: 2026-09-25 deciders: jochen reconstructed: false @@ -275,3 +275,11 @@ On acceptance, each of these is superseded or amended by this record, not edited everything a module needs is a requirement - [Issue 095](../04-ISSUES/095-a-module-assigned-after-genesis-has-no-broker-account/00-report.md), [issue 103](../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md): what fails today + +## Accepted, 2026-09-29, against what was built + +*Marked in a grooming pass.* The vault is a module providing `secret` at mesh scope, and six +modules in the catalogue require it — so a shared secret is a requirement answered by the vault, +which is what this record asks for. Private keys are still made where they are used and never +travel, which is the other half and was never in question. + diff --git a/02-DECISIONS/0115-one-assignment-of-a-module-per-node.md b/02-DECISIONS/0115-one-assignment-of-a-module-per-node.md index 341a3ff..5936529 100644 --- a/02-DECISIONS/0115-one-assignment-of-a-module-per-node.md +++ b/02-DECISIONS/0115-one-assignment-of-a-module-per-node.md @@ -1,6 +1,6 @@ --- topic: what runs on it -status: proposed +status: accepted date: 2026-09-26 deciders: jochen extends: 0112-a-module-definition-names-no-node-mesh-or-path.md @@ -47,3 +47,10 @@ other boundary already is: the module name. node runs one of each (ADR 0115)" — instead of failing on whichever name collides first. - Multi-tenant asks are answered in the catalogue (a second module definition), not in the control plane. + +## Accepted, 2026-09-29, against what was built + +*Marked in a grooming pass.* The rule is enforced where it cannot be forgotten: `assignment`'s +primary key is `(node, module)`, so a second assignment of one module to one machine is not a thing +the mesh can hold. The record read `proposed` while the schema had already settled it. + diff --git a/02-DECISIONS/README.md b/02-DECISIONS/README.md index 3e1cf39..cbf086a 100644 --- a/02-DECISIONS/README.md +++ b/02-DECISIONS/README.md @@ -207,9 +207,9 @@ python3 00-META/checks/index.py fail if stale - **0099** — [A step that runs once names what it reads, and runs again when it changed](0099-a-step-that-runs-once-names-what-it-reads.md) - **0110** — [A seat is held by one assignment, from a closed set, and it may deliver a provision](0110-a-seat-is-a-module-assignment-from-a-closed-set.md) - **0112** — [A module definition names no node, no mesh and no path: everything it needs is a requirement the mesh resolves](0112-a-module-definition-names-no-node-mesh-or-path.md) -- **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md) *(proposed)* +- **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md) - **0114** — [A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached](0114-a-shared-credential-rotates-over-two-credentials.md) *(proposed)* -- **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md) *(proposed)* +- **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md) - **0117** — [A machine's uplink is a seat: the mesh configures the manager, never the link](0117-a-machines-uplink-is-a-seat.md) - **0118** — [Undeclaring removes what the mesh made, and gives a unit back the state it was found in](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md) - **0120** — [A roster fact carries its format as a template: the mesh owns the data, the module owns the format](0120-a-roster-fact-carries-its-format-as-a-template.md) @@ -237,7 +237,7 @@ python3 00-META/checks/index.py fail if stale - **0014** — [No workspace — each module is a standalone package consuming published dependencies](0014-no-npm-workspace.md) - **0015** — [Applications live in their own repository; the monorepo is for the mesh](0015-applications-live-in-their-own-repository.md) - **0016** — [The lab](0016-the-lab.md) -- **0037** — [Where a module lives](0037-where-a-module-lives.md) *(proposed)* +- **0037** — [Where a module lives](0037-where-a-module-lives.md) - **0039** — [What the SDK holds, and what it refuses](0039-what-the-sdk-holds-and-refuses.md) - **0068** — [The lab takes requests, one at a time, and runs each from its own copy](0068-the-lab-takes-requests.md) *(proposed)* - **0069** — [A module is a repository and a path within it](0069-a-module-is-a-repository-and-a-path.md) diff --git a/04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md b/04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md new file mode 100644 index 0000000..1bd3f00 --- /dev/null +++ b/04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md @@ -0,0 +1,37 @@ +--- +status: open +opened: 2026-09-29 +located-in: [mesh-controller cmd/mesh-controller] +--- + +# 148 — a manifest outside this catalogue has no check + +## What was observed + +A module's manifest is validated by a **test** — `internal/catalogue`'s suite parses every manifest +in the catalogue checkout beside it and fails on one it cannot resolve. That works, and it is how +several real faults were caught before a machine saw them. + +It is available to exactly one repository: this one. Somebody describing their own application in +their own repository — the case +[ADR 0037](../../02-DECISIONS/0037-where-a-module-lives.md) calls *the case that matters most* — +has no check at all. They write a manifest, register it with a running mesh, and find out whether +it is valid when the mesh refuses it, or later, when a machine applies something that resolved and +should not have. + +The same record asks for the answer: **a `module check` command on the control plane's binary**, so +a manifest is checked by the tool rather than by a test that imports the tool's internals. + +## What would have prevented it + +Nothing prevents this; it was noticed and left. ADR 0037 named it on 2026-09-01 and the record sat +`proposed` until 2026-09-29, so the missing half was never anybody's task. + +## Evidence to carry into diagnosis + +- `mesh-controller/internal/catalogue` — `ParseManifest` and `CatalogueProblems` are the check, and + both are internal. +- The catalogue-wide test is `TestEveryCatalogueManifestDeclaresWhatItMounts` and its siblings; they + take a path from `MESH_CATALOG`, so the mechanism is already path-driven and not repository-bound. +- `mesh-controller module add` refuses a bad manifest at registration, which is the same check far + too late: by then it is in a running mesh's records. From 96bdffa9bc481daffd94bb9e7d1b0789a70dfda0 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 22:33:40 +0200 Subject: [PATCH 2/6] Two records were numbered 127; the second becomes 149, and is resolved A number identifies a record, and two were given 127 on 2026-09-27. The one three documents and three source files cite by number keeps it; the other becomes 149, says so in its own heading, and its two inbound references are repointed. It is also resolved: an empty declaration is sent carrying owns_nothing rather than skipped, and the host refuses an empty body that does not carry it, so emptiness cannot be read as a truncated declaration. --- ...s-a-unit-back-the-state-it-was-found-in.md | 2 +- .../00-report.md | 2 +- .../00-report.md | 25 ++++++++++++++++--- 3 files changed, 24 insertions(+), 5 deletions(-) rename 04-ISSUES/{127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent => 149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent}/00-report.md (57%) diff --git a/02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md b/02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md index 317ec90..c5958f1 100644 --- a/02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md +++ b/02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md @@ -12,7 +12,7 @@ extends: 0102-the-mesh-writes-into-a-shared-file-never-over-it.md ## Context When a resource stops being declared — its module unassigned, the node sent a -deliberately-empty declaration ([issue 127](../04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)), +deliberately-empty declaration ([issue 149](../04-ISSUES/149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)), or a new catalogue version renaming its id — the host undoes it. The host's own code states the rule it means to follow: **it removes what it made and leaves what it merely configured.** For almost every resource it does exactly that: diff --git a/04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md b/04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md index b9ea3ec..431a3f8 100644 --- a/04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md +++ b/04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md @@ -16,7 +16,7 @@ found that the host's `remove` path stops every `service` resource that is no lo delete". `store.Orphans` matches by id alone. So any of these stops the unit: - the module is unassigned — by mistake, or to switch it for another; -- the node is sent a deliberately-empty declaration ([issue 127](../127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)); +- the node is sent a deliberately-empty declaration ([issue 149](../149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)); - a later catalogue version renames the resource's `id`. That is right for a service the mesh brought into being. It is wrong for a unit the mesh diff --git a/04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md b/04-ISSUES/149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md similarity index 57% rename from 04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md rename to 04-ISSUES/149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md index 9a57cb6..b5d797b 100644 --- a/04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md +++ b/04-ISSUES/149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md @@ -1,10 +1,15 @@ --- -status: located +status: resolved opened: 2026-09-27 -located-in: [mesh-controller cmd/mesh-controller/push.go] +located-in: [mesh-controller cmd/mesh-controller/push.go, mesh-controller cmd/mesh-controller/sendable.go] +fixed-by: mesh-controller sendable.go and push.go — an empty declaration is sent carrying `owns_nothing`, and the host refuses an empty body that does not carry it --- -# A declaration that shrinks to empty is skipped, so the node keeps what it should drop +# 149 — a declaration that shrinks to empty is skipped, so the node keeps what it should drop + +*Opened as 127 and renumbered on 2026-09-29: two records were given that number on the same day.* +*The other kept it, because three documents and three source files cite it by number and nothing +cited this one but a decision and a sibling issue, both corrected with this move.* ## What was observed @@ -37,3 +42,17 @@ mean "own nothing", which the host already applies correctly when it receives on On ace, one command drops it permanently (the corrected controller never re-composes it): `sudo ufw delete allow 5671`. At ace's converge it would clear on its own. + +## Closed + +*2026-09-29, in a grooming pass.* Both halves are on `main` and both name this issue. + +- The control plane **sends** it: a declaration that composes to no resources goes out with + `owns_nothing`, and `push` says *sent, not skipped*. +- The host **refuses an empty body that does not carry it**, so a truncated or mis-composed + declaration can never be read as "own nothing" — which is the failure the fix had to avoid while + making the empty case expressible. + +Closed by reading the code rather than by watching a machine let go of a stray resource; the record +says so rather than implying a run. + From 2b5119ecd2dc29070f1fa9c2f5211cdb09b05a34 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 22:38:11 +0200 Subject: [PATCH 3/6] Issue 114 is answered: the controller is a process, by ADR 0142 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit It asked a question rather than reporting a defect, and the question was taken five days later — the mesh's own components are binaries on the machine, and third-party software stays a container because an image is the right way to carry somebody else's build. The delivery of them is issue 142. --- .../00-report.md | 22 +++++++++++++++++-- 1 file changed, 20 insertions(+), 2 deletions(-) diff --git a/04-ISSUES/114-should-the-controller-be-a-container-or-a-process/00-report.md b/04-ISSUES/114-should-the-controller-be-a-container-or-a-process/00-report.md index 6b03b96..194fbd7 100644 --- a/04-ISSUES/114-should-the-controller-be-a-container-or-a-process/00-report.md +++ b/04-ISSUES/114-should-the-controller-be-a-container-or-a-process/00-report.md @@ -1,8 +1,8 @@ --- -status: open +status: resolved opened: 2026-09-24 located-in: [mesh-controller module.json, mesh-host internal/apply] -fixed-by: +fixed-by: ADR 0142 — the mesh's own components are delivered as binaries on the machine, so the controller is a process; the delivery itself is issue 142 amended-design: --- @@ -84,3 +84,21 @@ restart and run-to-completion semantics — so this would not need host-side wor The two do not collapse into one. The controller is not a code-carrying sidecar, and `network: host` is what makes the asymmetry visible here and nowhere else. + +## Answered + +*2026-09-29, in a grooming pass.* This asked a question rather than reporting a defect, and the +question was taken: [ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md) +decides that the mesh's own components — the host, the controller, the catalogue, the builder, the +vault — are **binaries on the machine**, delivered by the mechanism +[ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) built, and that +third-party software (the store, the registry, the broker) stays a container because an image is the +right way to carry somebody else's build. + +So the operating experience this record was written from — every mutating command reached through +`docker exec mesh-controller` — is answered, and answered against the container. + +**The delivery is a separate matter and is not this record's.** Step 1 of it is built and no +component travels yet; that is +[issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md). + From 72eaf52867b34ebe100d83a48c0485bce6a4b792 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 23:13:54 +0200 Subject: [PATCH 4/6] Issue 152: a node whose plan will not compose withdraws its names from every machine ace numbered its two records 147 and 148, which this repository already uses; they become 150 and 151, as 127 became 149. The new record is why the control node could not stop applying: the roster alternates between two values because routeNamesInTheMesh swallows a per-node plan failure, and the roster is part of every container's identity. The loop closes through the control plane's own store, which each pass replaces. --- .../00-report.md | 2 +- .../00-report.md | 4 +- .../00-report.md | 100 ++++++++++++++++++ 3 files changed, 103 insertions(+), 3 deletions(-) rename 04-ISSUES/{147-a-route-is-contributed-before-its-module-is-taken => 150-a-route-is-contributed-before-its-module-is-taken}/00-report.md (97%) rename 04-ISSUES/{148-a-new-name-recreates-every-container-in-the-mesh => 151-a-new-name-recreates-every-container-in-the-mesh}/00-report.md (96%) create mode 100644 04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md diff --git a/04-ISSUES/147-a-route-is-contributed-before-its-module-is-taken/00-report.md b/04-ISSUES/150-a-route-is-contributed-before-its-module-is-taken/00-report.md similarity index 97% rename from 04-ISSUES/147-a-route-is-contributed-before-its-module-is-taken/00-report.md rename to 04-ISSUES/150-a-route-is-contributed-before-its-module-is-taken/00-report.md index f9afe62..77220e6 100644 --- a/04-ISSUES/147-a-route-is-contributed-before-its-module-is-taken/00-report.md +++ b/04-ISSUES/150-a-route-is-contributed-before-its-module-is-taken/00-report.md @@ -6,7 +6,7 @@ located-in: fixed-by: --- -# 147 — A route is contributed before its module is taken +# 150 — A route is contributed before its module is taken ## What was observed diff --git a/04-ISSUES/148-a-new-name-recreates-every-container-in-the-mesh/00-report.md b/04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md similarity index 96% rename from 04-ISSUES/148-a-new-name-recreates-every-container-in-the-mesh/00-report.md rename to 04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md index 8185c6a..bf44ae6 100644 --- a/04-ISSUES/148-a-new-name-recreates-every-container-in-the-mesh/00-report.md +++ b/04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md @@ -7,13 +7,13 @@ located-in: fixed-by: --- -# 148 — A new name recreates every container in the mesh +# 151 — A new name recreates every container in the mesh ## What was observed Migrating one small module on ace (searxng) took four routine controller actions: `node public-domain ace zurag.be`, `assign ace searxng` + push, `unassign ace searxng` + push (a rollback, see -[issue 147](../147-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take +[issue 150](../150-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take + push again. Each push to ace also pushed novox ("this push left g14, novox, shanks behind … sending it too"). novox's host then **replaced every container it runs, twice**: diff --git a/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md new file mode 100644 index 0000000..f65fa1a --- /dev/null +++ b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md @@ -0,0 +1,100 @@ +--- +status: located +opened: 2026-09-29 +located-in: + - mesh-controller cmd/mesh-controller/plan.go (routeNamesInTheMesh skips a node whose plan will not compose) +fixed-by: +amended-design: +--- + +# 152 — A node whose plan will not compose silently removes its names from every machine + +## What was observed + +For at least seventeen minutes after the last operator action, the control-node's host applied all +327 of its resources every ~6.5 minutes without pause, replacing every container on the machine each +time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail, +and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine +sat near 8. Nothing was converging: each pass ended and the next began four seconds later. + +The two machines carrying no containers were not churning. They were only knocked off the bus each +time the control node re-created it, reconnected, re-heard the same declaration and applied it again +as a no-op. + +## Why: the roster alternates between two values, and it is part of every container + +Two consecutive declarations were compared by reading the `--add-host` entries of four containers +the moment each pass created them: + +``` +23:02 ace.internal drive.novox.be g14.internal keycloak.novox.be novox.internal + office.novox.be portainer.novox.be shanks.internal umami.novox.be (9 names) +23:07 … the same nine, and searxng.zurag.be (10 names) +``` + +One routed name — belonging to a module on another machine entirely — leaves the roster and comes +back. Because the roster is part of every container's spec digest ([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)), +each flip is a different identity for every container on the machine, and a running container cannot +have its hosts changed. So every flip replaces all of them. + +**What makes it flip is a swallowed error.** `routeNamesInTheMesh` composes every node's plan to +find the names it serves, and when one will not compose it moves on: + +``` +plan, settings, err := planFor(ctx, open, n.Name) +if err != nil { + continue +} +``` + +A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does +not know", but "the mesh states these names do not exist", to every machine at once. + +## Why it cannot recover on its own + +The loop closes through the control plane's own database: + +1. An apply replaces `mesh-store` — the store the control plane reads — by removing the container, + so postgres comes back through crash recovery. +2. While it recovers it refuses connections: `FATAL: the database system is not yet accepting + connections / Consistent recovery state has not been yet reached` (observed, 21:08:39 UTC, every + pass). +3. `planFor` for the other machine fails against that store. `routeNamesInTheMesh` swallows it and + drops its routed name. +4. The roster changed, so all 327 resources differ, so all are replaced — including `mesh-store`, + and including the bus, which is why the host also cannot report: `applied, and could not tell the + mesh: reporting: nats: connection closed`. +5. Back to 1. + +It is stable in its instability: every pass destroys the evidence the next pass needs to decide it +has nothing to do. Nothing outside the machine has to be wrong for this to continue, and nothing +about it stops. + +## What it is not + +- Not the operator's four actions on the other machine. Those explain the first passes + ([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)); they were + finished seventeen minutes and three full passes before these measurements. +- Not a file that keeps changing. `/etc/hosts`, the bus's account list and the vault's export were + hashed across passes and are byte-identical, and the directory reported `mode 755 to 700` every + pass while already being `700`. Those resources are **misreported as changed** and are worth their + own question, but they are not what moves a container's identity. +- Not the lost report alone. A report that cannot be delivered explains a re-apply; it does not + explain a re-apply that finds 327 differences. + +## Why it matters beyond this outage + +The same `continue` makes every routed name in the mesh conditional on every node's plan composing at +the moment any machine is pushed to. One unreachable or half-migrated machine is enough to withdraw +its names from everywhere — and the withdrawal is indistinguishable, on the receiving machine, from +the operator having removed them. + +The codebase already states the rule this breaks, forty lines away, about the same kind of lookup: + +> A lookup failure is an error, never "not found": collapsing the two composed a declaration without +> the trust whenever the inventory hiccuped, delivered by a push that reported success. + +## How the fix is checked + +A test that composes the roster with one node's plan failing, and requires the compose to fail rather +than return a roster missing that node's names. From 3c2b4fc6b6b95bf3491a6e0b129b0a4a98104833 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 23:26:20 +0200 Subject: [PATCH 5/6] Issue 152 is fixed: a lookup failure is no longer an absence The three gatherers pass over a node whose set does not compose, and raise anything else. 151 stays open: this removes the false reasons a roster changes, not the fact that a real change still replaces every container. --- .../00-report.md | 24 +++++++++++++++---- 1 file changed, 19 insertions(+), 5 deletions(-) diff --git a/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md index f65fa1a..ccef9d9 100644 --- a/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md +++ b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md @@ -1,9 +1,9 @@ --- -status: located +status: resolved opened: 2026-09-29 located-in: - mesh-controller cmd/mesh-controller/plan.go (routeNamesInTheMesh skips a node whose plan will not compose) -fixed-by: +fixed-by: mesh-control fix/152-a-lookup-failure-is-not-an-absence amended-design: --- @@ -94,7 +94,21 @@ The codebase already states the rule this breaks, forty lines away, about the sa > A lookup failure is an error, never "not found": collapsing the two composed a declaration without > the trust whenever the inventory hiccuped, delivered by a push that reported success. -## How the fix is checked +## How it was fixed, and how the fix is checked -A test that composes the roster with one node's plan failing, and requires the compose to fail rather -than return a roster missing that node's names. +`planFor` now marks the two failures that really are a statement about the node — its set not +composing, and a setting that reaches nothing — and the three gatherers pass over those and only +those. Every other failure is raised, naming the machine and the read. + +Five tests hold it: a set that cannot compose is marked as the node's own; a store that cannot be +read is *not*; one incoherent node still does not cost the rest their names; a roster is never +returned beside an error; and the raised failure names what could not be read. + +The two sibling gatherers were audited and fixed the same way — the grant composer, which would have +withheld a consumer's credential, and the private-network membership, which would have taken a +machine off the overlay. Three other `planFor` callers were audited and left alone: they refuse or +report rather than silently withdraw, which is the safe direction. + +**This does not close [issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md).** +A roster that changes for a real reason still replaces every container in the mesh. This removes the +false reasons; whether the roster belongs in a container's identity at all is that record's question. From 18f37c25b2ef0cba6d56d51ebe697db6f9aef2d3 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 23:27:51 +0200 Subject: [PATCH 6/6] Issue 152: the loop is metastable, and it cleared at 23:11 It ran 22:46-23:11, five full replacements, and stopped when a pass happened to read the store during a window it was up. The record said nothing about it stops; that was wrong. Exiting by luck is the finding, not a mitigation. --- .../00-report.md | 24 +++++++++++++------ 1 file changed, 17 insertions(+), 7 deletions(-) diff --git a/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md index ccef9d9..6dcd31f 100644 --- a/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md +++ b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md @@ -11,9 +11,9 @@ amended-design: ## What was observed -For at least seventeen minutes after the last operator action, the control-node's host applied all -327 of its resources every ~6.5 minutes without pause, replacing every container on the machine each -time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail, +For twenty-five minutes, and for seventeen of them after the last operator action, the control-node's +host applied all 327 of its resources every ~6.5 minutes without pause, replacing every container on +the machine each time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail, and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine sat near 8. Nothing was converging: each pass ended and the next began four seconds later. @@ -50,7 +50,7 @@ if err != nil { A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does not know", but "the mesh states these names do not exist", to every machine at once. -## Why it cannot recover on its own +## Why it sustains itself The loop closes through the control plane's own database: @@ -66,9 +66,19 @@ The loop closes through the control plane's own database: mesh: reporting: nats: connection closed`. 5. Back to 1. -It is stable in its instability: every pass destroys the evidence the next pass needs to decide it -has nothing to do. Nothing outside the machine has to be wrong for this to continue, and nothing -about it stops. +Every pass destroys the evidence the next pass needs to decide it has nothing to do, and nothing +outside the machine has to be wrong for it to continue. + +**It is metastable, not permanent.** It ran from 22:46 to 23:11 — five full replacements of every +container on the machine — and then stopped on its own, when one pass happened to read the store +during a window it was up, composed the same roster twice running, and found nothing to do. Load fell +from 7.8 to 1.7 and the machine returned to its five-minute idle tick. + +That it ends by luck is the point, not a mitigation. The exit condition is a race the mesh does not +control, the operator cannot see, and nothing reports; the same four actions on a slower machine, or +a larger store, would not have found it. An outage that clears itself after twenty-five minutes and +five restarts of the forge, the directory and mail is not a smaller fault than one that does not — it +is the same fault, harder to catch. ## What it is not