From 72eaf52867b34ebe100d83a48c0485bce6a4b792 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 29 Sep 2026 23:13:54 +0200 Subject: [PATCH] Issue 152: a node whose plan will not compose withdraws its names from every machine ace numbered its two records 147 and 148, which this repository already uses; they become 150 and 151, as 127 became 149. The new record is why the control node could not stop applying: the roster alternates between two values because routeNamesInTheMesh swallows a per-node plan failure, and the roster is part of every container's identity. The loop closes through the control plane's own store, which each pass replaces. --- .../00-report.md | 2 +- .../00-report.md | 4 +- .../00-report.md | 100 ++++++++++++++++++ 3 files changed, 103 insertions(+), 3 deletions(-) rename 04-ISSUES/{147-a-route-is-contributed-before-its-module-is-taken => 150-a-route-is-contributed-before-its-module-is-taken}/00-report.md (97%) rename 04-ISSUES/{148-a-new-name-recreates-every-container-in-the-mesh => 151-a-new-name-recreates-every-container-in-the-mesh}/00-report.md (96%) create mode 100644 04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md diff --git a/04-ISSUES/147-a-route-is-contributed-before-its-module-is-taken/00-report.md b/04-ISSUES/150-a-route-is-contributed-before-its-module-is-taken/00-report.md similarity index 97% rename from 04-ISSUES/147-a-route-is-contributed-before-its-module-is-taken/00-report.md rename to 04-ISSUES/150-a-route-is-contributed-before-its-module-is-taken/00-report.md index f9afe62..77220e6 100644 --- a/04-ISSUES/147-a-route-is-contributed-before-its-module-is-taken/00-report.md +++ b/04-ISSUES/150-a-route-is-contributed-before-its-module-is-taken/00-report.md @@ -6,7 +6,7 @@ located-in: fixed-by: --- -# 147 — A route is contributed before its module is taken +# 150 — A route is contributed before its module is taken ## What was observed diff --git a/04-ISSUES/148-a-new-name-recreates-every-container-in-the-mesh/00-report.md b/04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md similarity index 96% rename from 04-ISSUES/148-a-new-name-recreates-every-container-in-the-mesh/00-report.md rename to 04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md index 8185c6a..bf44ae6 100644 --- a/04-ISSUES/148-a-new-name-recreates-every-container-in-the-mesh/00-report.md +++ b/04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md @@ -7,13 +7,13 @@ located-in: fixed-by: --- -# 148 — A new name recreates every container in the mesh +# 151 — A new name recreates every container in the mesh ## What was observed Migrating one small module on ace (searxng) took four routine controller actions: `node public-domain ace zurag.be`, `assign ace searxng` + push, `unassign ace searxng` + push (a rollback, see -[issue 147](../147-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take +[issue 150](../150-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take + push again. Each push to ace also pushed novox ("this push left g14, novox, shanks behind … sending it too"). novox's host then **replaced every container it runs, twice**: diff --git a/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md new file mode 100644 index 0000000..f65fa1a --- /dev/null +++ b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md @@ -0,0 +1,100 @@ +--- +status: located +opened: 2026-09-29 +located-in: + - mesh-controller cmd/mesh-controller/plan.go (routeNamesInTheMesh skips a node whose plan will not compose) +fixed-by: +amended-design: +--- + +# 152 — A node whose plan will not compose silently removes its names from every machine + +## What was observed + +For at least seventeen minutes after the last operator action, the control-node's host applied all +327 of its resources every ~6.5 minutes without pause, replacing every container on the machine each +time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail, +and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine +sat near 8. Nothing was converging: each pass ended and the next began four seconds later. + +The two machines carrying no containers were not churning. They were only knocked off the bus each +time the control node re-created it, reconnected, re-heard the same declaration and applied it again +as a no-op. + +## Why: the roster alternates between two values, and it is part of every container + +Two consecutive declarations were compared by reading the `--add-host` entries of four containers +the moment each pass created them: + +``` +23:02 ace.internal drive.novox.be g14.internal keycloak.novox.be novox.internal + office.novox.be portainer.novox.be shanks.internal umami.novox.be (9 names) +23:07 … the same nine, and searxng.zurag.be (10 names) +``` + +One routed name — belonging to a module on another machine entirely — leaves the roster and comes +back. Because the roster is part of every container's spec digest ([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)), +each flip is a different identity for every container on the machine, and a running container cannot +have its hosts changed. So every flip replaces all of them. + +**What makes it flip is a swallowed error.** `routeNamesInTheMesh` composes every node's plan to +find the names it serves, and when one will not compose it moves on: + +``` +plan, settings, err := planFor(ctx, open, n.Name) +if err != nil { + continue +} +``` + +A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does +not know", but "the mesh states these names do not exist", to every machine at once. + +## Why it cannot recover on its own + +The loop closes through the control plane's own database: + +1. An apply replaces `mesh-store` — the store the control plane reads — by removing the container, + so postgres comes back through crash recovery. +2. While it recovers it refuses connections: `FATAL: the database system is not yet accepting + connections / Consistent recovery state has not been yet reached` (observed, 21:08:39 UTC, every + pass). +3. `planFor` for the other machine fails against that store. `routeNamesInTheMesh` swallows it and + drops its routed name. +4. The roster changed, so all 327 resources differ, so all are replaced — including `mesh-store`, + and including the bus, which is why the host also cannot report: `applied, and could not tell the + mesh: reporting: nats: connection closed`. +5. Back to 1. + +It is stable in its instability: every pass destroys the evidence the next pass needs to decide it +has nothing to do. Nothing outside the machine has to be wrong for this to continue, and nothing +about it stops. + +## What it is not + +- Not the operator's four actions on the other machine. Those explain the first passes + ([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)); they were + finished seventeen minutes and three full passes before these measurements. +- Not a file that keeps changing. `/etc/hosts`, the bus's account list and the vault's export were + hashed across passes and are byte-identical, and the directory reported `mode 755 to 700` every + pass while already being `700`. Those resources are **misreported as changed** and are worth their + own question, but they are not what moves a container's identity. +- Not the lost report alone. A report that cannot be delivered explains a re-apply; it does not + explain a re-apply that finds 327 differences. + +## Why it matters beyond this outage + +The same `continue` makes every routed name in the mesh conditional on every node's plan composing at +the moment any machine is pushed to. One unreachable or half-migrated machine is enough to withdraw +its names from everywhere — and the withdrawal is indistinguishable, on the receiving machine, from +the operator having removed them. + +The codebase already states the rule this breaks, forty lines away, about the same kind of lookup: + +> A lookup failure is an error, never "not found": collapsing the two composed a declaration without +> the trust whenever the inventory hiccuped, delivered by a push that reported success. + +## How the fix is checked + +A test that composes the roster with one node's plan failing, and requires the compose to fail rather +than return a roster missing that node's names.