Files
hq/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md
T
jschoubben 47909c6b71 The records pointed at branches that no longer exist, and two fixes had no sequel
Three issues named the branch that fixed them, and a branch is deleted
when it merges — so every `fixed-by:` was a pointer that resolved to
nothing by the time anyone followed it. They name commits and pull
requests now, and playbook 03 says to.

Two records were missing the thing a reader arrives for. 146 did not say
that one of its fixes crash-looped the control plane on a running mesh,
which is the whole reason the delivery subject carries the stream and the
raise path was the only one exercised. 151 did not say that 152 removed
the false reasons its roster moved, or that it stays open for the real
ones.

ADR 0080 enumerates what cycle.py enforces and named four things; it
enforces five. A progressive insight names the fifth — the decision
stands, the list had gone stale. The checks README and playbook 03 gained
the same rule, and 155 points at all three.
2026-09-30 00:28:35 +02:00

6.8 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
resolved 2026-09-29
mesh-controller cmd/mesh-controller/plan.go (routeNamesInTheMesh skips a node whose plan will not compose)
mesh-controller 6c5dfd0 (PR 147)

152 — A node whose plan will not compose silently removes its names from every machine

What was observed

For twenty-five minutes, and for seventeen of them after the last operator action, the control-node's host applied all 327 of its resources every ~6.5 minutes without pause, replacing every container on the machine each time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail, and the bus the mesh runs on. The forge's web surface answered 502 throughout; load on the machine sat near 8. Nothing was converging: each pass ended and the next began four seconds later.

The two machines carrying no containers were not churning. They were only knocked off the bus each time the control node re-created it, reconnected, re-heard the same declaration and applied it again as a no-op.

Why: the roster alternates between two values, and it is part of every container

Two consecutive declarations were compared by reading the --add-host entries of four containers the moment each pass created them:

23:02  ace.internal  drive.novox.be  g14.internal  keycloak.novox.be  novox.internal
       office.novox.be  portainer.novox.be  shanks.internal  umami.novox.be          (9 names)
23:07  … the same nine, and searxng.zurag.be                                         (10 names)

One routed name — belonging to a module on another machine entirely — leaves the roster and comes back. Because the roster is part of every container's spec digest (issue 151), each flip is a different identity for every container on the machine, and a running container cannot have its hosts changed. So every flip replaces all of them.

What makes it flip is a swallowed error. routeNamesInTheMesh composes every node's plan to find the names it serves, and when one will not compose it moves on:

plan, settings, err := planFor(ctx, open, n.Name)
if err != nil {
        continue
}

A node whose plan cannot be composed right now therefore contributes no names — not "the mesh does not know", but "the mesh states these names do not exist", to every machine at once.

Why it sustains itself

The loop closes through the control plane's own database:

  1. An apply replaces mesh-store — the store the control plane reads — by removing the container, so postgres comes back through crash recovery.
  2. While it recovers it refuses connections: FATAL: the database system is not yet accepting connections / Consistent recovery state has not been yet reached (observed, 21:08:39 UTC, every pass).
  3. planFor for the other machine fails against that store. routeNamesInTheMesh swallows it and drops its routed name.
  4. The roster changed, so all 327 resources differ, so all are replaced — including mesh-store, and including the bus, which is why the host also cannot report: applied, and could not tell the mesh: reporting: nats: connection closed.
  5. Back to 1.

Every pass destroys the evidence the next pass needs to decide it has nothing to do, and nothing outside the machine has to be wrong for it to continue.

It is metastable, not permanent. It ran from 22:46 to 23:11 — five full replacements of every container on the machine — and then stopped on its own, when one pass happened to read the store during a window it was up, composed the same roster twice running, and found nothing to do. Load fell from 7.8 to 1.7 and the machine returned to its five-minute idle tick.

That it ends by luck is the point, not a mitigation. The exit condition is a race the mesh does not control, the operator cannot see, and nothing reports; the same four actions on a slower machine, or a larger store, would not have found it. An outage that clears itself after twenty-five minutes and five restarts of the forge, the directory and mail is not a smaller fault than one that does not — it is the same fault, harder to catch.

What it is not

  • Not the operator's four actions on the other machine. Those explain the first passes (issue 151); they were finished seventeen minutes and three full passes before these measurements.
  • Not a file that keeps changing. /etc/hosts, the bus's account list and the vault's export were hashed across passes and are byte-identical, and the directory reported mode 755 to 700 every pass while already being 700. Those resources are misreported as changed and are worth their own question, but they are not what moves a container's identity.
  • Not the lost report alone. A report that cannot be delivered explains a re-apply; it does not explain a re-apply that finds 327 differences.

Why it matters beyond this outage

The same continue makes every routed name in the mesh conditional on every node's plan composing at the moment any machine is pushed to. One unreachable or half-migrated machine is enough to withdraw its names from everywhere — and the withdrawal is indistinguishable, on the receiving machine, from the operator having removed them.

The codebase already states the rule this breaks, forty lines away, about the same kind of lookup:

A lookup failure is an error, never "not found": collapsing the two composed a declaration without the trust whenever the inventory hiccuped, delivered by a push that reported success.

How it was fixed, and how the fix is checked

planFor now marks the two failures that really are a statement about the node — its set not composing, and a setting that reaches nothing — and the three gatherers pass over those and only those. Every other failure is raised, naming the machine and the read.

Five tests hold it: a set that cannot compose is marked as the node's own; a store that cannot be read is not; one incoherent node still does not cost the rest their names; a roster is never returned beside an error; and the raised failure names what could not be read.

The two sibling gatherers were audited and fixed the same way — the grant composer, which would have withheld a consumer's credential, and the private-network membership, which would have taken a machine off the overlay. Three other planFor callers were audited and left alone: they refuse or report rather than silently withdraw, which is the safe direction.

This does not close issue 151. A roster that changes for a real reason still replaces every container in the mesh. This removes the false reasons; whether the roster belongs in a container's identity at all is that record's question.