Three issues named the branch that fixed them, and a branch is deleted when it merges — so every `fixed-by:` was a pointer that resolved to nothing by the time anyone followed it. They name commits and pull requests now, and playbook 03 says to. Two records were missing the thing a reader arrives for. 146 did not say that one of its fixes crash-looped the control plane on a running mesh, which is the whole reason the delivery subject carries the stream and the raise path was the only one exercised. 151 did not say that 152 removed the false reasons its roster moved, or that it stays open for the real ones. ADR 0080 enumerates what cycle.py enforces and named four things; it enforces five. A progressive insight names the fifth — the decision stands, the list had gone stale. The checks README and playbook 03 gained the same rule, and 155 points at all three.
6.8 KiB
status, opened, located-in, fixed-by, amended-design
| status | opened | located-in | fixed-by | amended-design | |
|---|---|---|---|---|---|
| resolved | 2026-09-29 |
|
mesh-controller 6c5dfd0 (PR 147) |
152 — A node whose plan will not compose silently removes its names from every machine
What was observed
For twenty-five minutes, and for seventeen of them after the last operator action, the control-node's
host applied all 327 of its resources every ~6.5 minutes without pause, replacing every container on
the machine each time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail,
and the bus the mesh runs on. The forge's web surface answered 502 throughout; load on the machine
sat near 8. Nothing was converging: each pass ended and the next began four seconds later.
The two machines carrying no containers were not churning. They were only knocked off the bus each time the control node re-created it, reconnected, re-heard the same declaration and applied it again as a no-op.
Why: the roster alternates between two values, and it is part of every container
Two consecutive declarations were compared by reading the --add-host entries of four containers
the moment each pass created them:
23:02 ace.internal drive.novox.be g14.internal keycloak.novox.be novox.internal
office.novox.be portainer.novox.be shanks.internal umami.novox.be (9 names)
23:07 … the same nine, and searxng.zurag.be (10 names)
One routed name — belonging to a module on another machine entirely — leaves the roster and comes back. Because the roster is part of every container's spec digest (issue 151), each flip is a different identity for every container on the machine, and a running container cannot have its hosts changed. So every flip replaces all of them.
What makes it flip is a swallowed error. routeNamesInTheMesh composes every node's plan to
find the names it serves, and when one will not compose it moves on:
plan, settings, err := planFor(ctx, open, n.Name)
if err != nil {
continue
}
A node whose plan cannot be composed right now therefore contributes no names — not "the mesh does not know", but "the mesh states these names do not exist", to every machine at once.
Why it sustains itself
The loop closes through the control plane's own database:
- An apply replaces
mesh-store— the store the control plane reads — by removing the container, so postgres comes back through crash recovery. - While it recovers it refuses connections:
FATAL: the database system is not yet accepting connections / Consistent recovery state has not been yet reached(observed, 21:08:39 UTC, every pass). planForfor the other machine fails against that store.routeNamesInTheMeshswallows it and drops its routed name.- The roster changed, so all 327 resources differ, so all are replaced — including
mesh-store, and including the bus, which is why the host also cannot report:applied, and could not tell the mesh: reporting: nats: connection closed. - Back to 1.
Every pass destroys the evidence the next pass needs to decide it has nothing to do, and nothing outside the machine has to be wrong for it to continue.
It is metastable, not permanent. It ran from 22:46 to 23:11 — five full replacements of every container on the machine — and then stopped on its own, when one pass happened to read the store during a window it was up, composed the same roster twice running, and found nothing to do. Load fell from 7.8 to 1.7 and the machine returned to its five-minute idle tick.
That it ends by luck is the point, not a mitigation. The exit condition is a race the mesh does not control, the operator cannot see, and nothing reports; the same four actions on a slower machine, or a larger store, would not have found it. An outage that clears itself after twenty-five minutes and five restarts of the forge, the directory and mail is not a smaller fault than one that does not — it is the same fault, harder to catch.
What it is not
- Not the operator's four actions on the other machine. Those explain the first passes (issue 151); they were finished seventeen minutes and three full passes before these measurements.
- Not a file that keeps changing.
/etc/hosts, the bus's account list and the vault's export were hashed across passes and are byte-identical, and the directory reportedmode 755 to 700every pass while already being700. Those resources are misreported as changed and are worth their own question, but they are not what moves a container's identity. - Not the lost report alone. A report that cannot be delivered explains a re-apply; it does not explain a re-apply that finds 327 differences.
Why it matters beyond this outage
The same continue makes every routed name in the mesh conditional on every node's plan composing at
the moment any machine is pushed to. One unreachable or half-migrated machine is enough to withdraw
its names from everywhere — and the withdrawal is indistinguishable, on the receiving machine, from
the operator having removed them.
The codebase already states the rule this breaks, forty lines away, about the same kind of lookup:
A lookup failure is an error, never "not found": collapsing the two composed a declaration without the trust whenever the inventory hiccuped, delivered by a push that reported success.
How it was fixed, and how the fix is checked
planFor now marks the two failures that really are a statement about the node — its set not
composing, and a setting that reaches nothing — and the three gatherers pass over those and only
those. Every other failure is raised, naming the machine and the read.
Five tests hold it: a set that cannot compose is marked as the node's own; a store that cannot be read is not; one incoherent node still does not cost the rest their names; a roster is never returned beside an error; and the raised failure names what could not be read.
The two sibling gatherers were audited and fixed the same way — the grant composer, which would have
withheld a consumer's credential, and the private-network membership, which would have taken a
machine off the overlay. Three other planFor callers were audited and left alone: they refuse or
report rather than silently withdraw, which is the safe direction.
This does not close issue 151. A roster that changes for a real reason still replaces every container in the mesh. This removes the false reasons; whether the roster belongs in a container's identity at all is that record's question.