The three gatherers pass over a node whose set does not compose, and raise anything else. 151 stays open: this removes the false reasons a roster changes, not the fact that a real change still replaces every container.
6.1 KiB
status, opened, located-in, fixed-by, amended-design
| status | opened | located-in | fixed-by | amended-design | |
|---|---|---|---|---|---|
| resolved | 2026-09-29 |
|
mesh-control fix/152-a-lookup-failure-is-not-an-absence |
152 — A node whose plan will not compose silently removes its names from every machine
What was observed
For at least seventeen minutes after the last operator action, the control-node's host applied all
327 of its resources every ~6.5 minutes without pause, replacing every container on the machine each
time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail,
and the bus the mesh runs on. The forge's web surface answered 502 throughout; load on the machine
sat near 8. Nothing was converging: each pass ended and the next began four seconds later.
The two machines carrying no containers were not churning. They were only knocked off the bus each time the control node re-created it, reconnected, re-heard the same declaration and applied it again as a no-op.
Why: the roster alternates between two values, and it is part of every container
Two consecutive declarations were compared by reading the --add-host entries of four containers
the moment each pass created them:
23:02 ace.internal drive.novox.be g14.internal keycloak.novox.be novox.internal
office.novox.be portainer.novox.be shanks.internal umami.novox.be (9 names)
23:07 … the same nine, and searxng.zurag.be (10 names)
One routed name — belonging to a module on another machine entirely — leaves the roster and comes back. Because the roster is part of every container's spec digest (issue 151), each flip is a different identity for every container on the machine, and a running container cannot have its hosts changed. So every flip replaces all of them.
What makes it flip is a swallowed error. routeNamesInTheMesh composes every node's plan to
find the names it serves, and when one will not compose it moves on:
plan, settings, err := planFor(ctx, open, n.Name)
if err != nil {
continue
}
A node whose plan cannot be composed right now therefore contributes no names — not "the mesh does not know", but "the mesh states these names do not exist", to every machine at once.
Why it cannot recover on its own
The loop closes through the control plane's own database:
- An apply replaces
mesh-store— the store the control plane reads — by removing the container, so postgres comes back through crash recovery. - While it recovers it refuses connections:
FATAL: the database system is not yet accepting connections / Consistent recovery state has not been yet reached(observed, 21:08:39 UTC, every pass). planForfor the other machine fails against that store.routeNamesInTheMeshswallows it and drops its routed name.- The roster changed, so all 327 resources differ, so all are replaced — including
mesh-store, and including the bus, which is why the host also cannot report:applied, and could not tell the mesh: reporting: nats: connection closed. - Back to 1.
It is stable in its instability: every pass destroys the evidence the next pass needs to decide it has nothing to do. Nothing outside the machine has to be wrong for this to continue, and nothing about it stops.
What it is not
- Not the operator's four actions on the other machine. Those explain the first passes (issue 151); they were finished seventeen minutes and three full passes before these measurements.
- Not a file that keeps changing.
/etc/hosts, the bus's account list and the vault's export were hashed across passes and are byte-identical, and the directory reportedmode 755 to 700every pass while already being700. Those resources are misreported as changed and are worth their own question, but they are not what moves a container's identity. - Not the lost report alone. A report that cannot be delivered explains a re-apply; it does not explain a re-apply that finds 327 differences.
Why it matters beyond this outage
The same continue makes every routed name in the mesh conditional on every node's plan composing at
the moment any machine is pushed to. One unreachable or half-migrated machine is enough to withdraw
its names from everywhere — and the withdrawal is indistinguishable, on the receiving machine, from
the operator having removed them.
The codebase already states the rule this breaks, forty lines away, about the same kind of lookup:
A lookup failure is an error, never "not found": collapsing the two composed a declaration without the trust whenever the inventory hiccuped, delivered by a push that reported success.
How it was fixed, and how the fix is checked
planFor now marks the two failures that really are a statement about the node — its set not
composing, and a setting that reaches nothing — and the three gatherers pass over those and only
those. Every other failure is raised, naming the machine and the read.
Five tests hold it: a set that cannot compose is marked as the node's own; a store that cannot be read is not; one incoherent node still does not cost the rest their names; a roster is never returned beside an error; and the raised failure names what could not be read.
The two sibling gatherers were audited and fixed the same way — the grant composer, which would have
withheld a consumer's credential, and the private-network membership, which would have taken a
machine off the overlay. Three other planFor callers were audited and left alone: they refuse or
report rather than silently withdraw, which is the safe direction.
This does not close issue 151. A roster that changes for a real reason still replaces every container in the mesh. This removes the false reasons; whether the roster belongs in a container's identity at all is that record's question.