Files
hq/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md
T
jschoubben 3c2b4fc6b6 Issue 152 is fixed: a lookup failure is no longer an absence
The three gatherers pass over a node whose set does not compose, and
raise anything else. 151 stays open: this removes the false reasons a
roster changes, not the fact that a real change still replaces every
container.
2026-09-29 23:26:20 +02:00

115 lines
6.1 KiB
Markdown

---
status: resolved
opened: 2026-09-29
located-in:
- mesh-controller cmd/mesh-controller/plan.go (routeNamesInTheMesh skips a node whose plan will not compose)
fixed-by: mesh-control fix/152-a-lookup-failure-is-not-an-absence
amended-design:
---
# 152 — A node whose plan will not compose silently removes its names from every machine
## What was observed
For at least seventeen minutes after the last operator action, the control-node's host applied all
327 of its resources every ~6.5 minutes without pause, replacing every container on the machine each
time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail,
and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine
sat near 8. Nothing was converging: each pass ended and the next began four seconds later.
The two machines carrying no containers were not churning. They were only knocked off the bus each
time the control node re-created it, reconnected, re-heard the same declaration and applied it again
as a no-op.
## Why: the roster alternates between two values, and it is part of every container
Two consecutive declarations were compared by reading the `--add-host` entries of four containers
the moment each pass created them:
```
23:02 ace.internal drive.novox.be g14.internal keycloak.novox.be novox.internal
office.novox.be portainer.novox.be shanks.internal umami.novox.be (9 names)
23:07 … the same nine, and searxng.zurag.be (10 names)
```
One routed name — belonging to a module on another machine entirely — leaves the roster and comes
back. Because the roster is part of every container's spec digest ([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)),
each flip is a different identity for every container on the machine, and a running container cannot
have its hosts changed. So every flip replaces all of them.
**What makes it flip is a swallowed error.** `routeNamesInTheMesh` composes every node's plan to
find the names it serves, and when one will not compose it moves on:
```
plan, settings, err := planFor(ctx, open, n.Name)
if err != nil {
continue
}
```
A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does
not know", but "the mesh states these names do not exist", to every machine at once.
## Why it cannot recover on its own
The loop closes through the control plane's own database:
1. An apply replaces `mesh-store` — the store the control plane reads — by removing the container,
so postgres comes back through crash recovery.
2. While it recovers it refuses connections: `FATAL: the database system is not yet accepting
connections / Consistent recovery state has not been yet reached` (observed, 21:08:39 UTC, every
pass).
3. `planFor` for the other machine fails against that store. `routeNamesInTheMesh` swallows it and
drops its routed name.
4. The roster changed, so all 327 resources differ, so all are replaced — including `mesh-store`,
and including the bus, which is why the host also cannot report: `applied, and could not tell the
mesh: reporting: nats: connection closed`.
5. Back to 1.
It is stable in its instability: every pass destroys the evidence the next pass needs to decide it
has nothing to do. Nothing outside the machine has to be wrong for this to continue, and nothing
about it stops.
## What it is not
- Not the operator's four actions on the other machine. Those explain the first passes
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)); they were
finished seventeen minutes and three full passes before these measurements.
- Not a file that keeps changing. `/etc/hosts`, the bus's account list and the vault's export were
hashed across passes and are byte-identical, and the directory reported `mode 755 to 700` every
pass while already being `700`. Those resources are **misreported as changed** and are worth their
own question, but they are not what moves a container's identity.
- Not the lost report alone. A report that cannot be delivered explains a re-apply; it does not
explain a re-apply that finds 327 differences.
## Why it matters beyond this outage
The same `continue` makes every routed name in the mesh conditional on every node's plan composing at
the moment any machine is pushed to. One unreachable or half-migrated machine is enough to withdraw
its names from everywhere — and the withdrawal is indistinguishable, on the receiving machine, from
the operator having removed them.
The codebase already states the rule this breaks, forty lines away, about the same kind of lookup:
> A lookup failure is an error, never "not found": collapsing the two composed a declaration without
> the trust whenever the inventory hiccuped, delivered by a push that reported success.
## How it was fixed, and how the fix is checked
`planFor` now marks the two failures that really are a statement about the node — its set not
composing, and a setting that reaches nothing — and the three gatherers pass over those and only
those. Every other failure is raised, naming the machine and the read.
Five tests hold it: a set that cannot compose is marked as the node's own; a store that cannot be
read is *not*; one incoherent node still does not cost the rest their names; a roster is never
returned beside an error; and the raised failure names what could not be read.
The two sibling gatherers were audited and fixed the same way — the grant composer, which would have
withheld a consumer's credential, and the private-network membership, which would have taken a
machine off the overlay. Three other `planFor` callers were audited and left alone: they refuse or
report rather than silently withdraw, which is the safe direction.
**This does not close [issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md).**
A roster that changes for a real reason still replaces every container in the mesh. This removes the
false reasons; whether the roster belongs in a container's identity at all is that record's question.