Files
hq/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md
T
jschoubben 47909c6b71 The records pointed at branches that no longer exist, and two fixes had no sequel
Three issues named the branch that fixed them, and a branch is deleted
when it merges — so every `fixed-by:` was a pointer that resolved to
nothing by the time anyone followed it. They name commits and pull
requests now, and playbook 03 says to.

Two records were missing the thing a reader arrives for. 146 did not say
that one of its fixes crash-looped the control plane on a running mesh,
which is the whole reason the delivery subject carries the stream and the
raise path was the only one exercised. 151 did not say that 152 removed
the false reasons its roster moved, or that it stays open for the real
ones.

ADR 0080 enumerates what cycle.py enforces and named four things; it
enforces five. A progressive insight names the fifth — the decision
stands, the list had gone stale. The checks README and playbook 03 gained
the same rule, and 155 points at all three.
2026-09-30 00:28:35 +02:00

125 lines
6.8 KiB
Markdown

---
status: resolved
opened: 2026-09-29
located-in:
- mesh-controller cmd/mesh-controller/plan.go (routeNamesInTheMesh skips a node whose plan will not compose)
fixed-by: mesh-controller 6c5dfd0 (PR 147)
amended-design:
---
# 152 — A node whose plan will not compose silently removes its names from every machine
## What was observed
For twenty-five minutes, and for seventeen of them after the last operator action, the control-node's
host applied all 327 of its resources every ~6.5 minutes without pause, replacing every container on
the machine each time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail,
and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine
sat near 8. Nothing was converging: each pass ended and the next began four seconds later.
The two machines carrying no containers were not churning. They were only knocked off the bus each
time the control node re-created it, reconnected, re-heard the same declaration and applied it again
as a no-op.
## Why: the roster alternates between two values, and it is part of every container
Two consecutive declarations were compared by reading the `--add-host` entries of four containers
the moment each pass created them:
```
23:02 ace.internal drive.novox.be g14.internal keycloak.novox.be novox.internal
office.novox.be portainer.novox.be shanks.internal umami.novox.be (9 names)
23:07 … the same nine, and searxng.zurag.be (10 names)
```
One routed name — belonging to a module on another machine entirely — leaves the roster and comes
back. Because the roster is part of every container's spec digest ([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)),
each flip is a different identity for every container on the machine, and a running container cannot
have its hosts changed. So every flip replaces all of them.
**What makes it flip is a swallowed error.** `routeNamesInTheMesh` composes every node's plan to
find the names it serves, and when one will not compose it moves on:
```
plan, settings, err := planFor(ctx, open, n.Name)
if err != nil {
continue
}
```
A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does
not know", but "the mesh states these names do not exist", to every machine at once.
## Why it sustains itself
The loop closes through the control plane's own database:
1. An apply replaces `mesh-store` — the store the control plane reads — by removing the container,
so postgres comes back through crash recovery.
2. While it recovers it refuses connections: `FATAL: the database system is not yet accepting
connections / Consistent recovery state has not been yet reached` (observed, 21:08:39 UTC, every
pass).
3. `planFor` for the other machine fails against that store. `routeNamesInTheMesh` swallows it and
drops its routed name.
4. The roster changed, so all 327 resources differ, so all are replaced — including `mesh-store`,
and including the bus, which is why the host also cannot report: `applied, and could not tell the
mesh: reporting: nats: connection closed`.
5. Back to 1.
Every pass destroys the evidence the next pass needs to decide it has nothing to do, and nothing
outside the machine has to be wrong for it to continue.
**It is metastable, not permanent.** It ran from 22:46 to 23:11 — five full replacements of every
container on the machine — and then stopped on its own, when one pass happened to read the store
during a window it was up, composed the same roster twice running, and found nothing to do. Load fell
from 7.8 to 1.7 and the machine returned to its five-minute idle tick.
That it ends by luck is the point, not a mitigation. The exit condition is a race the mesh does not
control, the operator cannot see, and nothing reports; the same four actions on a slower machine, or
a larger store, would not have found it. An outage that clears itself after twenty-five minutes and
five restarts of the forge, the directory and mail is not a smaller fault than one that does not — it
is the same fault, harder to catch.
## What it is not
- Not the operator's four actions on the other machine. Those explain the first passes
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)); they were
finished seventeen minutes and three full passes before these measurements.
- Not a file that keeps changing. `/etc/hosts`, the bus's account list and the vault's export were
hashed across passes and are byte-identical, and the directory reported `mode 755 to 700` every
pass while already being `700`. Those resources are **misreported as changed** and are worth their
own question, but they are not what moves a container's identity.
- Not the lost report alone. A report that cannot be delivered explains a re-apply; it does not
explain a re-apply that finds 327 differences.
## Why it matters beyond this outage
The same `continue` makes every routed name in the mesh conditional on every node's plan composing at
the moment any machine is pushed to. One unreachable or half-migrated machine is enough to withdraw
its names from everywhere — and the withdrawal is indistinguishable, on the receiving machine, from
the operator having removed them.
The codebase already states the rule this breaks, forty lines away, about the same kind of lookup:
> A lookup failure is an error, never "not found": collapsing the two composed a declaration without
> the trust whenever the inventory hiccuped, delivered by a push that reported success.
## How it was fixed, and how the fix is checked
`planFor` now marks the two failures that really are a statement about the node — its set not
composing, and a setting that reaches nothing — and the three gatherers pass over those and only
those. Every other failure is raised, naming the machine and the read.
Five tests hold it: a set that cannot compose is marked as the node's own; a store that cannot be
read is *not*; one incoherent node still does not cost the rest their names; a roster is never
returned beside an error; and the raised failure names what could not be read.
The two sibling gatherers were audited and fixed the same way — the grant composer, which would have
withheld a consumer's credential, and the private-network membership, which would have taken a
machine off the overlay. Three other `planFor` callers were audited and left alone: they refuse or
report rather than silently withdraw, which is the safe direction.
**This does not close [issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md).**
A roster that changes for a real reason still replaces every container in the mesh. This removes the
false reasons; whether the roster belongs in a container's identity at all is that record's question.