It ran 22:46-23:11, five full replacements, and stopped when a pass happened to read the store during a window it was up. The record said nothing about it stops; that was wrong. Exiting by luck is the finding, not a mitigation.
125 lines
6.8 KiB
Markdown
125 lines
6.8 KiB
Markdown
---
|
|
status: resolved
|
|
opened: 2026-09-29
|
|
located-in:
|
|
- mesh-controller cmd/mesh-controller/plan.go (routeNamesInTheMesh skips a node whose plan will not compose)
|
|
fixed-by: mesh-control fix/152-a-lookup-failure-is-not-an-absence
|
|
amended-design:
|
|
---
|
|
|
|
# 152 — A node whose plan will not compose silently removes its names from every machine
|
|
|
|
## What was observed
|
|
|
|
For twenty-five minutes, and for seventeen of them after the last operator action, the control-node's
|
|
host applied all 327 of its resources every ~6.5 minutes without pause, replacing every container on
|
|
the machine each time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail,
|
|
and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine
|
|
sat near 8. Nothing was converging: each pass ended and the next began four seconds later.
|
|
|
|
The two machines carrying no containers were not churning. They were only knocked off the bus each
|
|
time the control node re-created it, reconnected, re-heard the same declaration and applied it again
|
|
as a no-op.
|
|
|
|
## Why: the roster alternates between two values, and it is part of every container
|
|
|
|
Two consecutive declarations were compared by reading the `--add-host` entries of four containers
|
|
the moment each pass created them:
|
|
|
|
```
|
|
23:02 ace.internal drive.novox.be g14.internal keycloak.novox.be novox.internal
|
|
office.novox.be portainer.novox.be shanks.internal umami.novox.be (9 names)
|
|
23:07 … the same nine, and searxng.zurag.be (10 names)
|
|
```
|
|
|
|
One routed name — belonging to a module on another machine entirely — leaves the roster and comes
|
|
back. Because the roster is part of every container's spec digest ([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)),
|
|
each flip is a different identity for every container on the machine, and a running container cannot
|
|
have its hosts changed. So every flip replaces all of them.
|
|
|
|
**What makes it flip is a swallowed error.** `routeNamesInTheMesh` composes every node's plan to
|
|
find the names it serves, and when one will not compose it moves on:
|
|
|
|
```
|
|
plan, settings, err := planFor(ctx, open, n.Name)
|
|
if err != nil {
|
|
continue
|
|
}
|
|
```
|
|
|
|
A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does
|
|
not know", but "the mesh states these names do not exist", to every machine at once.
|
|
|
|
## Why it sustains itself
|
|
|
|
The loop closes through the control plane's own database:
|
|
|
|
1. An apply replaces `mesh-store` — the store the control plane reads — by removing the container,
|
|
so postgres comes back through crash recovery.
|
|
2. While it recovers it refuses connections: `FATAL: the database system is not yet accepting
|
|
connections / Consistent recovery state has not been yet reached` (observed, 21:08:39 UTC, every
|
|
pass).
|
|
3. `planFor` for the other machine fails against that store. `routeNamesInTheMesh` swallows it and
|
|
drops its routed name.
|
|
4. The roster changed, so all 327 resources differ, so all are replaced — including `mesh-store`,
|
|
and including the bus, which is why the host also cannot report: `applied, and could not tell the
|
|
mesh: reporting: nats: connection closed`.
|
|
5. Back to 1.
|
|
|
|
Every pass destroys the evidence the next pass needs to decide it has nothing to do, and nothing
|
|
outside the machine has to be wrong for it to continue.
|
|
|
|
**It is metastable, not permanent.** It ran from 22:46 to 23:11 — five full replacements of every
|
|
container on the machine — and then stopped on its own, when one pass happened to read the store
|
|
during a window it was up, composed the same roster twice running, and found nothing to do. Load fell
|
|
from 7.8 to 1.7 and the machine returned to its five-minute idle tick.
|
|
|
|
That it ends by luck is the point, not a mitigation. The exit condition is a race the mesh does not
|
|
control, the operator cannot see, and nothing reports; the same four actions on a slower machine, or
|
|
a larger store, would not have found it. An outage that clears itself after twenty-five minutes and
|
|
five restarts of the forge, the directory and mail is not a smaller fault than one that does not — it
|
|
is the same fault, harder to catch.
|
|
|
|
## What it is not
|
|
|
|
- Not the operator's four actions on the other machine. Those explain the first passes
|
|
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)); they were
|
|
finished seventeen minutes and three full passes before these measurements.
|
|
- Not a file that keeps changing. `/etc/hosts`, the bus's account list and the vault's export were
|
|
hashed across passes and are byte-identical, and the directory reported `mode 755 to 700` every
|
|
pass while already being `700`. Those resources are **misreported as changed** and are worth their
|
|
own question, but they are not what moves a container's identity.
|
|
- Not the lost report alone. A report that cannot be delivered explains a re-apply; it does not
|
|
explain a re-apply that finds 327 differences.
|
|
|
|
## Why it matters beyond this outage
|
|
|
|
The same `continue` makes every routed name in the mesh conditional on every node's plan composing at
|
|
the moment any machine is pushed to. One unreachable or half-migrated machine is enough to withdraw
|
|
its names from everywhere — and the withdrawal is indistinguishable, on the receiving machine, from
|
|
the operator having removed them.
|
|
|
|
The codebase already states the rule this breaks, forty lines away, about the same kind of lookup:
|
|
|
|
> A lookup failure is an error, never "not found": collapsing the two composed a declaration without
|
|
> the trust whenever the inventory hiccuped, delivered by a push that reported success.
|
|
|
|
## How it was fixed, and how the fix is checked
|
|
|
|
`planFor` now marks the two failures that really are a statement about the node — its set not
|
|
composing, and a setting that reaches nothing — and the three gatherers pass over those and only
|
|
those. Every other failure is raised, naming the machine and the read.
|
|
|
|
Five tests hold it: a set that cannot compose is marked as the node's own; a store that cannot be
|
|
read is *not*; one incoherent node still does not cost the rest their names; a roster is never
|
|
returned beside an error; and the raised failure names what could not be read.
|
|
|
|
The two sibling gatherers were audited and fixed the same way — the grant composer, which would have
|
|
withheld a consumer's credential, and the private-network membership, which would have taken a
|
|
machine off the overlay. Three other `planFor` callers were audited and left alone: they refuse or
|
|
report rather than silently withdraw, which is the safe direction.
|
|
|
|
**This does not close [issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md).**
|
|
A roster that changes for a real reason still replaces every container in the mesh. This removes the
|
|
false reasons; whether the roster belongs in a container's identity at all is that record's question.
|