Merge pull request 'Issue 152: a node the mesh could not read withdrew its names from every machine' (#191) from fix/152-a-lookup-failure-is-not-an-absence into main
This commit was merged in pull request #191.
This commit is contained in:
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: open
|
||||
status: resolved
|
||||
opened: 2026-09-24
|
||||
located-in: [mesh-controller module.json, mesh-host internal/apply]
|
||||
fixed-by:
|
||||
fixed-by: ADR 0142 — the mesh's own components are delivered as binaries on the machine, so the controller is a process; the delivery itself is issue 142
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -84,3 +84,21 @@ restart and run-to-completion semantics — so this would not need host-side wor
|
||||
|
||||
The two do not collapse into one. The controller is not a code-carrying sidecar, and `network: host`
|
||||
is what makes the asymmetry visible here and nowhere else.
|
||||
|
||||
## Answered
|
||||
|
||||
*2026-09-29, in a grooming pass.* This asked a question rather than reporting a defect, and the
|
||||
question was taken: [ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md)
|
||||
decides that the mesh's own components — the host, the controller, the catalogue, the builder, the
|
||||
vault — are **binaries on the machine**, delivered by the mechanism
|
||||
[ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) built, and that
|
||||
third-party software (the store, the registry, the broker) stays a container because an image is the
|
||||
right way to carry somebody else's build.
|
||||
|
||||
So the operating experience this record was written from — every mutating command reached through
|
||||
`docker exec mesh-controller` — is answered, and answered against the container.
|
||||
|
||||
**The delivery is a separate matter and is not this record's.** Step 1 of it is built and no
|
||||
component travels yet; that is
|
||||
[issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md).
|
||||
|
||||
|
||||
@@ -16,7 +16,7 @@ found that the host's `remove` path stops every `service` resource that is no lo
|
||||
delete". `store.Orphans` matches by id alone. So any of these stops the unit:
|
||||
|
||||
- the module is unassigned — by mistake, or to switch it for another;
|
||||
- the node is sent a deliberately-empty declaration ([issue 127](../127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md));
|
||||
- the node is sent a deliberately-empty declaration ([issue 149](../149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md));
|
||||
- a later catalogue version renames the resource's `id`.
|
||||
|
||||
That is right for a service the mesh brought into being. It is wrong for a unit the mesh
|
||||
|
||||
@@ -0,0 +1,37 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-29
|
||||
located-in: [mesh-controller cmd/mesh-controller]
|
||||
---
|
||||
|
||||
# 148 — a manifest outside this catalogue has no check
|
||||
|
||||
## What was observed
|
||||
|
||||
A module's manifest is validated by a **test** — `internal/catalogue`'s suite parses every manifest
|
||||
in the catalogue checkout beside it and fails on one it cannot resolve. That works, and it is how
|
||||
several real faults were caught before a machine saw them.
|
||||
|
||||
It is available to exactly one repository: this one. Somebody describing their own application in
|
||||
their own repository — the case
|
||||
[ADR 0037](../../02-DECISIONS/0037-where-a-module-lives.md) calls *the case that matters most* —
|
||||
has no check at all. They write a manifest, register it with a running mesh, and find out whether
|
||||
it is valid when the mesh refuses it, or later, when a machine applies something that resolved and
|
||||
should not have.
|
||||
|
||||
The same record asks for the answer: **a `module check` command on the control plane's binary**, so
|
||||
a manifest is checked by the tool rather than by a test that imports the tool's internals.
|
||||
|
||||
## What would have prevented it
|
||||
|
||||
Nothing prevents this; it was noticed and left. ADR 0037 named it on 2026-09-01 and the record sat
|
||||
`proposed` until 2026-09-29, so the missing half was never anybody's task.
|
||||
|
||||
## Evidence to carry into diagnosis
|
||||
|
||||
- `mesh-controller/internal/catalogue` — `ParseManifest` and `CatalogueProblems` are the check, and
|
||||
both are internal.
|
||||
- The catalogue-wide test is `TestEveryCatalogueManifestDeclaresWhatItMounts` and its siblings; they
|
||||
take a path from `MESH_CATALOG`, so the mechanism is already path-driven and not repository-bound.
|
||||
- `mesh-controller module add` refuses a bad manifest at registration, which is the same check far
|
||||
too late: by then it is in a running mesh's records.
|
||||
+22
-3
@@ -1,10 +1,15 @@
|
||||
---
|
||||
status: located
|
||||
status: resolved
|
||||
opened: 2026-09-27
|
||||
located-in: [mesh-controller cmd/mesh-controller/push.go]
|
||||
located-in: [mesh-controller cmd/mesh-controller/push.go, mesh-controller cmd/mesh-controller/sendable.go]
|
||||
fixed-by: mesh-controller sendable.go and push.go — an empty declaration is sent carrying `owns_nothing`, and the host refuses an empty body that does not carry it
|
||||
---
|
||||
|
||||
# A declaration that shrinks to empty is skipped, so the node keeps what it should drop
|
||||
# 149 — a declaration that shrinks to empty is skipped, so the node keeps what it should drop
|
||||
|
||||
*Opened as 127 and renumbered on 2026-09-29: two records were given that number on the same day.*
|
||||
*The other kept it, because three documents and three source files cite it by number and nothing
|
||||
cited this one but a decision and a sibling issue, both corrected with this move.*
|
||||
|
||||
## What was observed
|
||||
|
||||
@@ -37,3 +42,17 @@ mean "own nothing", which the host already applies correctly when it receives on
|
||||
|
||||
On ace, one command drops it permanently (the corrected controller never re-composes it):
|
||||
`sudo ufw delete allow 5671`. At ace's converge it would clear on its own.
|
||||
|
||||
## Closed
|
||||
|
||||
*2026-09-29, in a grooming pass.* Both halves are on `main` and both name this issue.
|
||||
|
||||
- The control plane **sends** it: a declaration that composes to no resources goes out with
|
||||
`owns_nothing`, and `push` says *sent, not skipped*.
|
||||
- The host **refuses an empty body that does not carry it**, so a truncated or mis-composed
|
||||
declaration can never be read as "own nothing" — which is the failure the fix had to avoid while
|
||||
making the empty case expressible.
|
||||
|
||||
Closed by reading the code rather than by watching a machine let go of a stray resource; the record
|
||||
says so rather than implying a run.
|
||||
|
||||
+1
-1
@@ -6,7 +6,7 @@ located-in:
|
||||
fixed-by:
|
||||
---
|
||||
|
||||
# 147 — A route is contributed before its module is taken
|
||||
# 150 — A route is contributed before its module is taken
|
||||
|
||||
## What was observed
|
||||
|
||||
+2
-2
@@ -7,13 +7,13 @@ located-in:
|
||||
fixed-by:
|
||||
---
|
||||
|
||||
# 148 — A new name recreates every container in the mesh
|
||||
# 151 — A new name recreates every container in the mesh
|
||||
|
||||
## What was observed
|
||||
|
||||
Migrating one small module on ace (searxng) took four routine controller actions: `node public-domain
|
||||
ace zurag.be`, `assign ace searxng` + push, `unassign ace searxng` + push (a rollback, see
|
||||
[issue 147](../147-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take
|
||||
[issue 150](../150-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take
|
||||
+ push again. Each push to ace also pushed novox ("this push left g14, novox, shanks behind … sending it
|
||||
too"). novox's host then **replaced every container it runs, twice**:
|
||||
|
||||
@@ -0,0 +1,124 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- mesh-controller cmd/mesh-controller/plan.go (routeNamesInTheMesh skips a node whose plan will not compose)
|
||||
fixed-by: mesh-control fix/152-a-lookup-failure-is-not-an-absence
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 152 — A node whose plan will not compose silently removes its names from every machine
|
||||
|
||||
## What was observed
|
||||
|
||||
For twenty-five minutes, and for seventeen of them after the last operator action, the control-node's
|
||||
host applied all 327 of its resources every ~6.5 minutes without pause, replacing every container on
|
||||
the machine each time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail,
|
||||
and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine
|
||||
sat near 8. Nothing was converging: each pass ended and the next began four seconds later.
|
||||
|
||||
The two machines carrying no containers were not churning. They were only knocked off the bus each
|
||||
time the control node re-created it, reconnected, re-heard the same declaration and applied it again
|
||||
as a no-op.
|
||||
|
||||
## Why: the roster alternates between two values, and it is part of every container
|
||||
|
||||
Two consecutive declarations were compared by reading the `--add-host` entries of four containers
|
||||
the moment each pass created them:
|
||||
|
||||
```
|
||||
23:02 ace.internal drive.novox.be g14.internal keycloak.novox.be novox.internal
|
||||
office.novox.be portainer.novox.be shanks.internal umami.novox.be (9 names)
|
||||
23:07 … the same nine, and searxng.zurag.be (10 names)
|
||||
```
|
||||
|
||||
One routed name — belonging to a module on another machine entirely — leaves the roster and comes
|
||||
back. Because the roster is part of every container's spec digest ([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)),
|
||||
each flip is a different identity for every container on the machine, and a running container cannot
|
||||
have its hosts changed. So every flip replaces all of them.
|
||||
|
||||
**What makes it flip is a swallowed error.** `routeNamesInTheMesh` composes every node's plan to
|
||||
find the names it serves, and when one will not compose it moves on:
|
||||
|
||||
```
|
||||
plan, settings, err := planFor(ctx, open, n.Name)
|
||||
if err != nil {
|
||||
continue
|
||||
}
|
||||
```
|
||||
|
||||
A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does
|
||||
not know", but "the mesh states these names do not exist", to every machine at once.
|
||||
|
||||
## Why it sustains itself
|
||||
|
||||
The loop closes through the control plane's own database:
|
||||
|
||||
1. An apply replaces `mesh-store` — the store the control plane reads — by removing the container,
|
||||
so postgres comes back through crash recovery.
|
||||
2. While it recovers it refuses connections: `FATAL: the database system is not yet accepting
|
||||
connections / Consistent recovery state has not been yet reached` (observed, 21:08:39 UTC, every
|
||||
pass).
|
||||
3. `planFor` for the other machine fails against that store. `routeNamesInTheMesh` swallows it and
|
||||
drops its routed name.
|
||||
4. The roster changed, so all 327 resources differ, so all are replaced — including `mesh-store`,
|
||||
and including the bus, which is why the host also cannot report: `applied, and could not tell the
|
||||
mesh: reporting: nats: connection closed`.
|
||||
5. Back to 1.
|
||||
|
||||
Every pass destroys the evidence the next pass needs to decide it has nothing to do, and nothing
|
||||
outside the machine has to be wrong for it to continue.
|
||||
|
||||
**It is metastable, not permanent.** It ran from 22:46 to 23:11 — five full replacements of every
|
||||
container on the machine — and then stopped on its own, when one pass happened to read the store
|
||||
during a window it was up, composed the same roster twice running, and found nothing to do. Load fell
|
||||
from 7.8 to 1.7 and the machine returned to its five-minute idle tick.
|
||||
|
||||
That it ends by luck is the point, not a mitigation. The exit condition is a race the mesh does not
|
||||
control, the operator cannot see, and nothing reports; the same four actions on a slower machine, or
|
||||
a larger store, would not have found it. An outage that clears itself after twenty-five minutes and
|
||||
five restarts of the forge, the directory and mail is not a smaller fault than one that does not — it
|
||||
is the same fault, harder to catch.
|
||||
|
||||
## What it is not
|
||||
|
||||
- Not the operator's four actions on the other machine. Those explain the first passes
|
||||
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)); they were
|
||||
finished seventeen minutes and three full passes before these measurements.
|
||||
- Not a file that keeps changing. `/etc/hosts`, the bus's account list and the vault's export were
|
||||
hashed across passes and are byte-identical, and the directory reported `mode 755 to 700` every
|
||||
pass while already being `700`. Those resources are **misreported as changed** and are worth their
|
||||
own question, but they are not what moves a container's identity.
|
||||
- Not the lost report alone. A report that cannot be delivered explains a re-apply; it does not
|
||||
explain a re-apply that finds 327 differences.
|
||||
|
||||
## Why it matters beyond this outage
|
||||
|
||||
The same `continue` makes every routed name in the mesh conditional on every node's plan composing at
|
||||
the moment any machine is pushed to. One unreachable or half-migrated machine is enough to withdraw
|
||||
its names from everywhere — and the withdrawal is indistinguishable, on the receiving machine, from
|
||||
the operator having removed them.
|
||||
|
||||
The codebase already states the rule this breaks, forty lines away, about the same kind of lookup:
|
||||
|
||||
> A lookup failure is an error, never "not found": collapsing the two composed a declaration without
|
||||
> the trust whenever the inventory hiccuped, delivered by a push that reported success.
|
||||
|
||||
## How it was fixed, and how the fix is checked
|
||||
|
||||
`planFor` now marks the two failures that really are a statement about the node — its set not
|
||||
composing, and a setting that reaches nothing — and the three gatherers pass over those and only
|
||||
those. Every other failure is raised, naming the machine and the read.
|
||||
|
||||
Five tests hold it: a set that cannot compose is marked as the node's own; a store that cannot be
|
||||
read is *not*; one incoherent node still does not cost the rest their names; a roster is never
|
||||
returned beside an error; and the raised failure names what could not be read.
|
||||
|
||||
The two sibling gatherers were audited and fixed the same way — the grant composer, which would have
|
||||
withheld a consumer's credential, and the private-network membership, which would have taken a
|
||||
machine off the overlay. Three other `planFor` callers were audited and left alone: they refuse or
|
||||
report rather than silently withdraw, which is the safe direction.
|
||||
|
||||
**This does not close [issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md).**
|
||||
A roster that changes for a real reason still replaces every container in the mesh. This removes the
|
||||
false reasons; whether the roster belongs in a container's identity at all is that record's question.
|
||||
Reference in New Issue
Block a user