Issue 110's cause was not the filter: the runtime had never been told, and the resolver dropped a query arriving on a bridge. ADR 0148 step 3 landed once it did (109, 151 resolved). ADR 0151 composes a route's internal name under the serving node and drops the suffixed alias (139, 157 resolved). Design 08 amended; a fact in 0148 corrected.
107 lines
6.3 KiB
Markdown
107 lines
6.3 KiB
Markdown
---
|
||
status: resolved
|
||
opened: 2026-09-29
|
||
located-in:
|
||
- mesh-host internal/apply/apply.go (containerSpecReading hashes every `host` entry)
|
||
- mesh-controller internal/catalogue/declaration.go (withMeshNames gives every container the mesh's names)
|
||
fixed-by: mesh-controller PR 161 — the roster left every container's declaration and so its digest (ADR 0148 step 3, 2026-09-30)
|
||
---
|
||
|
||
# 151 — A new name recreates every container in the mesh
|
||
|
||
## What was observed
|
||
|
||
Migrating one small module on ace (searxng) took four routine controller actions: `node public-domain
|
||
ace zurag.be`, `assign ace searxng` + push, `unassign ace searxng` + push (a rollback, see
|
||
[issue 150](../150-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take
|
||
+ push again. Each push to ace also pushed novox ("this push left g14, novox, shanks behind … sending it
|
||
too"). novox's host then **replaced every container it runs, twice**:
|
||
|
||
```
|
||
22:46 … updated postgres.server (mesh-store): replaced; a container's configuration is fixed when it is created
|
||
22:46 … updated distribution.store (mesh-registry): replaced; …
|
||
22:46 … updated route-proxy.server (route-proxy): recreated …
|
||
22:51 … updated postgres.server (mesh-store): replaced; …
|
||
22:52 … updated gitea.server (gitea): replaced; …
|
||
22:52–22:55 mesh-vault, builder, all of mailu, keycloak, minio, mongodb, invoicing, photos, umami, …
|
||
```
|
||
|
||
Replacements per minute on novox, 22:46–22:55: 12, 8, 13, 5, 6, 8, 8, 8, 8, 4. The control plane was
|
||
unreachable twice while its own store came back through crash recovery
|
||
(`FATAL: the database system is starting up`), and dependants (umami) crash-looped until it did. None
|
||
of the replaced containers belonged to the module being migrated, or to ace.
|
||
|
||
## Why (confirmed part)
|
||
|
||
Every container the mesh runs is given the mesh's names as `--add-host` entries
|
||
(`withMeshNames`), and the host puts every entry into the container's spec digest — deliberately,
|
||
since [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md): a container left alone when the roster moved kept a five-day-old
|
||
address and restarted 2286 times. A running container cannot have its hosts changed, so a changed
|
||
digest means a replace.
|
||
|
||
The consequence is that **the roster is part of every container everywhere**: anything that adds,
|
||
removes or moves one name — a node's public domain, a routed name, an unassign — replaces every
|
||
container on every machine that carries the list. On the hub that includes the control plane's store,
|
||
the registry, the edge and mail.
|
||
|
||
## Not yet established
|
||
|
||
Which name moved in each pass. gitea's hosts after the fact list the `.internal` names and novox's
|
||
internally routed `*.novox.be` names, and **not** `searxng.zurag.be` — so it is not simply "a routed
|
||
name was added". Two passes suggest the set changed at assign and changed back at unassign; the
|
||
declarations before and after would say, and nothing on the machine records the previous one.
|
||
|
||
## Why it matters now
|
||
|
||
The node-by-node migration adds names one module at a time. On ace alone that is ~25 web modules; if
|
||
each assignment moves the roster, each is a full restart of every hub service, and a rollback is
|
||
another. The migration is paused on this.
|
||
|
||
## What has since been ruled out as a cause (2026-09-30)
|
||
|
||
The roster was also moving for a reason that was not a name changing at all: a node whose plan would
|
||
not compose had its routed names silently dropped from the roster handed to every machine, so a
|
||
briefly unreachable store withdrew and restored a name on alternating passes. That is
|
||
[issue 152](../152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md), and it is
|
||
fixed — it is what made the churn on the control node repeat every six minutes for twenty-five
|
||
minutes rather than once per operator action.
|
||
|
||
**It does not close this record.** A roster that changes for a real reason — a machine joining, a
|
||
module assigned, a public domain set — still replaces every container on every machine that carries
|
||
the list. 152 removed the false reasons; the question below is still open.
|
||
|
||
## What would be right (for diagnosis)
|
||
|
||
Keep 135's guarantee (no container runs with a stale address) without making the roster part of every
|
||
container's identity — e.g. resolve mesh names through a resolver the container asks at lookup time
|
||
rather than baked entries, or scope each container's entries to the names it actually binds.
|
||
|
||
## Answered (2026-09-30): the first of those two
|
||
|
||
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) takes
|
||
the resolver, not the scoping. Scoping was the close call and was rejected for contradicting the
|
||
standing intent that anything on the mesh can call anything on it: it would make a name reachable only
|
||
where the mesh was told in advance it would be wanted, and it would leave the roster in the digest, so
|
||
the churn returns whenever a widely-bound name moves.
|
||
|
||
Resolving at lookup time makes staleness impossible rather than detected — 109 and 135 stop being bugs
|
||
that were fixed and become a shape that does not exist — and makes a name's blast radius nothing.
|
||
|
||
**This record stays open**, because the record answers it and the code does not. Nothing may stop
|
||
copying names until a container can reach the resolver from any of the runtime's networks
|
||
([issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)),
|
||
which it cannot on two of four machines today. Removing the copy first reintroduces 109 and 135
|
||
silently, on a live mesh, which is how both were found. The order is in the record.
|
||
|
||
## Resolved (2026-09-30)
|
||
|
||
Step 3 landed the day 110 did. mesh-controller PR 161 stops writing the roster into any container, so
|
||
a container's digest no longer carries a name that is not its own. The controller's tests hold the
|
||
record's check — a container's declaration does not move when the mesh's roster does, and does move
|
||
when the module's own declared entries do.
|
||
|
||
The first push after the change recreated every container once, because every digest lost its host
|
||
entries at the same moment. That was the last such event: from here a name added or moved on one
|
||
machine changes no container anywhere, and the record's second check — add a routed name, watch every
|
||
other machine's apply report say nothing changed — is what the next module assignment will show.
|