Files
hq/04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md
T
jschoubben 04c9500b5b Group 2 is resolved: containers resolve, nothing is copied, a route's name says where it arrives
Issue 110's cause was not the filter: the runtime had never been told,
and the resolver dropped a query arriving on a bridge. ADR 0148 step 3
landed once it did (109, 151 resolved). ADR 0151 composes a route's
internal name under the serving node and drops the suffixed alias
(139, 157 resolved). Design 08 amended; a fact in 0148 corrected.
2026-09-30 14:56:43 +02:00

107 lines
6.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
status: resolved
opened: 2026-09-29
located-in:
- mesh-host internal/apply/apply.go (containerSpecReading hashes every `host` entry)
- mesh-controller internal/catalogue/declaration.go (withMeshNames gives every container the mesh's names)
fixed-by: mesh-controller PR 161 — the roster left every container's declaration and so its digest (ADR 0148 step 3, 2026-09-30)
---
# 151 — A new name recreates every container in the mesh
## What was observed
Migrating one small module on ace (searxng) took four routine controller actions: `node public-domain
ace zurag.be`, `assign ace searxng` + push, `unassign ace searxng` + push (a rollback, see
[issue 150](../150-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take
+ push again. Each push to ace also pushed novox ("this push left g14, novox, shanks behind … sending it
too"). novox's host then **replaced every container it runs, twice**:
```
22:46 … updated postgres.server (mesh-store): replaced; a container's configuration is fixed when it is created
22:46 … updated distribution.store (mesh-registry): replaced; …
22:46 … updated route-proxy.server (route-proxy): recreated …
22:51 … updated postgres.server (mesh-store): replaced; …
22:52 … updated gitea.server (gitea): replaced; …
22:52–22:55 mesh-vault, builder, all of mailu, keycloak, minio, mongodb, invoicing, photos, umami, …
```
Replacements per minute on novox, 22:46–22:55: 12, 8, 13, 5, 6, 8, 8, 8, 8, 4. The control plane was
unreachable twice while its own store came back through crash recovery
(`FATAL: the database system is starting up`), and dependants (umami) crash-looped until it did. None
of the replaced containers belonged to the module being migrated, or to ace.
## Why (confirmed part)
Every container the mesh runs is given the mesh's names as `--add-host` entries
(`withMeshNames`), and the host puts every entry into the container's spec digest — deliberately,
since [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md): a container left alone when the roster moved kept a five-day-old
address and restarted 2286 times. A running container cannot have its hosts changed, so a changed
digest means a replace.
The consequence is that **the roster is part of every container everywhere**: anything that adds,
removes or moves one name — a node's public domain, a routed name, an unassign — replaces every
container on every machine that carries the list. On the hub that includes the control plane's store,
the registry, the edge and mail.
## Not yet established
Which name moved in each pass. gitea's hosts after the fact list the `.internal` names and novox's
internally routed `*.novox.be` names, and **not** `searxng.zurag.be` — so it is not simply "a routed
name was added". Two passes suggest the set changed at assign and changed back at unassign; the
declarations before and after would say, and nothing on the machine records the previous one.
## Why it matters now
The node-by-node migration adds names one module at a time. On ace alone that is ~25 web modules; if
each assignment moves the roster, each is a full restart of every hub service, and a rollback is
another. The migration is paused on this.
## What has since been ruled out as a cause (2026-09-30)
The roster was also moving for a reason that was not a name changing at all: a node whose plan would
not compose had its routed names silently dropped from the roster handed to every machine, so a
briefly unreachable store withdrew and restored a name on alternating passes. That is
[issue 152](../152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md), and it is
fixed — it is what made the churn on the control node repeat every six minutes for twenty-five
minutes rather than once per operator action.
**It does not close this record.** A roster that changes for a real reason — a machine joining, a
module assigned, a public domain set — still replaces every container on every machine that carries
the list. 152 removed the false reasons; the question below is still open.
## What would be right (for diagnosis)
Keep 135's guarantee (no container runs with a stale address) without making the roster part of every
container's identity — e.g. resolve mesh names through a resolver the container asks at lookup time
rather than baked entries, or scope each container's entries to the names it actually binds.
## Answered (2026-09-30): the first of those two
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) takes
the resolver, not the scoping. Scoping was the close call and was rejected for contradicting the
standing intent that anything on the mesh can call anything on it: it would make a name reachable only
where the mesh was told in advance it would be wanted, and it would leave the roster in the digest, so
the churn returns whenever a widely-bound name moves.
Resolving at lookup time makes staleness impossible rather than detected — 109 and 135 stop being bugs
that were fixed and become a shape that does not exist — and makes a name's blast radius nothing.
**This record stays open**, because the record answers it and the code does not. Nothing may stop
copying names until a container can reach the resolver from any of the runtime's networks
([issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)),
which it cannot on two of four machines today. Removing the copy first reintroduces 109 and 135
silently, on a live mesh, which is how both were found. The order is in the record.
## Resolved (2026-09-30)
Step 3 landed the day 110 did. mesh-controller PR 161 stops writing the roster into any container, so
a container's digest no longer carries a name that is not its own. The controller's tests hold the
record's check — a container's declaration does not move when the mesh's roster does, and does move
when the module's own declared entries do.
The first push after the change recreated every container once, because every digest lost its host
entries at the same moment. That was the last such event: from here a name added or moved on one
machine changes no container anywhere, and the record's second check — add a routed name, watch every
other machine's apply report say nothing changed — is what the next module assignment will show.