Answers issue 151. Copying the roster into each container made the roster part of each container's identity, so one name moving replaced every container in the mesh — and it never stopped the staleness it was for, since a copy taken at creation is stale the moment the roster moves (109, 135). A container resolves through its machine's resolver instead, and nothing is copied. Staleness stops being possible rather than detected, and a name's blast radius becomes nothing. Scoping each container to the names it binds was the close call and is rejected: it contradicts anything-calls-anything, and leaves the roster in the digest so the churn returns for a widely-bound name. Gated on issue 110 — a container on the runtime's default network has no DNS at all today. Removing the copy first reintroduces 109 and 135 silently on a live mesh. 151 stays open until the code lands; design 08's file-not-resolver passage is narrowed to the machine's own roster.
5.4 KiB
status, opened, located-in, fixed-by
| status | opened | located-in | fixed-by | ||
|---|---|---|---|---|---|
| open | 2026-09-29 |
|
151 — A new name recreates every container in the mesh
What was observed
Migrating one small module on ace (searxng) took four routine controller actions: node public-domain ace zurag.be, assign ace searxng + push, unassign ace searxng + push (a rollback, see
issue 150), and assign + take
- push again. Each push to ace also pushed novox ("this push left g14, novox, shanks behind … sending it too"). novox's host then replaced every container it runs, twice:
22:46 … updated postgres.server (mesh-store): replaced; a container's configuration is fixed when it is created
22:46 … updated distribution.store (mesh-registry): replaced; …
22:46 … updated route-proxy.server (route-proxy): recreated …
22:51 … updated postgres.server (mesh-store): replaced; …
22:52 … updated gitea.server (gitea): replaced; …
22:52–22:55 mesh-vault, builder, all of mailu, keycloak, minio, mongodb, invoicing, photos, umami, …
Replacements per minute on novox, 22:46–22:55: 12, 8, 13, 5, 6, 8, 8, 8, 8, 4. The control plane was
unreachable twice while its own store came back through crash recovery
(FATAL: the database system is starting up), and dependants (umami) crash-looped until it did. None
of the replaced containers belonged to the module being migrated, or to ace.
Why (confirmed part)
Every container the mesh runs is given the mesh's names as --add-host entries
(withMeshNames), and the host puts every entry into the container's spec digest — deliberately,
since issue 135: a container left alone when the roster moved kept a five-day-old
address and restarted 2286 times. A running container cannot have its hosts changed, so a changed
digest means a replace.
The consequence is that the roster is part of every container everywhere: anything that adds, removes or moves one name — a node's public domain, a routed name, an unassign — replaces every container on every machine that carries the list. On the hub that includes the control plane's store, the registry, the edge and mail.
Not yet established
Which name moved in each pass. gitea's hosts after the fact list the .internal names and novox's
internally routed *.novox.be names, and not searxng.zurag.be — so it is not simply "a routed
name was added". Two passes suggest the set changed at assign and changed back at unassign; the
declarations before and after would say, and nothing on the machine records the previous one.
Why it matters now
The node-by-node migration adds names one module at a time. On ace alone that is ~25 web modules; if each assignment moves the roster, each is a full restart of every hub service, and a rollback is another. The migration is paused on this.
What has since been ruled out as a cause (2026-09-30)
The roster was also moving for a reason that was not a name changing at all: a node whose plan would not compose had its routed names silently dropped from the roster handed to every machine, so a briefly unreachable store withdrew and restored a name on alternating passes. That is issue 152, and it is fixed — it is what made the churn on the control node repeat every six minutes for twenty-five minutes rather than once per operator action.
It does not close this record. A roster that changes for a real reason — a machine joining, a module assigned, a public domain set — still replaces every container on every machine that carries the list. 152 removed the false reasons; the question below is still open.
What would be right (for diagnosis)
Keep 135's guarantee (no container runs with a stale address) without making the roster part of every container's identity — e.g. resolve mesh names through a resolver the container asks at lookup time rather than baked entries, or scope each container's entries to the names it actually binds.
Answered (2026-09-30): the first of those two
ADR 0148 takes the resolver, not the scoping. Scoping was the close call and was rejected for contradicting the standing intent that anything on the mesh can call anything on it: it would make a name reachable only where the mesh was told in advance it would be wanted, and it would leave the roster in the digest, so the churn returns whenever a widely-bound name moves.
Resolving at lookup time makes staleness impossible rather than detected — 109 and 135 stop being bugs that were fixed and become a shape that does not exist — and makes a name's blast radius nothing.
This record stays open, because the record answers it and the code does not. Nothing may stop copying names until a container can reach the resolver from any of the runtime's networks (issue 110), which it cannot on two of four machines today. Removing the copy first reintroduces 109 and 135 silently, on a live mesh, which is how both were found. The order is in the record.