--- status: open opened: 2026-09-29 located-in: - mesh-host internal/apply/apply.go (containerSpecReading hashes every `host` entry) - mesh-controller internal/catalogue/declaration.go (withMeshNames gives every container the mesh's names) fixed-by: --- # 151 — A new name recreates every container in the mesh ## What was observed Migrating one small module on ace (searxng) took four routine controller actions: `node public-domain ace zurag.be`, `assign ace searxng` + push, `unassign ace searxng` + push (a rollback, see [issue 150](../150-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take + push again. Each push to ace also pushed novox ("this push left g14, novox, shanks behind … sending it too"). novox's host then **replaced every container it runs, twice**: ``` 22:46 … updated postgres.server (mesh-store): replaced; a container's configuration is fixed when it is created 22:46 … updated distribution.store (mesh-registry): replaced; … 22:46 … updated route-proxy.server (route-proxy): recreated … 22:51 … updated postgres.server (mesh-store): replaced; … 22:52 … updated gitea.server (gitea): replaced; … 22:52–22:55 mesh-vault, builder, all of mailu, keycloak, minio, mongodb, invoicing, photos, umami, … ``` Replacements per minute on novox, 22:46–22:55: 12, 8, 13, 5, 6, 8, 8, 8, 8, 4. The control plane was unreachable twice while its own store came back through crash recovery (`FATAL: the database system is starting up`), and dependants (umami) crash-looped until it did. None of the replaced containers belonged to the module being migrated, or to ace. ## Why (confirmed part) Every container the mesh runs is given the mesh's names as `--add-host` entries (`withMeshNames`), and the host puts every entry into the container's spec digest — deliberately, since [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md): a container left alone when the roster moved kept a five-day-old address and restarted 2286 times. A running container cannot have its hosts changed, so a changed digest means a replace. The consequence is that **the roster is part of every container everywhere**: anything that adds, removes or moves one name — a node's public domain, a routed name, an unassign — replaces every container on every machine that carries the list. On the hub that includes the control plane's store, the registry, the edge and mail. ## Not yet established Which name moved in each pass. gitea's hosts after the fact list the `.internal` names and novox's internally routed `*.novox.be` names, and **not** `searxng.zurag.be` — so it is not simply "a routed name was added". Two passes suggest the set changed at assign and changed back at unassign; the declarations before and after would say, and nothing on the machine records the previous one. ## Why it matters now The node-by-node migration adds names one module at a time. On ace alone that is ~25 web modules; if each assignment moves the roster, each is a full restart of every hub service, and a rollback is another. The migration is paused on this. ## What has since been ruled out as a cause (2026-09-30) The roster was also moving for a reason that was not a name changing at all: a node whose plan would not compose had its routed names silently dropped from the roster handed to every machine, so a briefly unreachable store withdrew and restored a name on alternating passes. That is [issue 152](../152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md), and it is fixed — it is what made the churn on the control node repeat every six minutes for twenty-five minutes rather than once per operator action. **It does not close this record.** A roster that changes for a real reason — a machine joining, a module assigned, a public domain set — still replaces every container on every machine that carries the list. 152 removed the false reasons; the question below is still open. ## What would be right (for diagnosis) Keep 135's guarantee (no container runs with a stale address) without making the roster part of every container's identity — e.g. resolve mesh names through a resolver the container asks at lookup time rather than baked entries, or scope each container's entries to the names it actually binds. ## Answered (2026-09-30): the first of those two [ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) takes the resolver, not the scoping. Scoping was the close call and was rejected for contradicting the standing intent that anything on the mesh can call anything on it: it would make a name reachable only where the mesh was told in advance it would be wanted, and it would leave the roster in the digest, so the churn returns whenever a widely-bound name moves. Resolving at lookup time makes staleness impossible rather than detected — 109 and 135 stop being bugs that were fixed and become a shape that does not exist — and makes a name's blast radius nothing. **This record stays open**, because the record answers it and the code does not. Nothing may stop copying names until a container can reach the resolver from any of the runtime's networks ([issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)), which it cannot on two of four machines today. Removing the copy first reintroduces 109 and 135 silently, on a live mesh, which is how both were found. The order is in the record.