ace numbered its two records 147 and 148, which this repository already uses; they become 150 and 151, as 127 became 149. The new record is why the control node could not stop applying: the roster alternates between two values because routeNamesInTheMesh swallows a per-node plan failure, and the roster is part of every container's identity. The loop closes through the control plane's own store, which each pass replaces.
3.4 KiB
status, opened, located-in, fixed-by
| status | opened | located-in | fixed-by | ||
|---|---|---|---|---|---|
| open | 2026-09-29 |
|
151 — A new name recreates every container in the mesh
What was observed
Migrating one small module on ace (searxng) took four routine controller actions: node public-domain ace zurag.be, assign ace searxng + push, unassign ace searxng + push (a rollback, see
issue 150), and assign + take
- push again. Each push to ace also pushed novox ("this push left g14, novox, shanks behind … sending it too"). novox's host then replaced every container it runs, twice:
22:46 … updated postgres.server (mesh-store): replaced; a container's configuration is fixed when it is created
22:46 … updated distribution.store (mesh-registry): replaced; …
22:46 … updated route-proxy.server (route-proxy): recreated …
22:51 … updated postgres.server (mesh-store): replaced; …
22:52 … updated gitea.server (gitea): replaced; …
22:52–22:55 mesh-vault, builder, all of mailu, keycloak, minio, mongodb, invoicing, photos, umami, …
Replacements per minute on novox, 22:46–22:55: 12, 8, 13, 5, 6, 8, 8, 8, 8, 4. The control plane was
unreachable twice while its own store came back through crash recovery
(FATAL: the database system is starting up), and dependants (umami) crash-looped until it did. None
of the replaced containers belonged to the module being migrated, or to ace.
Why (confirmed part)
Every container the mesh runs is given the mesh's names as --add-host entries
(withMeshNames), and the host puts every entry into the container's spec digest — deliberately,
since issue 135: a container left alone when the roster moved kept a five-day-old
address and restarted 2286 times. A running container cannot have its hosts changed, so a changed
digest means a replace.
The consequence is that the roster is part of every container everywhere: anything that adds, removes or moves one name — a node's public domain, a routed name, an unassign — replaces every container on every machine that carries the list. On the hub that includes the control plane's store, the registry, the edge and mail.
Not yet established
Which name moved in each pass. gitea's hosts after the fact list the .internal names and novox's
internally routed *.novox.be names, and not searxng.zurag.be — so it is not simply "a routed
name was added". Two passes suggest the set changed at assign and changed back at unassign; the
declarations before and after would say, and nothing on the machine records the previous one.
Why it matters now
The node-by-node migration adds names one module at a time. On ace alone that is ~25 web modules; if each assignment moves the roster, each is a full restart of every hub service, and a rollback is another. The migration is paused on this.
What would be right (for diagnosis)
Keep 135's guarantee (no container runs with a stale address) without making the roster part of every container's identity — e.g. resolve mesh names through a resolver the container asks at lookup time rather than baked entries, or scope each container's entries to the names it actually binds.