Files
hq/04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md
T
jschoubben 47909c6b71 The records pointed at branches that no longer exist, and two fixes had no sequel
Three issues named the branch that fixed them, and a branch is deleted
when it merges — so every `fixed-by:` was a pointer that resolved to
nothing by the time anyone followed it. They name commits and pull
requests now, and playbook 03 says to.

Two records were missing the thing a reader arrives for. 146 did not say
that one of its fixes crash-looped the control plane on a running mesh,
which is the whole reason the delivery subject carries the stream and the
raise path was the only one exercised. 151 did not say that 152 removed
the false reasons its roster moved, or that it stays open for the real
ones.

ADR 0080 enumerates what cycle.py enforces and named four things; it
enforces five. A progressive insight names the fifth — the decision
stands, the list had gone stale. The checks README and playbook 03 gained
the same rule, and 155 points at all three.
2026-09-30 00:28:35 +02:00

4.2 KiB
Raw Blame History

status, opened, located-in, fixed-by
status opened located-in fixed-by
open 2026-09-29
mesh-host internal/apply/apply.go (containerSpecReading hashes every `host` entry)
mesh-controller internal/catalogue/declaration.go (withMeshNames gives every container the mesh's names)

151 — A new name recreates every container in the mesh

What was observed

Migrating one small module on ace (searxng) took four routine controller actions: node public-domain ace zurag.be, assign ace searxng + push, unassign ace searxng + push (a rollback, see issue 150), and assign + take

  • push again. Each push to ace also pushed novox ("this push left g14, novox, shanks behind … sending it too"). novox's host then replaced every container it runs, twice:
22:46  … updated postgres.server (mesh-store): replaced; a container's configuration is fixed when it is created
22:46  … updated distribution.store (mesh-registry): replaced; …
22:46  … updated route-proxy.server (route-proxy): recreated …
22:51  … updated postgres.server (mesh-store): replaced; …
22:52  … updated gitea.server (gitea): replaced; …
22:52–22:55  mesh-vault, builder, all of mailu, keycloak, minio, mongodb, invoicing, photos, umami, …

Replacements per minute on novox, 22:46–22:55: 12, 8, 13, 5, 6, 8, 8, 8, 8, 4. The control plane was unreachable twice while its own store came back through crash recovery (FATAL: the database system is starting up), and dependants (umami) crash-looped until it did. None of the replaced containers belonged to the module being migrated, or to ace.

Why (confirmed part)

Every container the mesh runs is given the mesh's names as --add-host entries (withMeshNames), and the host puts every entry into the container's spec digest — deliberately, since issue 135: a container left alone when the roster moved kept a five-day-old address and restarted 2286 times. A running container cannot have its hosts changed, so a changed digest means a replace.

The consequence is that the roster is part of every container everywhere: anything that adds, removes or moves one name — a node's public domain, a routed name, an unassign — replaces every container on every machine that carries the list. On the hub that includes the control plane's store, the registry, the edge and mail.

Not yet established

Which name moved in each pass. gitea's hosts after the fact list the .internal names and novox's internally routed *.novox.be names, and not searxng.zurag.be — so it is not simply "a routed name was added". Two passes suggest the set changed at assign and changed back at unassign; the declarations before and after would say, and nothing on the machine records the previous one.

Why it matters now

The node-by-node migration adds names one module at a time. On ace alone that is ~25 web modules; if each assignment moves the roster, each is a full restart of every hub service, and a rollback is another. The migration is paused on this.

What has since been ruled out as a cause (2026-09-30)

The roster was also moving for a reason that was not a name changing at all: a node whose plan would not compose had its routed names silently dropped from the roster handed to every machine, so a briefly unreachable store withdrew and restored a name on alternating passes. That is issue 152, and it is fixed — it is what made the churn on the control node repeat every six minutes for twenty-five minutes rather than once per operator action.

It does not close this record. A roster that changes for a real reason — a machine joining, a module assigned, a public domain set — still replaces every container on every machine that carries the list. 152 removed the false reasons; the question below is still open.

What would be right (for diagnosis)

Keep 135's guarantee (no container runs with a stale address) without making the roster part of every container's identity — e.g. resolve mesh names through a resolver the container asks at lookup time rather than baked entries, or scope each container's entries to the names it actually binds.