Three issues named the branch that fixed them, and a branch is deleted when it merges — so every `fixed-by:` was a pointer that resolved to nothing by the time anyone followed it. They name commits and pull requests now, and playbook 03 says to. Two records were missing the thing a reader arrives for. 146 did not say that one of its fixes crash-looped the control plane on a running mesh, which is the whole reason the delivery subject carries the stream and the raise path was the only one exercised. 151 did not say that 152 removed the false reasons its roster moved, or that it stays open for the real ones. ADR 0080 enumerates what cycle.py enforces and named four things; it enforces five. A progressive insight names the fifth — the decision stands, the list had gone stale. The checks README and playbook 03 gained the same rule, and 155 points at all three.
78 lines
4.2 KiB
Markdown
78 lines
4.2 KiB
Markdown
---
|
||
status: open
|
||
opened: 2026-09-29
|
||
located-in:
|
||
- mesh-host internal/apply/apply.go (containerSpecReading hashes every `host` entry)
|
||
- mesh-controller internal/catalogue/declaration.go (withMeshNames gives every container the mesh's names)
|
||
fixed-by:
|
||
---
|
||
|
||
# 151 — A new name recreates every container in the mesh
|
||
|
||
## What was observed
|
||
|
||
Migrating one small module on ace (searxng) took four routine controller actions: `node public-domain
|
||
ace zurag.be`, `assign ace searxng` + push, `unassign ace searxng` + push (a rollback, see
|
||
[issue 150](../150-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take
|
||
+ push again. Each push to ace also pushed novox ("this push left g14, novox, shanks behind … sending it
|
||
too"). novox's host then **replaced every container it runs, twice**:
|
||
|
||
```
|
||
22:46 … updated postgres.server (mesh-store): replaced; a container's configuration is fixed when it is created
|
||
22:46 … updated distribution.store (mesh-registry): replaced; …
|
||
22:46 … updated route-proxy.server (route-proxy): recreated …
|
||
22:51 … updated postgres.server (mesh-store): replaced; …
|
||
22:52 … updated gitea.server (gitea): replaced; …
|
||
22:52–22:55 mesh-vault, builder, all of mailu, keycloak, minio, mongodb, invoicing, photos, umami, …
|
||
```
|
||
|
||
Replacements per minute on novox, 22:46–22:55: 12, 8, 13, 5, 6, 8, 8, 8, 8, 4. The control plane was
|
||
unreachable twice while its own store came back through crash recovery
|
||
(`FATAL: the database system is starting up`), and dependants (umami) crash-looped until it did. None
|
||
of the replaced containers belonged to the module being migrated, or to ace.
|
||
|
||
## Why (confirmed part)
|
||
|
||
Every container the mesh runs is given the mesh's names as `--add-host` entries
|
||
(`withMeshNames`), and the host puts every entry into the container's spec digest — deliberately,
|
||
since [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md): a container left alone when the roster moved kept a five-day-old
|
||
address and restarted 2286 times. A running container cannot have its hosts changed, so a changed
|
||
digest means a replace.
|
||
|
||
The consequence is that **the roster is part of every container everywhere**: anything that adds,
|
||
removes or moves one name — a node's public domain, a routed name, an unassign — replaces every
|
||
container on every machine that carries the list. On the hub that includes the control plane's store,
|
||
the registry, the edge and mail.
|
||
|
||
## Not yet established
|
||
|
||
Which name moved in each pass. gitea's hosts after the fact list the `.internal` names and novox's
|
||
internally routed `*.novox.be` names, and **not** `searxng.zurag.be` — so it is not simply "a routed
|
||
name was added". Two passes suggest the set changed at assign and changed back at unassign; the
|
||
declarations before and after would say, and nothing on the machine records the previous one.
|
||
|
||
## Why it matters now
|
||
|
||
The node-by-node migration adds names one module at a time. On ace alone that is ~25 web modules; if
|
||
each assignment moves the roster, each is a full restart of every hub service, and a rollback is
|
||
another. The migration is paused on this.
|
||
|
||
## What has since been ruled out as a cause (2026-09-30)
|
||
|
||
The roster was also moving for a reason that was not a name changing at all: a node whose plan would
|
||
not compose had its routed names silently dropped from the roster handed to every machine, so a
|
||
briefly unreachable store withdrew and restored a name on alternating passes. That is
|
||
[issue 152](../152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md), and it is
|
||
fixed — it is what made the churn on the control node repeat every six minutes for twenty-five
|
||
minutes rather than once per operator action.
|
||
|
||
**It does not close this record.** A roster that changes for a real reason — a machine joining, a
|
||
module assigned, a public domain set — still replaces every container on every machine that carries
|
||
the list. 152 removed the false reasons; the question below is still open.
|
||
|
||
## What would be right (for diagnosis)
|
||
|
||
Keep 135's guarantee (no container runs with a stale address) without making the roster part of every
|
||
container's identity — e.g. resolve mesh names through a resolver the container asks at lookup time
|
||
rather than baked entries, or scope each container's entries to the names it actually binds.
|