Issues 147 and 148: a route before its module is taken; a new name recreates every container

Both found migrating the first module on ace (searxng), adopted with route-adapter:
147 — assign is not a no-op for a routed module: its route reaches the predecessor's
proxy while its container is held, and the name serves a dead backend until take.
148 — every push that moved the mesh's names replaced every container on novox, twice,
the control plane's own store included.
This commit is contained in:
2026-09-29 23:01:51 +02:00
parent 0b08cdfce1
commit 3bd6f34de3
2 changed files with 116 additions and 0 deletions
@@ -0,0 +1,52 @@
---
status: open
opened: 2026-09-29
located-in:
- mesh-controller internal/catalogue/declaration.go (contributions are emitted for an assigned module, taken or not)
fixed-by:
---
# 147 — A route is contributed before its module is taken
## What was observed
ace is adopted and runs `route-adapter` beside the predecessor's traefik (ADR 0104). The first web
module migrated there was searxng. `assign ace searxng` + `push ace` — the step that is supposed to
change nothing on an adopted machine (ADR 0100, hq 125: assign holds, take replaces) — took
`searxng.zurag.be` down:
```
https://searxng.zurag.be/ → 502
```
for about five minutes, until the module was unassigned again.
## Why
Assign held everything it found on the machine — the predecessor's `searxng` container, its
directories — exactly as designed. But the module's **route contribution** is not a resource on the
machine, so nothing held it: it reached `route-adapter` at once, which wrote
`mesh-searxng.zurag.be.yml` into traefik's file provider pointing at the mesh's assigned machine port
(`http://ace.internal:20000`) — where nothing listened, because the container that would was held.
traefik's file router for the name then won over the predecessor's docker-label router for the same
name, and the name served a dead backend.
On this occasion the window was lengthened by an unrelated failure (the machine's docker address
pools were exhausted, so the module's network could not be created), but the fault does not depend on
it: **between assign and take, every routed module's public name points at a backend the mesh has
deliberately not started.** The runbook's §6 step 2 ("check the preparation — HAL's service still
serves") is false for every routed module on a node running route-adapter.
## What the operator did
Rolled back per runbook (unassign before take), then re-ran as assign → push → wait for the node to
report its holds → take → push, keeping the window to the ~10 s between the two pushes. A workaround,
not a fix: a module whose data must move between assign and take (runbook §6 step 3) cannot shrink
its window this way.
## What would be right
A contribution from a module that is assigned but **not taken** on an adopted node should be held like
the module's resources are — withheld from the provider until take — or the provider should be told
the contributor is held and leave the predecessor's route alone. Either keeps "assign changes nothing"
true for routed modules.
@@ -0,0 +1,64 @@
---
status: open
opened: 2026-09-29
located-in:
- mesh-host internal/apply/apply.go (containerSpecReading hashes every `host` entry)
- mesh-controller internal/catalogue/declaration.go (withMeshNames gives every container the mesh's names)
fixed-by:
---
# 148 — A new name recreates every container in the mesh
## What was observed
Migrating one small module on ace (searxng) took four routine controller actions: `node public-domain
ace zurag.be`, `assign ace searxng` + push, `unassign ace searxng` + push (a rollback, see
[issue 147](../147-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take
+ push again. Each push to ace also pushed novox ("this push left g14, novox, shanks behind … sending it
too"). novox's host then **replaced every container it runs, twice**:
```
22:46 … updated postgres.server (mesh-store): replaced; a container's configuration is fixed when it is created
22:46 … updated distribution.store (mesh-registry): replaced; …
22:46 … updated route-proxy.server (route-proxy): recreated …
22:51 … updated postgres.server (mesh-store): replaced; …
22:52 … updated gitea.server (gitea): replaced; …
22:52–22:55 mesh-vault, builder, all of mailu, keycloak, minio, mongodb, invoicing, photos, umami, …
```
Replacements per minute on novox, 22:46–22:55: 12, 8, 13, 5, 6, 8, 8, 8, 8, 4. The control plane was
unreachable twice while its own store came back through crash recovery
(`FATAL: the database system is starting up`), and dependants (umami) crash-looped until it did. None
of the replaced containers belonged to the module being migrated, or to ace.
## Why (confirmed part)
Every container the mesh runs is given the mesh's names as `--add-host` entries
(`withMeshNames`), and the host puts every entry into the container's spec digest — deliberately,
since [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md): a container left alone when the roster moved kept a five-day-old
address and restarted 2286 times. A running container cannot have its hosts changed, so a changed
digest means a replace.
The consequence is that **the roster is part of every container everywhere**: anything that adds,
removes or moves one name — a node's public domain, a routed name, an unassign — replaces every
container on every machine that carries the list. On the hub that includes the control plane's store,
the registry, the edge and mail.
## Not yet established
Which name moved in each pass. gitea's hosts after the fact list the `.internal` names and novox's
internally routed `*.novox.be` names, and **not** `searxng.zurag.be` — so it is not simply "a routed
name was added". Two passes suggest the set changed at assign and changed back at unassign; the
declarations before and after would say, and nothing on the machine records the previous one.
## Why it matters now
The node-by-node migration adds names one module at a time. On ace alone that is ~25 web modules; if
each assignment moves the roster, each is a full restart of every hub service, and a rollback is
another. The migration is paused on this.
## What would be right (for diagnosis)
Keep 135's guarantee (no container runs with a stale address) without making the roster part of every
container's identity — e.g. resolve mesh names through a resolver the container asks at lookup time
rather than baked entries, or scope each container's entries to the names it actually binds.