Both found migrating the first module on ace (searxng), adopted with route-adapter: 147 — assign is not a no-op for a routed module: its route reaches the predecessor's proxy while its container is held, and the name serves a dead backend until take. 148 — every push that moved the mesh's names replaced every container on novox, twice, the control plane's own store included.
2.5 KiB
status, opened, located-in, fixed-by
| status | opened | located-in | fixed-by | |
|---|---|---|---|---|
| open | 2026-09-29 |
|
147 — A route is contributed before its module is taken
What was observed
ace is adopted and runs route-adapter beside the predecessor's traefik (ADR 0104). The first web
module migrated there was searxng. assign ace searxng + push ace — the step that is supposed to
change nothing on an adopted machine (ADR 0100, hq 125: assign holds, take replaces) — took
searxng.zurag.be down:
https://searxng.zurag.be/ → 502
for about five minutes, until the module was unassigned again.
Why
Assign held everything it found on the machine — the predecessor's searxng container, its
directories — exactly as designed. But the module's route contribution is not a resource on the
machine, so nothing held it: it reached route-adapter at once, which wrote
mesh-searxng.zurag.be.yml into traefik's file provider pointing at the mesh's assigned machine port
(http://ace.internal:20000) — where nothing listened, because the container that would was held.
traefik's file router for the name then won over the predecessor's docker-label router for the same
name, and the name served a dead backend.
On this occasion the window was lengthened by an unrelated failure (the machine's docker address pools were exhausted, so the module's network could not be created), but the fault does not depend on it: between assign and take, every routed module's public name points at a backend the mesh has deliberately not started. The runbook's §6 step 2 ("check the preparation — HAL's service still serves") is false for every routed module on a node running route-adapter.
What the operator did
Rolled back per runbook (unassign before take), then re-ran as assign → push → wait for the node to report its holds → take → push, keeping the window to the ~10 s between the two pushes. A workaround, not a fix: a module whose data must move between assign and take (runbook §6 step 3) cannot shrink its window this way.
What would be right
A contribution from a module that is assigned but not taken on an adopted node should be held like the module's resources are — withheld from the provider until take — or the provider should be told the contributor is held and leave the predecessor's route alone. Either keeps "assign changes nothing" true for routed modules.