ace numbered its two records 147 and 148, which this repository already uses; they become 150 and 151, as 127 became 149. The new record is why the control node could not stop applying: the roster alternates between two values because routeNamesInTheMesh swallows a per-node plan failure, and the roster is part of every container's identity. The loop closes through the control plane's own store, which each pass replaces.
2.5 KiB
status, opened, located-in, fixed-by
| status | opened | located-in | fixed-by | |
|---|---|---|---|---|
| open | 2026-09-29 |
|
150 — A route is contributed before its module is taken
What was observed
ace is adopted and runs route-adapter beside the predecessor's traefik (ADR 0104). The first web
module migrated there was searxng. assign ace searxng + push ace — the step that is supposed to
change nothing on an adopted machine (ADR 0100, hq 125: assign holds, take replaces) — took
searxng.zurag.be down:
https://searxng.zurag.be/ → 502
for about five minutes, until the module was unassigned again.
Why
Assign held everything it found on the machine — the predecessor's searxng container, its
directories — exactly as designed. But the module's route contribution is not a resource on the
machine, so nothing held it: it reached route-adapter at once, which wrote
mesh-searxng.zurag.be.yml into traefik's file provider pointing at the mesh's assigned machine port
(http://ace.internal:20000) — where nothing listened, because the container that would was held.
traefik's file router for the name then won over the predecessor's docker-label router for the same
name, and the name served a dead backend.
On this occasion the window was lengthened by an unrelated failure (the machine's docker address pools were exhausted, so the module's network could not be created), but the fault does not depend on it: between assign and take, every routed module's public name points at a backend the mesh has deliberately not started. The runbook's §6 step 2 ("check the preparation — HAL's service still serves") is false for every routed module on a node running route-adapter.
What the operator did
Rolled back per runbook (unassign before take), then re-ran as assign → push → wait for the node to report its holds → take → push, keeping the window to the ~10 s between the two pushes. A workaround, not a fix: a module whose data must move between assign and take (runbook §6 step 3) cannot shrink its window this way.
What would be right
A contribution from a module that is assigned but not taken on an adopted node should be held like the module's resources are — withheld from the provider until take — or the provider should be told the contributor is held and leave the predecessor's route alone. Either keeps "assign changes nothing" true for routed modules.