Files
hq/04-ISSUES/147-a-route-is-contributed-before-its-module-is-taken/00-report.md
T
jschoubben 3bd6f34de3 Issues 147 and 148: a route before its module is taken; a new name recreates every container
Both found migrating the first module on ace (searxng), adopted with route-adapter:
147 — assign is not a no-op for a routed module: its route reaches the predecessor's
proxy while its container is held, and the name serves a dead backend until take.
148 — every push that moved the mesh's names replaced every container on novox, twice,
the control plane's own store included.
2026-09-29 23:01:51 +02:00

2.5 KiB

status, opened, located-in, fixed-by
status opened located-in fixed-by
open 2026-09-29
mesh-controller internal/catalogue/declaration.go (contributions are emitted for an assigned module, taken or not)

147 — A route is contributed before its module is taken

What was observed

ace is adopted and runs route-adapter beside the predecessor's traefik (ADR 0104). The first web module migrated there was searxng. assign ace searxng + push ace — the step that is supposed to change nothing on an adopted machine (ADR 0100, hq 125: assign holds, take replaces) — took searxng.zurag.be down:

https://searxng.zurag.be/  →  502

for about five minutes, until the module was unassigned again.

Why

Assign held everything it found on the machine — the predecessor's searxng container, its directories — exactly as designed. But the module's route contribution is not a resource on the machine, so nothing held it: it reached route-adapter at once, which wrote mesh-searxng.zurag.be.yml into traefik's file provider pointing at the mesh's assigned machine port (http://ace.internal:20000) — where nothing listened, because the container that would was held. traefik's file router for the name then won over the predecessor's docker-label router for the same name, and the name served a dead backend.

On this occasion the window was lengthened by an unrelated failure (the machine's docker address pools were exhausted, so the module's network could not be created), but the fault does not depend on it: between assign and take, every routed module's public name points at a backend the mesh has deliberately not started. The runbook's §6 step 2 ("check the preparation — HAL's service still serves") is false for every routed module on a node running route-adapter.

What the operator did

Rolled back per runbook (unassign before take), then re-ran as assign → push → wait for the node to report its holds → take → push, keeping the window to the ~10 s between the two pushes. A workaround, not a fix: a module whose data must move between assign and take (runbook §6 step 3) cannot shrink its window this way.

What would be right

A contribution from a module that is assigned but not taken on an adopted node should be held like the module's resources are — withheld from the provider until take — or the provider should be told the contributor is held and leave the predecessor's route alone. Either keeps "assign changes nothing" true for routed modules.