Both found migrating the first module on ace (searxng), adopted with route-adapter: 147 — assign is not a no-op for a routed module: its route reaches the predecessor's proxy while its container is held, and the name serves a dead backend until take. 148 — every push that moved the mesh's names replaced every container on novox, twice, the control plane's own store included.
53 lines
2.5 KiB
Markdown
53 lines
2.5 KiB
Markdown
---
|
|
status: open
|
|
opened: 2026-09-29
|
|
located-in:
|
|
- mesh-controller internal/catalogue/declaration.go (contributions are emitted for an assigned module, taken or not)
|
|
fixed-by:
|
|
---
|
|
|
|
# 147 — A route is contributed before its module is taken
|
|
|
|
## What was observed
|
|
|
|
ace is adopted and runs `route-adapter` beside the predecessor's traefik (ADR 0104). The first web
|
|
module migrated there was searxng. `assign ace searxng` + `push ace` — the step that is supposed to
|
|
change nothing on an adopted machine (ADR 0100, hq 125: assign holds, take replaces) — took
|
|
`searxng.zurag.be` down:
|
|
|
|
```
|
|
https://searxng.zurag.be/ → 502
|
|
```
|
|
|
|
for about five minutes, until the module was unassigned again.
|
|
|
|
## Why
|
|
|
|
Assign held everything it found on the machine — the predecessor's `searxng` container, its
|
|
directories — exactly as designed. But the module's **route contribution** is not a resource on the
|
|
machine, so nothing held it: it reached `route-adapter` at once, which wrote
|
|
`mesh-searxng.zurag.be.yml` into traefik's file provider pointing at the mesh's assigned machine port
|
|
(`http://ace.internal:20000`) — where nothing listened, because the container that would was held.
|
|
traefik's file router for the name then won over the predecessor's docker-label router for the same
|
|
name, and the name served a dead backend.
|
|
|
|
On this occasion the window was lengthened by an unrelated failure (the machine's docker address
|
|
pools were exhausted, so the module's network could not be created), but the fault does not depend on
|
|
it: **between assign and take, every routed module's public name points at a backend the mesh has
|
|
deliberately not started.** The runbook's §6 step 2 ("check the preparation — HAL's service still
|
|
serves") is false for every routed module on a node running route-adapter.
|
|
|
|
## What the operator did
|
|
|
|
Rolled back per runbook (unassign before take), then re-ran as assign → push → wait for the node to
|
|
report its holds → take → push, keeping the window to the ~10 s between the two pushes. A workaround,
|
|
not a fix: a module whose data must move between assign and take (runbook §6 step 3) cannot shrink
|
|
its window this way.
|
|
|
|
## What would be right
|
|
|
|
A contribution from a module that is assigned but **not taken** on an adopted node should be held like
|
|
the module's resources are — withheld from the provider until take — or the provider should be told
|
|
the contributor is held and leave the predecessor's route alone. Either keeps "assign changes nothing"
|
|
true for routed modules.
|