ace numbered its two records 147 and 148, which this repository already uses; they become 150 and 151, as 127 became 149. The new record is why the control node could not stop applying: the roster alternates between two values because routeNamesInTheMesh swallows a per-node plan failure, and the roster is part of every container's identity. The loop closes through the control plane's own store, which each pass replaces.
53 lines
2.5 KiB
Markdown
53 lines
2.5 KiB
Markdown
---
|
|
status: open
|
|
opened: 2026-09-29
|
|
located-in:
|
|
- mesh-controller internal/catalogue/declaration.go (contributions are emitted for an assigned module, taken or not)
|
|
fixed-by:
|
|
---
|
|
|
|
# 150 — A route is contributed before its module is taken
|
|
|
|
## What was observed
|
|
|
|
ace is adopted and runs `route-adapter` beside the predecessor's traefik (ADR 0104). The first web
|
|
module migrated there was searxng. `assign ace searxng` + `push ace` — the step that is supposed to
|
|
change nothing on an adopted machine (ADR 0100, hq 125: assign holds, take replaces) — took
|
|
`searxng.zurag.be` down:
|
|
|
|
```
|
|
https://searxng.zurag.be/ → 502
|
|
```
|
|
|
|
for about five minutes, until the module was unassigned again.
|
|
|
|
## Why
|
|
|
|
Assign held everything it found on the machine — the predecessor's `searxng` container, its
|
|
directories — exactly as designed. But the module's **route contribution** is not a resource on the
|
|
machine, so nothing held it: it reached `route-adapter` at once, which wrote
|
|
`mesh-searxng.zurag.be.yml` into traefik's file provider pointing at the mesh's assigned machine port
|
|
(`http://ace.internal:20000`) — where nothing listened, because the container that would was held.
|
|
traefik's file router for the name then won over the predecessor's docker-label router for the same
|
|
name, and the name served a dead backend.
|
|
|
|
On this occasion the window was lengthened by an unrelated failure (the machine's docker address
|
|
pools were exhausted, so the module's network could not be created), but the fault does not depend on
|
|
it: **between assign and take, every routed module's public name points at a backend the mesh has
|
|
deliberately not started.** The runbook's §6 step 2 ("check the preparation — HAL's service still
|
|
serves") is false for every routed module on a node running route-adapter.
|
|
|
|
## What the operator did
|
|
|
|
Rolled back per runbook (unassign before take), then re-ran as assign → push → wait for the node to
|
|
report its holds → take → push, keeping the window to the ~10 s between the two pushes. A workaround,
|
|
not a fix: a module whose data must move between assign and take (runbook §6 step 3) cannot shrink
|
|
its window this way.
|
|
|
|
## What would be right
|
|
|
|
A contribution from a module that is assigned but **not taken** on an adopted node should be held like
|
|
the module's resources are — withheld from the provider until take — or the provider should be told
|
|
the contributor is held and leave the predecessor's route alone. Either keeps "assign changes nothing"
|
|
true for routed modules.
|