diff --git a/04-ISSUES/147-a-route-is-contributed-before-its-module-is-taken/00-report.md b/04-ISSUES/147-a-route-is-contributed-before-its-module-is-taken/00-report.md new file mode 100644 index 0000000..f9afe62 --- /dev/null +++ b/04-ISSUES/147-a-route-is-contributed-before-its-module-is-taken/00-report.md @@ -0,0 +1,52 @@ +--- +status: open +opened: 2026-09-29 +located-in: + - mesh-controller internal/catalogue/declaration.go (contributions are emitted for an assigned module, taken or not) +fixed-by: +--- + +# 147 — A route is contributed before its module is taken + +## What was observed + +ace is adopted and runs `route-adapter` beside the predecessor's traefik (ADR 0104). The first web +module migrated there was searxng. `assign ace searxng` + `push ace` — the step that is supposed to +change nothing on an adopted machine (ADR 0100, hq 125: assign holds, take replaces) — took +`searxng.zurag.be` down: + +``` +https://searxng.zurag.be/ → 502 +``` + +for about five minutes, until the module was unassigned again. + +## Why + +Assign held everything it found on the machine — the predecessor's `searxng` container, its +directories — exactly as designed. But the module's **route contribution** is not a resource on the +machine, so nothing held it: it reached `route-adapter` at once, which wrote +`mesh-searxng.zurag.be.yml` into traefik's file provider pointing at the mesh's assigned machine port +(`http://ace.internal:20000`) — where nothing listened, because the container that would was held. +traefik's file router for the name then won over the predecessor's docker-label router for the same +name, and the name served a dead backend. + +On this occasion the window was lengthened by an unrelated failure (the machine's docker address +pools were exhausted, so the module's network could not be created), but the fault does not depend on +it: **between assign and take, every routed module's public name points at a backend the mesh has +deliberately not started.** The runbook's §6 step 2 ("check the preparation — HAL's service still +serves") is false for every routed module on a node running route-adapter. + +## What the operator did + +Rolled back per runbook (unassign before take), then re-ran as assign → push → wait for the node to +report its holds → take → push, keeping the window to the ~10 s between the two pushes. A workaround, +not a fix: a module whose data must move between assign and take (runbook §6 step 3) cannot shrink +its window this way. + +## What would be right + +A contribution from a module that is assigned but **not taken** on an adopted node should be held like +the module's resources are — withheld from the provider until take — or the provider should be told +the contributor is held and leave the predecessor's route alone. Either keeps "assign changes nothing" +true for routed modules. diff --git a/04-ISSUES/148-a-new-name-recreates-every-container-in-the-mesh/00-report.md b/04-ISSUES/148-a-new-name-recreates-every-container-in-the-mesh/00-report.md new file mode 100644 index 0000000..8185c6a --- /dev/null +++ b/04-ISSUES/148-a-new-name-recreates-every-container-in-the-mesh/00-report.md @@ -0,0 +1,64 @@ +--- +status: open +opened: 2026-09-29 +located-in: + - mesh-host internal/apply/apply.go (containerSpecReading hashes every `host` entry) + - mesh-controller internal/catalogue/declaration.go (withMeshNames gives every container the mesh's names) +fixed-by: +--- + +# 148 — A new name recreates every container in the mesh + +## What was observed + +Migrating one small module on ace (searxng) took four routine controller actions: `node public-domain +ace zurag.be`, `assign ace searxng` + push, `unassign ace searxng` + push (a rollback, see +[issue 147](../147-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take ++ push again. Each push to ace also pushed novox ("this push left g14, novox, shanks behind … sending it +too"). novox's host then **replaced every container it runs, twice**: + +``` +22:46 … updated postgres.server (mesh-store): replaced; a container's configuration is fixed when it is created +22:46 … updated distribution.store (mesh-registry): replaced; … +22:46 … updated route-proxy.server (route-proxy): recreated … +22:51 … updated postgres.server (mesh-store): replaced; … +22:52 … updated gitea.server (gitea): replaced; … +22:52–22:55 mesh-vault, builder, all of mailu, keycloak, minio, mongodb, invoicing, photos, umami, … +``` + +Replacements per minute on novox, 22:46–22:55: 12, 8, 13, 5, 6, 8, 8, 8, 8, 4. The control plane was +unreachable twice while its own store came back through crash recovery +(`FATAL: the database system is starting up`), and dependants (umami) crash-looped until it did. None +of the replaced containers belonged to the module being migrated, or to ace. + +## Why (confirmed part) + +Every container the mesh runs is given the mesh's names as `--add-host` entries +(`withMeshNames`), and the host puts every entry into the container's spec digest — deliberately, +since [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md): a container left alone when the roster moved kept a five-day-old +address and restarted 2286 times. A running container cannot have its hosts changed, so a changed +digest means a replace. + +The consequence is that **the roster is part of every container everywhere**: anything that adds, +removes or moves one name — a node's public domain, a routed name, an unassign — replaces every +container on every machine that carries the list. On the hub that includes the control plane's store, +the registry, the edge and mail. + +## Not yet established + +Which name moved in each pass. gitea's hosts after the fact list the `.internal` names and novox's +internally routed `*.novox.be` names, and **not** `searxng.zurag.be` — so it is not simply "a routed +name was added". Two passes suggest the set changed at assign and changed back at unassign; the +declarations before and after would say, and nothing on the machine records the previous one. + +## Why it matters now + +The node-by-node migration adds names one module at a time. On ace alone that is ~25 web modules; if +each assignment moves the roster, each is a full restart of every hub service, and a rollback is +another. The migration is paused on this. + +## What would be right (for diagnosis) + +Keep 135's guarantee (no container runs with a stale address) without making the roster part of every +container's identity — e.g. resolve mesh names through a resolver the container asks at lookup time +rather than baked entries, or scope each container's entries to the names it actually binds.