The record gets its built note; designs 05 and 09 the revisions; 086, 098, 099, 100 and 101 stay located because every machine is converged and the record's live row — a take read on an adopted machine — has not been run; 090 is built in part, its network difference left for the take to say.
61 lines
3.1 KiB
Markdown
61 lines
3.1 KiB
Markdown
---
|
||
status: resolved
|
||
opened: 2026-09-26
|
||
located-in: [mesh-host internal/apply]
|
||
fixed-by: mesh-host 63 (every written field compared), mesh-controller 201 (build says the policy)
|
||
amended-design:
|
||
---
|
||
|
||
# 126 — a volume path is not in the spec comparison, and a roll-out raced a data move
|
||
|
||
## What was observed
|
||
|
||
Landing the "module data lives in /var/lib" change on novox (mesh-catalog #97), two
|
||
distinct faults surfaced in one hour:
|
||
|
||
1. **Building a module with a roll-out upgrade policy IS deploying it.** gitea's policy
|
||
was roll-out; the `build` that registered its repathed manifest sent it to the node
|
||
immediately, which recreated the container mounting the *not-yet-renamed* (empty)
|
||
`/var/lib/gitea/data`. The forge came back as its own install page, fresh host keys
|
||
and all, and every subsequent pipeline build died on `repository not found` — which
|
||
also blocked the fix, since re-registering the other modules needed the forge. The
|
||
operator narrative "build, then move data, then push" is only safe under the record
|
||
policy; nothing warned that one module in the batch would skip the pause.
|
||
|
||
2. **Changing a container's volume paths does not recreate the container.** After the
|
||
final push, five of the six repathed modules kept their old containers running
|
||
("Up 13–26 hours") — the new declaration's volume paths differ from the running
|
||
containers' mounts, and the apply judged them current. Same class as mesh-host #27
|
||
(`dns`/`ip` absent from the comparison): a field the comparison does not read is a
|
||
field that can never change a running container. Benign here only because a rename
|
||
on one filesystem preserves the mounted inode — the running containers keep serving
|
||
the same bytes the new path names, and the next natural recreation converges. A
|
||
cross-filesystem move, or a path change to *different* data, would have silently
|
||
split the module between two worlds.
|
||
|
||
## What would have prevented it
|
||
|
||
- `build` printing the module's upgrade policy when that policy will act on the result
|
||
("gitea rolls out on build — the node will receive this immediately"), or a
|
||
`--register-only` flag for exactly this choreography.
|
||
- Volumes (and every other container field) in the spec comparison, or the honest
|
||
refusal: "this field changed and I cannot apply it without recreation."
|
||
|
||
## Recovery that worked
|
||
|
||
Instant renames both ways broke the circular dependency (forge needed for builds,
|
||
builds needed for the push, push needed for the forge): data back to the old path,
|
||
old-spec forge started, artifacts rebuilt, data renamed forward, push. Nothing lost;
|
||
the install-page junk was discarded twice.
|
||
|
||
## Decided, 2026-10-01
|
||
|
||
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rules 5 and 7: every field compared; build says the policy. Building follows,
|
||
host first, then the controller's `take`.
|
||
|
||
## Resolved, 2026-10-02
|
||
|
||
mesh-host 63: every field the host writes is compared before a container is called current, volumes
|
||
and paths included. mesh-controller 201: `build` and the take-in say when a module's policy rolls a
|
||
result out at once; under ADR 0162 the plan says it too.
|