Files
hq/04-ISSUES/126-a-volume-path-is-not-in-the-spec-comparison/00-report.md
T
jschoubben 7842457d4b Issues 125 and 126: two apply-layer gaps the novox session hit live
125: a hold is not a line in the apply report — sixteen resources held
for an untaken module while four surfaces reported success, and the
operator stopped the edge's predecessor on their word (the route-proxy
flip outage). 126: a changed volume path neither recreates a running
container nor warns, and a roll-out upgrade policy makes a build a
deployment — together they turned a data-path migration into a forge
outage (the /var/lib move). Filed as 119/121 in the novox session
before syncing; renumbered past the other session's 119-124.
2026-09-26 17:25:27 +02:00

48 lines
2.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
status: open
opened: 2026-09-26
located-in: [mesh-host internal/apply]
---
# 126 — a volume path is not in the spec comparison, and a roll-out raced a data move
## What was observed
Landing the "module data lives in /var/lib" change on novox (mesh-catalog #97), two
distinct faults surfaced in one hour:
1. **Building a module with a roll-out upgrade policy IS deploying it.** gitea's policy
was roll-out; the `build` that registered its repathed manifest sent it to the node
immediately, which recreated the container mounting the *not-yet-renamed* (empty)
`/var/lib/gitea/data`. The forge came back as its own install page, fresh host keys
and all, and every subsequent pipeline build died on `repository not found` — which
also blocked the fix, since re-registering the other modules needed the forge. The
operator narrative "build, then move data, then push" is only safe under the record
policy; nothing warned that one module in the batch would skip the pause.
2. **Changing a container's volume paths does not recreate the container.** After the
final push, five of the six repathed modules kept their old containers running
("Up 13–26 hours") — the new declaration's volume paths differ from the running
containers' mounts, and the apply judged them current. Same class as mesh-host #27
(`dns`/`ip` absent from the comparison): a field the comparison does not read is a
field that can never change a running container. Benign here only because a rename
on one filesystem preserves the mounted inode — the running containers keep serving
the same bytes the new path names, and the next natural recreation converges. A
cross-filesystem move, or a path change to *different* data, would have silently
split the module between two worlds.
## What would have prevented it
- `build` printing the module's upgrade policy when that policy will act on the result
("gitea rolls out on build — the node will receive this immediately"), or a
`--register-only` flag for exactly this choreography.
- Volumes (and every other container field) in the spec comparison, or the honest
refusal: "this field changed and I cannot apply it without recreation."
## Recovery that worked
Instant renames both ways broke the circular dependency (forge needed for builds,
builds needed for the push, push needed for the forge): data back to the old path,
old-spec forge started, artifacts rebuilt, data renamed forward, push. Nothing lost;
the install-page junk was discarded twice.