Files
hq/04-ISSUES/126-a-volume-path-is-not-in-the-spec-comparison/00-report.md
T
jschoubben 7842457d4b Issues 125 and 126: two apply-layer gaps the novox session hit live
125: a hold is not a line in the apply report — sixteen resources held
for an untaken module while four surfaces reported success, and the
operator stopped the edge's predecessor on their word (the route-proxy
flip outage). 126: a changed volume path neither recreates a running
container nor warns, and a roll-out upgrade policy makes a build a
deployment — together they turned a data-path migration into a forge
outage (the /var/lib move). Filed as 119/121 in the novox session
before syncing; renumbered past the other session's 119-124.
2026-09-26 17:25:27 +02:00

2.5 KiB
Raw Blame History

status, opened, located-in
status opened located-in
open 2026-09-26
mesh-host internal/apply

126 — a volume path is not in the spec comparison, and a roll-out raced a data move

What was observed

Landing the "module data lives in /var/lib" change on novox (mesh-catalog #97), two distinct faults surfaced in one hour:

  1. Building a module with a roll-out upgrade policy IS deploying it. gitea's policy was roll-out; the build that registered its repathed manifest sent it to the node immediately, which recreated the container mounting the not-yet-renamed (empty) /var/lib/gitea/data. The forge came back as its own install page, fresh host keys and all, and every subsequent pipeline build died on repository not found — which also blocked the fix, since re-registering the other modules needed the forge. The operator narrative "build, then move data, then push" is only safe under the record policy; nothing warned that one module in the batch would skip the pause.

  2. Changing a container's volume paths does not recreate the container. After the final push, five of the six repathed modules kept their old containers running ("Up 13–26 hours") — the new declaration's volume paths differ from the running containers' mounts, and the apply judged them current. Same class as mesh-host #27 (dns/ip absent from the comparison): a field the comparison does not read is a field that can never change a running container. Benign here only because a rename on one filesystem preserves the mounted inode — the running containers keep serving the same bytes the new path names, and the next natural recreation converges. A cross-filesystem move, or a path change to different data, would have silently split the module between two worlds.

What would have prevented it

  • build printing the module's upgrade policy when that policy will act on the result ("gitea rolls out on build — the node will receive this immediately"), or a --register-only flag for exactly this choreography.
  • Volumes (and every other container field) in the spec comparison, or the honest refusal: "this field changed and I cannot apply it without recreation."

Recovery that worked

Instant renames both ways broke the circular dependency (forge needed for builds, builds needed for the push, push needed for the forge): data back to the old path, old-spec forge started, artifacts rebuilt, data renamed forward, push. Nothing lost; the install-page junk was discarded twice.