Issues 125 and 126: two apply-layer gaps the novox session hit live #131
@@ -0,0 +1,59 @@
|
|||||||
|
---
|
||||||
|
status: open
|
||||||
|
opened: 2026-09-26
|
||||||
|
located-in: [mesh-host internal/apply, mesh-controller]
|
||||||
|
---
|
||||||
|
|
||||||
|
# 125 — a hold is not a line in the apply report, and an operator flew blind into an outage
|
||||||
|
|
||||||
|
## What was observed
|
||||||
|
|
||||||
|
During the route-proxy edge cutover on novox (2026-09-26): the module was assigned, the
|
||||||
|
push reported success, `status` said the node was doing everything it was told — and the
|
||||||
|
module's three containers did not exist. The operator stopped the predecessor's proxy on
|
||||||
|
the strength of those reports, and every public name on the node went dark until rollback.
|
||||||
|
|
||||||
|
The cause was correct behaviour, invisibly reported. The first (rolled-back) route-proxy
|
||||||
|
attempt had left `/var/lib/route-proxy/*` on disk; on re-assign, the adopted node *found*
|
||||||
|
those directories, held them (ADR 0100, exactly as designed), and held every container
|
||||||
|
that mounts them — `"would mount /var/lib/route-proxy/ca, found on this adopted node;
|
||||||
|
not run until route-proxy is taken"`. All of that lived only in `state.json`. What the
|
||||||
|
operator saw:
|
||||||
|
|
||||||
|
- the push: `sent novox 346 resource(s)` — the controller's count of what it sent;
|
||||||
|
- the node's journal: `applied 330 resource(s)` — sixteen fewer, with no line saying
|
||||||
|
which sixteen or why;
|
||||||
|
- `status`: green — a held resource is not "wrong", so nothing was flagged;
|
||||||
|
- `node show novox`: the holds list did NOT include route-proxy's (it showed only holds
|
||||||
|
the *controller* knew about from take-time listings, not what the node decided at
|
||||||
|
apply-time).
|
||||||
|
|
||||||
|
Four surfaces, none carrying the one sentence that mattered: *route-proxy is assigned
|
||||||
|
but not taken, and its containers will not run until it is.*
|
||||||
|
|
||||||
|
## Why this is a real fault and not operator error alone
|
||||||
|
|
||||||
|
The operator error (an edge-flip runbook that omitted `take`) was only possible because
|
||||||
|
every surface reported success. A system whose correct refusals are indistinguishable
|
||||||
|
from completed work will keep converting small procedural gaps into outages. The
|
||||||
|
`sent 346 / applied 330` discrepancy was the single visible symptom, and interpreting it
|
||||||
|
required reading `state.json` by hand.
|
||||||
|
|
||||||
|
## What would have prevented it
|
||||||
|
|
||||||
|
Any one of:
|
||||||
|
|
||||||
|
1. **The apply report says what it held.** `applied 330 resource(s), 16 held for
|
||||||
|
untaken modules (route-proxy: 13, …)` — one line in the journal.
|
||||||
|
2. **`status` counts holds against untaken-but-assigned modules.** A module assigned,
|
||||||
|
pushed, and running zero of its containers is at minimum worth a "waiting on take"
|
||||||
|
line — it is never converged in any useful sense.
|
||||||
|
3. **`node show <node>` shows the node's own held list**, not only what take-time
|
||||||
|
computed — the node already records it in `state.json` with reasons.
|
||||||
|
|
||||||
|
## Precedent
|
||||||
|
|
||||||
|
The photos cutover hit the same semantics benignly the same week (assign → held
|
||||||
|
containers in `Created` state → take), and the mailu cutover documented "take is the
|
||||||
|
verb, and ADR 0100 meant it". The semantics are consistent and right; the reporting is
|
||||||
|
what let them be forgotten at the worst moment.
|
||||||
@@ -0,0 +1,47 @@
|
|||||||
|
---
|
||||||
|
status: open
|
||||||
|
opened: 2026-09-26
|
||||||
|
located-in: [mesh-host internal/apply]
|
||||||
|
---
|
||||||
|
|
||||||
|
# 126 — a volume path is not in the spec comparison, and a roll-out raced a data move
|
||||||
|
|
||||||
|
## What was observed
|
||||||
|
|
||||||
|
Landing the "module data lives in /var/lib" change on novox (mesh-catalog #97), two
|
||||||
|
distinct faults surfaced in one hour:
|
||||||
|
|
||||||
|
1. **Building a module with a roll-out upgrade policy IS deploying it.** gitea's policy
|
||||||
|
was roll-out; the `build` that registered its repathed manifest sent it to the node
|
||||||
|
immediately, which recreated the container mounting the *not-yet-renamed* (empty)
|
||||||
|
`/var/lib/gitea/data`. The forge came back as its own install page, fresh host keys
|
||||||
|
and all, and every subsequent pipeline build died on `repository not found` — which
|
||||||
|
also blocked the fix, since re-registering the other modules needed the forge. The
|
||||||
|
operator narrative "build, then move data, then push" is only safe under the record
|
||||||
|
policy; nothing warned that one module in the batch would skip the pause.
|
||||||
|
|
||||||
|
2. **Changing a container's volume paths does not recreate the container.** After the
|
||||||
|
final push, five of the six repathed modules kept their old containers running
|
||||||
|
("Up 13–26 hours") — the new declaration's volume paths differ from the running
|
||||||
|
containers' mounts, and the apply judged them current. Same class as mesh-host #27
|
||||||
|
(`dns`/`ip` absent from the comparison): a field the comparison does not read is a
|
||||||
|
field that can never change a running container. Benign here only because a rename
|
||||||
|
on one filesystem preserves the mounted inode — the running containers keep serving
|
||||||
|
the same bytes the new path names, and the next natural recreation converges. A
|
||||||
|
cross-filesystem move, or a path change to *different* data, would have silently
|
||||||
|
split the module between two worlds.
|
||||||
|
|
||||||
|
## What would have prevented it
|
||||||
|
|
||||||
|
- `build` printing the module's upgrade policy when that policy will act on the result
|
||||||
|
("gitea rolls out on build — the node will receive this immediately"), or a
|
||||||
|
`--register-only` flag for exactly this choreography.
|
||||||
|
- Volumes (and every other container field) in the spec comparison, or the honest
|
||||||
|
refusal: "this field changed and I cannot apply it without recreation."
|
||||||
|
|
||||||
|
## Recovery that worked
|
||||||
|
|
||||||
|
Instant renames both ways broke the circular dependency (forge needed for builds,
|
||||||
|
builds needed for the push, push needed for the forge): data back to the old path,
|
||||||
|
old-spec forge started, artifacts rebuilt, data renamed forward, push. Nothing lost;
|
||||||
|
the install-page junk was discarded twice.
|
||||||
Reference in New Issue
Block a user