Files
hq/04-ISSUES/125-a-hold-is-not-a-line-in-the-apply-report/00-report.md
T
jschoubben 7842457d4b Issues 125 and 126: two apply-layer gaps the novox session hit live
125: a hold is not a line in the apply report — sixteen resources held
for an untaken module while four surfaces reported success, and the
operator stopped the edge's predecessor on their word (the route-proxy
flip outage). 126: a changed volume path neither recreates a running
container nor warns, and a roll-out upgrade policy makes a build a
deployment — together they turned a data-path migration into a forge
outage (the /var/lib move). Filed as 119/121 in the novox session
before syncing; renumbered past the other session's 119-124.
2026-09-26 17:25:27 +02:00

60 lines
2.9 KiB
Markdown

---
status: open
opened: 2026-09-26
located-in: [mesh-host internal/apply, mesh-controller]
---
# 125 — a hold is not a line in the apply report, and an operator flew blind into an outage
## What was observed
During the route-proxy edge cutover on novox (2026-09-26): the module was assigned, the
push reported success, `status` said the node was doing everything it was told — and the
module's three containers did not exist. The operator stopped the predecessor's proxy on
the strength of those reports, and every public name on the node went dark until rollback.
The cause was correct behaviour, invisibly reported. The first (rolled-back) route-proxy
attempt had left `/var/lib/route-proxy/*` on disk; on re-assign, the adopted node *found*
those directories, held them (ADR 0100, exactly as designed), and held every container
that mounts them — `"would mount /var/lib/route-proxy/ca, found on this adopted node;
not run until route-proxy is taken"`. All of that lived only in `state.json`. What the
operator saw:
- the push: `sent novox 346 resource(s)` — the controller's count of what it sent;
- the node's journal: `applied 330 resource(s)` — sixteen fewer, with no line saying
which sixteen or why;
- `status`: green — a held resource is not "wrong", so nothing was flagged;
- `node show novox`: the holds list did NOT include route-proxy's (it showed only holds
the *controller* knew about from take-time listings, not what the node decided at
apply-time).
Four surfaces, none carrying the one sentence that mattered: *route-proxy is assigned
but not taken, and its containers will not run until it is.*
## Why this is a real fault and not operator error alone
The operator error (an edge-flip runbook that omitted `take`) was only possible because
every surface reported success. A system whose correct refusals are indistinguishable
from completed work will keep converting small procedural gaps into outages. The
`sent 346 / applied 330` discrepancy was the single visible symptom, and interpreting it
required reading `state.json` by hand.
## What would have prevented it
Any one of:
1. **The apply report says what it held.** `applied 330 resource(s), 16 held for
untaken modules (route-proxy: 13, …)` — one line in the journal.
2. **`status` counts holds against untaken-but-assigned modules.** A module assigned,
pushed, and running zero of its containers is at minimum worth a "waiting on take"
line — it is never converged in any useful sense.
3. **`node show <node>` shows the node's own held list**, not only what take-time
computed — the node already records it in `state.json` with reasons.
## Precedent
The photos cutover hit the same semantics benignly the same week (assign → held
containers in `Created` state → take), and the mailu cutover documented "take is the
verb, and ADR 0100 meant it". The semantics are consistent and right; the reporting is
what let them be forgotten at the worst moment.