125: a hold is not a line in the apply report — sixteen resources held for an untaken module while four surfaces reported success, and the operator stopped the edge's predecessor on their word (the route-proxy flip outage). 126: a changed volume path neither recreates a running container nor warns, and a roll-out upgrade policy makes a build a deployment — together they turned a data-path migration into a forge outage (the /var/lib move). Filed as 119/121 in the novox session before syncing; renumbered past the other session's 119-124.
60 lines
2.9 KiB
Markdown
60 lines
2.9 KiB
Markdown
---
|
|
status: open
|
|
opened: 2026-09-26
|
|
located-in: [mesh-host internal/apply, mesh-controller]
|
|
---
|
|
|
|
# 125 — a hold is not a line in the apply report, and an operator flew blind into an outage
|
|
|
|
## What was observed
|
|
|
|
During the route-proxy edge cutover on novox (2026-09-26): the module was assigned, the
|
|
push reported success, `status` said the node was doing everything it was told — and the
|
|
module's three containers did not exist. The operator stopped the predecessor's proxy on
|
|
the strength of those reports, and every public name on the node went dark until rollback.
|
|
|
|
The cause was correct behaviour, invisibly reported. The first (rolled-back) route-proxy
|
|
attempt had left `/var/lib/route-proxy/*` on disk; on re-assign, the adopted node *found*
|
|
those directories, held them (ADR 0100, exactly as designed), and held every container
|
|
that mounts them — `"would mount /var/lib/route-proxy/ca, found on this adopted node;
|
|
not run until route-proxy is taken"`. All of that lived only in `state.json`. What the
|
|
operator saw:
|
|
|
|
- the push: `sent novox 346 resource(s)` — the controller's count of what it sent;
|
|
- the node's journal: `applied 330 resource(s)` — sixteen fewer, with no line saying
|
|
which sixteen or why;
|
|
- `status`: green — a held resource is not "wrong", so nothing was flagged;
|
|
- `node show novox`: the holds list did NOT include route-proxy's (it showed only holds
|
|
the *controller* knew about from take-time listings, not what the node decided at
|
|
apply-time).
|
|
|
|
Four surfaces, none carrying the one sentence that mattered: *route-proxy is assigned
|
|
but not taken, and its containers will not run until it is.*
|
|
|
|
## Why this is a real fault and not operator error alone
|
|
|
|
The operator error (an edge-flip runbook that omitted `take`) was only possible because
|
|
every surface reported success. A system whose correct refusals are indistinguishable
|
|
from completed work will keep converting small procedural gaps into outages. The
|
|
`sent 346 / applied 330` discrepancy was the single visible symptom, and interpreting it
|
|
required reading `state.json` by hand.
|
|
|
|
## What would have prevented it
|
|
|
|
Any one of:
|
|
|
|
1. **The apply report says what it held.** `applied 330 resource(s), 16 held for
|
|
untaken modules (route-proxy: 13, …)` — one line in the journal.
|
|
2. **`status` counts holds against untaken-but-assigned modules.** A module assigned,
|
|
pushed, and running zero of its containers is at minimum worth a "waiting on take"
|
|
line — it is never converged in any useful sense.
|
|
3. **`node show <node>` shows the node's own held list**, not only what take-time
|
|
computed — the node already records it in `state.json` with reasons.
|
|
|
|
## Precedent
|
|
|
|
The photos cutover hit the same semantics benignly the same week (assign → held
|
|
containers in `Created` state → take), and the mailu cutover documented "take is the
|
|
verb, and ADR 0100 meant it". The semantics are consistent and right; the reporting is
|
|
what let them be forgotten at the worst moment.
|