issue 119: a hold is not a line in the apply report
This commit is contained in:
@@ -0,0 +1,59 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-26
|
||||
located-in: [mesh-host internal/apply, mesh-controller]
|
||||
---
|
||||
|
||||
# 119 — a hold is not a line in the apply report, and an operator flew blind into an outage
|
||||
|
||||
## What was observed
|
||||
|
||||
During the route-proxy edge cutover on novox (2026-09-26): the module was assigned, the
|
||||
push reported success, `status` said the node was doing everything it was told — and the
|
||||
module's three containers did not exist. The operator stopped the predecessor's proxy on
|
||||
the strength of those reports, and every public name on the node went dark until rollback.
|
||||
|
||||
The cause was correct behaviour, invisibly reported. The first (rolled-back) route-proxy
|
||||
attempt had left `/var/lib/route-proxy/*` on disk; on re-assign, the adopted node *found*
|
||||
those directories, held them (ADR 0100, exactly as designed), and held every container
|
||||
that mounts them — `"would mount /var/lib/route-proxy/ca, found on this adopted node;
|
||||
not run until route-proxy is taken"`. All of that lived only in `state.json`. What the
|
||||
operator saw:
|
||||
|
||||
- the push: `sent novox 346 resource(s)` — the controller's count of what it sent;
|
||||
- the node's journal: `applied 330 resource(s)` — sixteen fewer, with no line saying
|
||||
which sixteen or why;
|
||||
- `status`: green — a held resource is not "wrong", so nothing was flagged;
|
||||
- `node show novox`: the holds list did NOT include route-proxy's (it showed only holds
|
||||
the *controller* knew about from take-time listings, not what the node decided at
|
||||
apply-time).
|
||||
|
||||
Four surfaces, none carrying the one sentence that mattered: *route-proxy is assigned
|
||||
but not taken, and its containers will not run until it is.*
|
||||
|
||||
## Why this is a real fault and not operator error alone
|
||||
|
||||
The operator error (an edge-flip runbook that omitted `take`) was only possible because
|
||||
every surface reported success. A system whose correct refusals are indistinguishable
|
||||
from completed work will keep converting small procedural gaps into outages. The
|
||||
`sent 346 / applied 330` discrepancy was the single visible symptom, and interpreting it
|
||||
required reading `state.json` by hand.
|
||||
|
||||
## What would have prevented it
|
||||
|
||||
Any one of:
|
||||
|
||||
1. **The apply report says what it held.** `applied 330 resource(s), 16 held for
|
||||
untaken modules (route-proxy: 13, …)` — one line in the journal.
|
||||
2. **`status` counts holds against untaken-but-assigned modules.** A module assigned,
|
||||
pushed, and running zero of its containers is at minimum worth a "waiting on take"
|
||||
line — it is never converged in any useful sense.
|
||||
3. **`node show <node>` shows the node's own held list**, not only what take-time
|
||||
computed — the node already records it in `state.json` with reasons.
|
||||
|
||||
## Precedent
|
||||
|
||||
The photos cutover hit the same semantics benignly the same week (assign → held
|
||||
containers in `Created` state → take), and the mailu cutover documented "take is the
|
||||
verb, and ADR 0100 meant it". The semantics are consistent and right; the reporting is
|
||||
what let them be forgotten at the worst moment.
|
||||
Reference in New Issue
Block a user