From f3611bbe63e2e64188c69741b81a90f70aa50ccc Mon Sep 17 00:00:00 2001 From: jochen Date: Sat, 26 Sep 2026 14:17:00 +0200 Subject: [PATCH] issue 119: a hold is not a line in the apply report --- .../00-report.md | 59 +++++++++++++++++++ 1 file changed, 59 insertions(+) create mode 100644 04-ISSUES/119-a-hold-is-not-a-line-in-the-apply-report/00-report.md diff --git a/04-ISSUES/119-a-hold-is-not-a-line-in-the-apply-report/00-report.md b/04-ISSUES/119-a-hold-is-not-a-line-in-the-apply-report/00-report.md new file mode 100644 index 0000000..55cd10c --- /dev/null +++ b/04-ISSUES/119-a-hold-is-not-a-line-in-the-apply-report/00-report.md @@ -0,0 +1,59 @@ +--- +status: open +opened: 2026-09-26 +located-in: [mesh-host internal/apply, mesh-controller] +--- + +# 119 — a hold is not a line in the apply report, and an operator flew blind into an outage + +## What was observed + +During the route-proxy edge cutover on novox (2026-09-26): the module was assigned, the +push reported success, `status` said the node was doing everything it was told — and the +module's three containers did not exist. The operator stopped the predecessor's proxy on +the strength of those reports, and every public name on the node went dark until rollback. + +The cause was correct behaviour, invisibly reported. The first (rolled-back) route-proxy +attempt had left `/var/lib/route-proxy/*` on disk; on re-assign, the adopted node *found* +those directories, held them (ADR 0100, exactly as designed), and held every container +that mounts them — `"would mount /var/lib/route-proxy/ca, found on this adopted node; +not run until route-proxy is taken"`. All of that lived only in `state.json`. What the +operator saw: + +- the push: `sent novox 346 resource(s)` — the controller's count of what it sent; +- the node's journal: `applied 330 resource(s)` — sixteen fewer, with no line saying + which sixteen or why; +- `status`: green — a held resource is not "wrong", so nothing was flagged; +- `node show novox`: the holds list did NOT include route-proxy's (it showed only holds + the *controller* knew about from take-time listings, not what the node decided at + apply-time). + +Four surfaces, none carrying the one sentence that mattered: *route-proxy is assigned +but not taken, and its containers will not run until it is.* + +## Why this is a real fault and not operator error alone + +The operator error (an edge-flip runbook that omitted `take`) was only possible because +every surface reported success. A system whose correct refusals are indistinguishable +from completed work will keep converting small procedural gaps into outages. The +`sent 346 / applied 330` discrepancy was the single visible symptom, and interpreting it +required reading `state.json` by hand. + +## What would have prevented it + +Any one of: + +1. **The apply report says what it held.** `applied 330 resource(s), 16 held for + untaken modules (route-proxy: 13, …)` — one line in the journal. +2. **`status` counts holds against untaken-but-assigned modules.** A module assigned, + pushed, and running zero of its containers is at minimum worth a "waiting on take" + line — it is never converged in any useful sense. +3. **`node show ` shows the node's own held list**, not only what take-time + computed — the node already records it in `state.json` with reasons. + +## Precedent + +The photos cutover hit the same semantics benignly the same week (assign → held +containers in `Created` state → take), and the mailu cutover documented "take is the +verb, and ADR 0100 meant it". The semantics are consistent and right; the reporting is +what let them be forgotten at the worst moment.