Files
hq/04-ISSUES/125-a-hold-is-not-a-line-in-the-apply-report/00-report.md
T
jschoubben f7a37ee4f3 Issue 125 is resolved: a hold is a line in the report, and it breaks "all well"
Two of the four surfaces the report named already carried it — the host
has reported Held since ADR 0100, and node show reads the machine's own
list with an `as of` beside it. Recorded as checked rather than assumed.

Two did not. The apply line counted what it applied and said nothing
about the difference; status read the mesh's take-time listing, so a
module assigned after it showed nothing at all.

Both now say it, and the part that carries the weight: a hold suppresses
"all doing what they were told, all heard from, running what the mesh
would send them". That sentence was true for the whole outage, and acting
on it is what stopped the predecessor's proxy. Being adopted still does
not suppress it — a mode somebody chose is not a half-finished action.

Status does not call a hold a fault, deliberately. It is correct
behaviour, and a reader trained to see red for something the mesh did
right stops reading.
2026-09-30 08:47:36 +02:00

3.1 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
resolved 2026-09-26
mesh-host internal/link (the apply report line)
mesh-controller cmd/mesh-controller (status and its all-well condition)
mesh-host cbf5018, mesh-controller bfd983e

125 — a hold is not a line in the apply report, and an operator flew blind into an outage

What was observed

During the route-proxy edge cutover on novox (2026-09-26): the module was assigned, the push reported success, status said the node was doing everything it was told — and the module's three containers did not exist. The operator stopped the predecessor's proxy on the strength of those reports, and every public name on the node went dark until rollback.

The cause was correct behaviour, invisibly reported. The first (rolled-back) route-proxy attempt had left /var/lib/route-proxy/* on disk; on re-assign, the adopted node found those directories, held them (ADR 0100, exactly as designed), and held every container that mounts them — "would mount /var/lib/route-proxy/ca, found on this adopted node; not run until route-proxy is taken". All of that lived only in state.json. What the operator saw:

  • the push: sent novox 346 resource(s) — the controller's count of what it sent;
  • the node's journal: applied 330 resource(s) — sixteen fewer, with no line saying which sixteen or why;
  • status: green — a held resource is not "wrong", so nothing was flagged;
  • node show novox: the holds list did NOT include route-proxy's (it showed only holds the controller knew about from take-time listings, not what the node decided at apply-time).

Four surfaces, none carrying the one sentence that mattered: route-proxy is assigned but not taken, and its containers will not run until it is.

Why this is a real fault and not operator error alone

The operator error (an edge-flip runbook that omitted take) was only possible because every surface reported success. A system whose correct refusals are indistinguishable from completed work will keep converting small procedural gaps into outages. The sent 346 / applied 330 discrepancy was the single visible symptom, and interpreting it required reading state.json by hand.

What would have prevented it

Any one of:

  1. The apply report says what it held. applied 330 resource(s), 16 held for untaken modules (route-proxy: 13, …) — one line in the journal.
  2. status counts holds against untaken-but-assigned modules. A module assigned, pushed, and running zero of its containers is at minimum worth a "waiting on take" line — it is never converged in any useful sense.
  3. node show <node> shows the node's own held list, not only what take-time computed — the node already records it in state.json with reasons.

Precedent

The photos cutover hit the same semantics benignly the same week (assign → held containers in Created state → take), and the mailu cutover documented "take is the verb, and ADR 0100 meant it". The semantics are consistent and right; the reporting is what let them be forgotten at the worst moment.