The controller kept one report per machine, replaced, so a resource nothing can ever apply looked like a failure that had just happened, every few minutes, for ever. It now counts identical reports and status says stuck after three.
1.2 KiB
1.2 KiB
Diagnosis — 2026-09-21
- The controller keeps one report per machine, replaced on every report, by design: the question is the machine's current state and a history would bury it. A host reports after every apply and applies on its reconcile interval, so a permanent failure is a row that says "failed" at a fresh time every few minutes, indistinguishable from a failure that just happened.
- The four open questions, answered in turn. Escalate: yes, as a word in
status, which is where "what is wrong" is already read. The signal: consecutive identical reports, not a duration — a machine shut for a week has had one attempt. Where: the controller, which sees every node and can tell one stuck machine from a mesh-wide fault; the host does not know whether a failure is permanent and must keep trying. A gating failure: unchanged by this, and left open in the report.
Located in: the controller's node report and status. The fix keeps, beside the last report,
when the current failure began and how many reports in a row have said it; three make the machine
stuck, said in status and in its JSON. Decided in
ADR 0090.