Files
jschoubben fcf34727fe ADR 0090: a failure that repeats is said to be stuck; issue 065 resolved
The controller kept one report per machine, replaced, so a resource nothing can
ever apply looked like a failure that had just happened, every few minutes, for
ever. It now counts identical reports and status says stuck after three.
2026-09-21 17:43:02 +02:00

1.2 KiB

Diagnosis — 2026-09-21

  1. The controller keeps one report per machine, replaced on every report, by design: the question is the machine's current state and a history would bury it. A host reports after every apply and applies on its reconcile interval, so a permanent failure is a row that says "failed" at a fresh time every few minutes, indistinguishable from a failure that just happened.
  2. The four open questions, answered in turn. Escalate: yes, as a word in status, which is where "what is wrong" is already read. The signal: consecutive identical reports, not a duration — a machine shut for a week has had one attempt. Where: the controller, which sees every node and can tell one stuck machine from a mesh-wide fault; the host does not know whether a failure is permanent and must keep trying. A gating failure: unchanged by this, and left open in the report.

Located in: the controller's node report and status. The fix keeps, beside the last report, when the current failure began and how many reports in a row have said it; three make the machine stuck, said in status and in its JSON. Decided in ADR 0090.