The controller kept one report per machine, replaced, so a resource nothing can ever apply looked like a failure that had just happened, every few minutes, for ever. It now counts identical reports and status says stuck after three.
18 lines
1.2 KiB
Markdown
18 lines
1.2 KiB
Markdown
# Diagnosis — 2026-09-21
|
|
|
|
1. The controller keeps one report per machine, replaced on every report, by design: the question
|
|
is the machine's current state and a history would bury it. A host reports after every apply
|
|
and applies on its reconcile interval, so a permanent failure is a row that says "failed" at a
|
|
fresh time every few minutes, indistinguishable from a failure that just happened.
|
|
2. The four open questions, answered in turn. Escalate: yes, as a word in `status`, which is
|
|
where "what is wrong" is already read. The signal: consecutive identical reports, not a
|
|
duration — a machine shut for a week has had one attempt. Where: the controller, which sees
|
|
every node and can tell one stuck machine from a mesh-wide fault; the host does not know
|
|
whether a failure is permanent and must keep trying. A gating failure: unchanged by this, and
|
|
left open in the report.
|
|
|
|
**Located in:** the controller's node report and `status`. The fix keeps, beside the last report,
|
|
when the current failure began and how many reports in a row have said it; three make the machine
|
|
stuck, said in `status` and in its JSON. Decided in
|
|
[ADR 0090](../../02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md).
|