Files
jschoubben fcf34727fe ADR 0090: a failure that repeats is said to be stuck; issue 065 resolved
The controller kept one report per machine, replaced, so a resource nothing can
ever apply looked like a failure that had just happened, every few minutes, for
ever. It now counts identical reports and status says stuck after three.
2026-09-21 17:43:02 +02:00

18 lines
1.2 KiB
Markdown

# Diagnosis — 2026-09-21
1. The controller keeps one report per machine, replaced on every report, by design: the question
is the machine's current state and a history would bury it. A host reports after every apply
and applies on its reconcile interval, so a permanent failure is a row that says "failed" at a
fresh time every few minutes, indistinguishable from a failure that just happened.
2. The four open questions, answered in turn. Escalate: yes, as a word in `status`, which is
where "what is wrong" is already read. The signal: consecutive identical reports, not a
duration — a machine shut for a week has had one attempt. Where: the controller, which sees
every node and can tell one stuck machine from a mesh-wide fault; the host does not know
whether a failure is permanent and must keep trying. A gating failure: unchanged by this, and
left open in the report.
**Located in:** the controller's node report and `status`. The fix keeps, beside the last report,
when the current failure began and how many reports in a row have said it; three make the machine
stuck, said in `status` and in its JSON. Decided in
[ADR 0090](../../02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md).