A failure that repeats is said to be stuck

The mesh kept one report per machine, replaced, so a resource nothing can ever apply
looked like a failure that had just happened, every reconcile interval, for ever.
The row now keeps when the current failure began and how many reports in a row have
said it — the same outcome, refusal and failed resources; anything different starts
again and a clean apply clears it. Three make the machine stuck, and status says so
beside the failure, in words and in JSON (novox/hq 04-ISSUES/065, ADR 0090).
This commit is contained in:
2026-09-21 17:43:24 +02:00
parent 69d0b94395
commit fcaad7271a
6 changed files with 183 additions and 8 deletions
@@ -0,0 +1,15 @@
-- A failure that repeats is told apart from one that just happened (novox/hq 04-ISSUES/065).
--
-- A node re-applies what it holds on a steady interval and reports each time (novox/hq ADR 0010),
-- which is right for a failure that goes away by itself -- the overlay not up yet, a registry
-- briefly unreachable -- and makes a failure that will never go away look exactly the same: one
-- row, replaced, saying "failed" at a fresh time. Nothing distinguished "failed once, will succeed
-- when its dependency arrives" from "failed identically for ever", and nothing escalated the second.
--
-- Still one row per node. What is added is how long the CURRENT failure has been the same one:
-- when it first appeared, and how many reports in a row have said it. A report that says something
-- different starts the count again; a clean apply clears it.
alter table node_report
add column failing_since timestamptz,
add column failures int not null default 0;