The mesh kept one report per machine, replaced, so a resource nothing can ever apply looked like a failure that had just happened, every reconcile interval, for ever. The row now keeps when the current failure began and how many reports in a row have said it — the same outcome, refusal and failed resources; anything different starts again and a clean apply clears it. Three make the machine stuck, and status says so beside the failure, in words and in JSON (novox/hq 04-ISSUES/065, ADR 0090).
16 lines
971 B
SQL
16 lines
971 B
SQL
-- A failure that repeats is told apart from one that just happened (novox/hq 04-ISSUES/065).
|
|
--
|
|
-- A node re-applies what it holds on a steady interval and reports each time (novox/hq ADR 0010),
|
|
-- which is right for a failure that goes away by itself -- the overlay not up yet, a registry
|
|
-- briefly unreachable -- and makes a failure that will never go away look exactly the same: one
|
|
-- row, replaced, saying "failed" at a fresh time. Nothing distinguished "failed once, will succeed
|
|
-- when its dependency arrives" from "failed identically for ever", and nothing escalated the second.
|
|
--
|
|
-- Still one row per node. What is added is how long the CURRENT failure has been the same one:
|
|
-- when it first appeared, and how many reports in a row have said it. A report that says something
|
|
-- different starts the count again; a clean apply clears it.
|
|
|
|
alter table node_report
|
|
add column failing_since timestamptz,
|
|
add column failures int not null default 0;
|