A failure that repeats is said to be stuck
The mesh kept one report per machine, replaced, so a resource nothing can ever apply looked like a failure that had just happened, every reconcile interval, for ever. The row now keeps when the current failure began and how many reports in a row have said it — the same outcome, refusal and failed resources; anything different starts again and a clean apply clears it. Three make the machine stuck, and status says so beside the failure, in words and in JSON (novox/hq 04-ISSUES/065, ADR 0090).
This commit is contained in:
@@ -0,0 +1,15 @@
|
||||
-- A failure that repeats is told apart from one that just happened (novox/hq 04-ISSUES/065).
|
||||
--
|
||||
-- A node re-applies what it holds on a steady interval and reports each time (novox/hq ADR 0010),
|
||||
-- which is right for a failure that goes away by itself -- the overlay not up yet, a registry
|
||||
-- briefly unreachable -- and makes a failure that will never go away look exactly the same: one
|
||||
-- row, replaced, saying "failed" at a fresh time. Nothing distinguished "failed once, will succeed
|
||||
-- when its dependency arrives" from "failed identically for ever", and nothing escalated the second.
|
||||
--
|
||||
-- Still one row per node. What is added is how long the CURRENT failure has been the same one:
|
||||
-- when it first appeared, and how many reports in a row have said it. A report that says something
|
||||
-- different starts the count again; a clean apply clears it.
|
||||
|
||||
alter table node_report
|
||||
add column failing_since timestamptz,
|
||||
add column failures int not null default 0;
|
||||
Reference in New Issue
Block a user