Files
mesh-controller/internal/inventory/migrations/0025-a-failure-that-repeats-is-said-to-be-stuck.sql
T
jschoubben fcaad7271a A failure that repeats is said to be stuck
The mesh kept one report per machine, replaced, so a resource nothing can ever apply
looked like a failure that had just happened, every reconcile interval, for ever.
The row now keeps when the current failure began and how many reports in a row have
said it — the same outcome, refusal and failed resources; anything different starts
again and a clean apply clears it. Three make the machine stuck, and status says so
beside the failure, in words and in JSON (novox/hq 04-ISSUES/065, ADR 0090).
2026-09-21 17:43:24 +02:00

16 lines
971 B
SQL

-- A failure that repeats is told apart from one that just happened (novox/hq 04-ISSUES/065).
--
-- A node re-applies what it holds on a steady interval and reports each time (novox/hq ADR 0010),
-- which is right for a failure that goes away by itself -- the overlay not up yet, a registry
-- briefly unreachable -- and makes a failure that will never go away look exactly the same: one
-- row, replaced, saying "failed" at a fresh time. Nothing distinguished "failed once, will succeed
-- when its dependency arrives" from "failed identically for ever", and nothing escalated the second.
--
-- Still one row per node. What is added is how long the CURRENT failure has been the same one:
-- when it first appeared, and how many reports in a row have said it. A report that says something
-- different starts the count again; a clean apply clears it.
alter table node_report
add column failing_since timestamptz,
add column failures int not null default 0;