Files
mesh-controller/internal/inventory/migrations/0025-a-failure-that-repeats-is-said-to-be-stuck.sql
T
jschoubben 53a79a17a1 Review: a failure is the same by resource id, not by the host's words; bound files and /run/docker.sock are declared
The host's error text may carry a duration or a counter, and a resource looping on
it would never have read as stuck. The previous row is read and compared here.
Stuck needs a start to say. A container may mount the file a binding lands in; the
runtime socket is declared under both of its spellings; the catalogue-wide test
takes MESH_CATALOG.
2026-09-21 19:23:02 +02:00

17 lines
1.1 KiB
SQL

-- A failure that repeats is told apart from one that just happened (novox/hq 04-ISSUES/065).
--
-- A node re-applies what it holds on a steady interval and reports each time (novox/hq ADR 0010),
-- which is right for a failure that goes away by itself -- the overlay not up yet, a registry
-- briefly unreachable -- and makes a failure that will never go away look exactly the same: one
-- row, replaced, saying "failed" at a fresh time. Nothing distinguished "failed once, will succeed
-- when its dependency arrives" from "failed identically for ever", and nothing escalated the second.
--
-- Still one row per node. What is added is how long the CURRENT failure has been the same one:
-- when it first appeared, and how many reports in a row have said it -- the same outcome, the same
-- refusal, the same failed resources by id (not by the host's words, which may carry a duration).
-- A report that says something different starts the count again; a clean apply clears it.
alter table node_report
add column failing_since timestamptz,
add column failures int not null default 0;