Files
hq/02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md
T
jschoubben fcf34727fe ADR 0090: a failure that repeats is said to be stuck; issue 065 resolved
The controller kept one report per machine, replaced, so a resource nothing can
ever apply looked like a failure that had just happened, every few minutes, for
ever. It now counts identical reports and status says stuck after three.
2026-09-21 17:43:02 +02:00

3.3 KiB

topic, status, date, deciders, reconstructed, extends
topic status date deciders reconstructed extends
the mesh accepted 2026-09-21 jochen false 02-DECISIONS/0010-delivery.md

90. A failure that repeats is said to be stuck

Context

Delivery is a comparison, not a one-shot (ADR 0010): a node re-applies the declaration it holds on a steady interval and reports each time. That is right for a failure that goes away by itself — the overlay not up yet, a registry briefly unreachable — and it makes a failure that will never go away look exactly the same. A resource nothing can ever apply is attempted, fails, is reported, and is attempted again every few minutes, indefinitely; the mesh keeps one report per machine, replaced, so each attempt arrives as "failed" at a fresh time. Nothing distinguished "failed once, will succeed when its dependency arrives" from "failed identically for ever", and nothing escalated the second (issue 065).

Considered Options

  1. The host gives up after some number of attempts. Rejected: the host does not know whether a failure is permanent — that a registry has not answered three times is not evidence it never will — and a host that stops trying is a node that must be pushed to again by hand.
  2. A duration since the failure was first seen. Rejected as the signal: a laptop shut for a week has had one attempt, and a week is not evidence of anything.
  3. The controller counts identical reports, and says when there are enough of them. Adopted.

Decision

The mesh keeps, beside each machine's last report, when the current failure was first reported and how many reports in a row have said it — the same outcome, the same refusal, the same failed resources with the same words. A report that says something different starts the count again; a clean apply clears it. Three identical reports in a row make a machine stuck: status says so beside the failure, with the count and the time it began, and the machine-readable status carries the same three facts. The host keeps retrying; being stuck is a statement about the mesh's knowledge, not an instruction to the machine.

The controller counts rather than the host, because only it sees every node: one stuck machine and a mesh-wide fault are different situations, and the host cannot tell them apart.

Consequences

A resource that will never apply is visible from status after three reconcile intervals, to anyone who looks, without being asked for. What got harder: nothing on the machine changes — a gating failure that stops what follows still stops it, and this only makes the wait visible. Whether a stuck machine should also be raised as an event, and what a gating failure should do, stay open in the issue's own questions.

How it is checked

An inventory test records the same failure three times and asserts the count and the unchanged start; then a different failure, and asserts the count restarted; then a clean apply, and asserts both cleared. The status command's own test asserts a stuck machine is said to be one.

References