3.5 KiB
topic, status, date, deciders, reconstructed, extends
| topic | status | date | deciders | reconstructed | extends |
|---|---|---|---|---|---|
| the mesh | accepted | 2026-09-21 | jochen | false | 02-DECISIONS/0010-delivery.md |
90. A failure that repeats is said to be stuck
Context
Delivery is a comparison, not a one-shot (ADR 0010): a node re-applies the declaration it holds on a steady interval and reports each time. That is right for a failure that goes away by itself — the overlay not up yet, a registry briefly unreachable — and it makes a failure that will never go away look exactly the same. A resource nothing can ever apply is attempted, fails, is reported, and is attempted again every few minutes, indefinitely; the mesh keeps one report per machine, replaced, so each attempt arrives as "failed" at a fresh time. Nothing distinguished "failed once, will succeed when its dependency arrives" from "failed identically for ever", and nothing escalated the second (issue 065).
Considered Options
- The host gives up after some number of attempts. Rejected: the host does not know whether a failure is permanent — that a registry has not answered three times is not evidence it never will — and a host that stops trying is a node that must be pushed to again by hand.
- A duration since the failure was first seen. Rejected as the signal: a laptop shut for a week has had one attempt, and a week is not evidence of anything.
- The controller counts identical reports, and says when there are enough of them. Adopted.
Decision
The mesh keeps, beside each machine's last report, when the current failure was first reported
and how many reports in a row have said it — the same outcome, the same refusal, the same failed
resources by id. Not by the host's words: an error carrying a duration or a counter would read as
new on every report, and the resource looping on it is exactly what this is for. A report that
says something different starts the count again; a clean apply clears it. Three identical reports in a row make a machine stuck: status says
so beside the failure, with the count and the time it began, and the machine-readable status
carries the same three facts. The host keeps retrying; being stuck is a statement about the
mesh's knowledge, not an instruction to the machine.
The controller counts rather than the host, because only it sees every node: one stuck machine and a mesh-wide fault are different situations, and the host cannot tell them apart.
Consequences
A resource that will never apply is visible from status after three reconcile intervals, to
anyone who looks, without being asked for. What got harder: nothing on the machine changes — a
gating failure that stops what follows still stops it, and this only makes the wait visible.
Whether a stuck machine should also be raised as an event, and what a gating failure should do,
stay open in the issue's own questions.
How it is checked
An inventory test records the same failure three times and asserts the count and the unchanged start; the same resource failing in other words, and asserts the count went on; a different failure, and asserts it restarted; a clean apply, and asserts both cleared. The status command's own test asserts a stuck machine is said to be one.