--- topic: the mesh status: accepted date: 2026-09-21 deciders: jochen reconstructed: false extends: 02-DECISIONS/0010-delivery.md --- # 90. A failure that repeats is said to be stuck ## Context Delivery is a comparison, not a one-shot ([ADR 0010](0010-delivery.md)): a node re-applies the declaration it holds on a steady interval and reports each time. That is right for a failure that goes away by itself — the overlay not up yet, a registry briefly unreachable — and it makes a failure that will never go away look exactly the same. A resource nothing can ever apply is attempted, fails, is reported, and is attempted again every few minutes, indefinitely; the mesh keeps one report per machine, replaced, so each attempt arrives as "failed" at a fresh time. Nothing distinguished "failed once, will succeed when its dependency arrives" from "failed identically for ever", and nothing escalated the second ([issue 065](../04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md)). ## Considered Options 1. **The host gives up** after some number of attempts. Rejected: the host does not know whether a failure is permanent — that a registry has not answered three times is not evidence it never will — and a host that stops trying is a node that must be pushed to again by hand. 2. **A duration** since the failure was first seen. Rejected as the signal: a laptop shut for a week has had one attempt, and a week is not evidence of anything. 3. **The controller counts identical reports**, and says when there are enough of them. Adopted. ## Decision The mesh keeps, beside each machine's last report, when the current failure was first reported and how many reports in a row have said it — the same outcome, the same refusal, the same failed resources by id. Not by the host's words: an error carrying a duration or a counter would read as new on every report, and the resource looping on it is exactly what this is for. A report that says something different starts the count again; a clean apply clears it. **Three identical reports in a row make a machine stuck**: `status` says so beside the failure, with the count and the time it began, and the machine-readable status carries the same three facts. The host keeps retrying; being stuck is a statement about the mesh's knowledge, not an instruction to the machine. The controller counts rather than the host, because only it sees every node: one stuck machine and a mesh-wide fault are different situations, and the host cannot tell them apart. ## Consequences A resource that will never apply is visible from `status` after three reconcile intervals, to anyone who looks, without being asked for. What got harder: nothing on the machine changes — a gating failure that stops what follows still stops it, and this only makes the wait visible. Whether a stuck machine should also be raised as an event, and what a gating failure should do, stay open in the issue's own questions. ## How it is checked An inventory test records the same failure three times and asserts the count and the unchanged start; the same resource failing in other words, and asserts the count went on; a different failure, and asserts it restarted; a clean apply, and asserts both cleared. The status command's own test asserts a stuck machine is said to be one. ## References - [issue 065](../04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md) - [ADR 0010](0010-delivery.md) - [`03-DESIGN/01-to-be/10-delivery.md`](../03-DESIGN/01-to-be/10-delivery.md)