Files
hq/02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md
T

66 lines
3.5 KiB
Markdown

---
topic: the mesh
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0010-delivery.md
---
# 90. A failure that repeats is said to be stuck
## Context
Delivery is a comparison, not a one-shot ([ADR 0010](0010-delivery.md)): a node re-applies the
declaration it holds on a steady interval and reports each time. That is right for a failure that
goes away by itself — the overlay not up yet, a registry briefly unreachable — and it makes a
failure that will never go away look exactly the same. A resource nothing can ever apply is
attempted, fails, is reported, and is attempted again every few minutes, indefinitely; the mesh
keeps one report per machine, replaced, so each attempt arrives as "failed" at a fresh time.
Nothing distinguished "failed once, will succeed when its dependency arrives" from "failed
identically for ever", and nothing escalated the second
([issue 065](../04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md)).
## Considered Options
1. **The host gives up** after some number of attempts. Rejected: the host does not know whether
a failure is permanent — that a registry has not answered three times is not evidence it never
will — and a host that stops trying is a node that must be pushed to again by hand.
2. **A duration** since the failure was first seen. Rejected as the signal: a laptop shut for a
week has had one attempt, and a week is not evidence of anything.
3. **The controller counts identical reports**, and says when there are enough of them. Adopted.
## Decision
The mesh keeps, beside each machine's last report, when the current failure was first reported
and how many reports in a row have said it — the same outcome, the same refusal, the same failed
resources by id. Not by the host's words: an error carrying a duration or a counter would read as
new on every report, and the resource looping on it is exactly what this is for. A report that
says something different starts the count again; a clean apply clears it. **Three identical reports in a row make a machine stuck**: `status` says
so beside the failure, with the count and the time it began, and the machine-readable status
carries the same three facts. The host keeps retrying; being stuck is a statement about the
mesh's knowledge, not an instruction to the machine.
The controller counts rather than the host, because only it sees every node: one stuck machine
and a mesh-wide fault are different situations, and the host cannot tell them apart.
## Consequences
A resource that will never apply is visible from `status` after three reconcile intervals, to
anyone who looks, without being asked for. What got harder: nothing on the machine changes — a
gating failure that stops what follows still stops it, and this only makes the wait visible.
Whether a stuck machine should also be raised as an event, and what a gating failure should do,
stay open in the issue's own questions.
## How it is checked
An inventory test records the same failure three times and asserts the count and the unchanged
start; the same resource failing in other words, and asserts the count went on; a different
failure, and asserts it restarted; a clean apply, and asserts both cleared. The status command's own test asserts a stuck machine is said to be one.
## References
- [issue 065](../04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md)
- [ADR 0010](0010-delivery.md)
- [`03-DESIGN/01-to-be/10-delivery.md`](../03-DESIGN/01-to-be/10-delivery.md)