The controller kept one report per machine, replaced, so a resource nothing can ever apply looked like a failure that had just happened, every few minutes, for ever. It now counts identical reports and status says stuck after three.
65 lines
3.3 KiB
Markdown
65 lines
3.3 KiB
Markdown
---
|
|
topic: the mesh
|
|
status: accepted
|
|
date: 2026-09-21
|
|
deciders: jochen
|
|
reconstructed: false
|
|
extends: 02-DECISIONS/0010-delivery.md
|
|
---
|
|
|
|
# 90. A failure that repeats is said to be stuck
|
|
|
|
## Context
|
|
|
|
Delivery is a comparison, not a one-shot ([ADR 0010](0010-delivery.md)): a node re-applies the
|
|
declaration it holds on a steady interval and reports each time. That is right for a failure that
|
|
goes away by itself — the overlay not up yet, a registry briefly unreachable — and it makes a
|
|
failure that will never go away look exactly the same. A resource nothing can ever apply is
|
|
attempted, fails, is reported, and is attempted again every few minutes, indefinitely; the mesh
|
|
keeps one report per machine, replaced, so each attempt arrives as "failed" at a fresh time.
|
|
Nothing distinguished "failed once, will succeed when its dependency arrives" from "failed
|
|
identically for ever", and nothing escalated the second
|
|
([issue 065](../04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md)).
|
|
|
|
## Considered Options
|
|
|
|
1. **The host gives up** after some number of attempts. Rejected: the host does not know whether
|
|
a failure is permanent — that a registry has not answered three times is not evidence it never
|
|
will — and a host that stops trying is a node that must be pushed to again by hand.
|
|
2. **A duration** since the failure was first seen. Rejected as the signal: a laptop shut for a
|
|
week has had one attempt, and a week is not evidence of anything.
|
|
3. **The controller counts identical reports**, and says when there are enough of them. Adopted.
|
|
|
|
## Decision
|
|
|
|
The mesh keeps, beside each machine's last report, when the current failure was first reported
|
|
and how many reports in a row have said it — the same outcome, the same refusal, the same failed
|
|
resources with the same words. A report that says something different starts the count again; a
|
|
clean apply clears it. **Three identical reports in a row make a machine stuck**: `status` says
|
|
so beside the failure, with the count and the time it began, and the machine-readable status
|
|
carries the same three facts. The host keeps retrying; being stuck is a statement about the
|
|
mesh's knowledge, not an instruction to the machine.
|
|
|
|
The controller counts rather than the host, because only it sees every node: one stuck machine
|
|
and a mesh-wide fault are different situations, and the host cannot tell them apart.
|
|
|
|
## Consequences
|
|
|
|
A resource that will never apply is visible from `status` after three reconcile intervals, to
|
|
anyone who looks, without being asked for. What got harder: nothing on the machine changes — a
|
|
gating failure that stops what follows still stops it, and this only makes the wait visible.
|
|
Whether a stuck machine should also be raised as an event, and what a gating failure should do,
|
|
stay open in the issue's own questions.
|
|
|
|
## How it is checked
|
|
|
|
An inventory test records the same failure three times and asserts the count and the unchanged
|
|
start; then a different failure, and asserts the count restarted; then a clean apply, and asserts
|
|
both cleared. The status command's own test asserts a stuck machine is said to be one.
|
|
|
|
## References
|
|
|
|
- [issue 065](../04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md)
|
|
- [ADR 0010](0010-delivery.md)
|
|
- [`03-DESIGN/01-to-be/10-delivery.md`](../03-DESIGN/01-to-be/10-delivery.md)
|