diff --git a/02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md b/02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md new file mode 100644 index 0000000..33b684e --- /dev/null +++ b/02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md @@ -0,0 +1,64 @@ +--- +topic: the mesh +status: accepted +date: 2026-09-21 +deciders: jochen +reconstructed: false +extends: 02-DECISIONS/0010-delivery.md +--- + +# 90. A failure that repeats is said to be stuck + +## Context + +Delivery is a comparison, not a one-shot ([ADR 0010](0010-delivery.md)): a node re-applies the +declaration it holds on a steady interval and reports each time. That is right for a failure that +goes away by itself — the overlay not up yet, a registry briefly unreachable — and it makes a +failure that will never go away look exactly the same. A resource nothing can ever apply is +attempted, fails, is reported, and is attempted again every few minutes, indefinitely; the mesh +keeps one report per machine, replaced, so each attempt arrives as "failed" at a fresh time. +Nothing distinguished "failed once, will succeed when its dependency arrives" from "failed +identically for ever", and nothing escalated the second +([issue 065](../04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md)). + +## Considered Options + +1. **The host gives up** after some number of attempts. Rejected: the host does not know whether + a failure is permanent — that a registry has not answered three times is not evidence it never + will — and a host that stops trying is a node that must be pushed to again by hand. +2. **A duration** since the failure was first seen. Rejected as the signal: a laptop shut for a + week has had one attempt, and a week is not evidence of anything. +3. **The controller counts identical reports**, and says when there are enough of them. Adopted. + +## Decision + +The mesh keeps, beside each machine's last report, when the current failure was first reported +and how many reports in a row have said it — the same outcome, the same refusal, the same failed +resources with the same words. A report that says something different starts the count again; a +clean apply clears it. **Three identical reports in a row make a machine stuck**: `status` says +so beside the failure, with the count and the time it began, and the machine-readable status +carries the same three facts. The host keeps retrying; being stuck is a statement about the +mesh's knowledge, not an instruction to the machine. + +The controller counts rather than the host, because only it sees every node: one stuck machine +and a mesh-wide fault are different situations, and the host cannot tell them apart. + +## Consequences + +A resource that will never apply is visible from `status` after three reconcile intervals, to +anyone who looks, without being asked for. What got harder: nothing on the machine changes — a +gating failure that stops what follows still stops it, and this only makes the wait visible. +Whether a stuck machine should also be raised as an event, and what a gating failure should do, +stay open in the issue's own questions. + +## How it is checked + +An inventory test records the same failure three times and asserts the count and the unchanged +start; then a different failure, and asserts the count restarted; then a clean apply, and asserts +both cleared. The status command's own test asserts a stuck machine is said to be one. + +## References + +- [issue 065](../04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md) +- [ADR 0010](0010-delivery.md) +- [`03-DESIGN/01-to-be/10-delivery.md`](../03-DESIGN/01-to-be/10-delivery.md) diff --git a/02-DECISIONS/README.md b/02-DECISIONS/README.md index 1bdb3ba..148d930 100644 --- a/02-DECISIONS/README.md +++ b/02-DECISIONS/README.md @@ -87,6 +87,7 @@ python3 00-META/checks/index.py fail if stale - **0077** — [The parts are named controller, foundation, node — not control plane, substrate, master](0077-the-controller-and-the-foundation.md) - **0083** — [One push leaves the mesh consistent](0083-one-push-leaves-the-mesh-consistent.md) - **0088** — [The foundation filters before anything listens](0088-the-foundation-filters-before-anything-listens.md) +- **0090** — [A failure that repeats is said to be stuck](0090-a-failure-that-repeats-is-said-to-be-stuck.md) ### Its tiers, from the bottom up diff --git a/03-DESIGN/01-to-be/10-delivery.md b/03-DESIGN/01-to-be/10-delivery.md index e463161..6ca0cfd 100644 --- a/03-DESIGN/01-to-be/10-delivery.md +++ b/03-DESIGN/01-to-be/10-delivery.md @@ -5,8 +5,9 @@ code: - mesh-controller internal/builder - mesh-controller cmd/mesh-controller (build, build --behind, push, status) - mesh-controller internal/inventory/builds.go -updated: 2026-08-31 +updated: 2026-09-21 decisions: + - 02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md - 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md - 02-DECISIONS/0010-delivery.md - 02-DECISIONS/0009-modules-and-the-graph.md @@ -216,3 +217,12 @@ the remedy is the same push: **`status` says it and `push --behind` acts on it**, and both because the alternative is a flag that knows something the person reading the status does not. + +**A failure that repeats is said to be stuck** +([ADR 0090](../../02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md)). A machine +re-applies on its interval and reports each time, so a resource nothing can ever apply arrives as +the same failure over and over, at a fresh time each time. The mesh keeps, beside the last report, +when the current failure began and how many reports in a row have said it; three make the machine +stuck, and `status` says so beside the failure. The host keeps trying — stuck is what the mesh +knows, not what the machine is told. *How it is checked:* an inventory test counts three identical +reports, a different one, and a clean apply; the status test asserts the word appears. diff --git a/04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md b/04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md index 6b77818..b6931cf 100644 --- a/04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md +++ b/04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md @@ -1,9 +1,9 @@ --- -status: open +status: resolved opened: 2026-09-20 -located-in: [] -fixed-by: -amended-design: +located-in: [mesh-controller internal/inventory (node_report), mesh-controller cmd/mesh-controller (status)] +fixed-by: ADR 0090; mesh-controller feat/multiple-fixes (the report counts identical failures; status says stuck after three) +amended-design: 03-DESIGN/01-to-be/10-delivery.md --- # 065 — A permanently failing resource is retried for ever with no escalation diff --git a/04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/01-diagnosis.md b/04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/01-diagnosis.md new file mode 100644 index 0000000..cb9e2ea --- /dev/null +++ b/04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/01-diagnosis.md @@ -0,0 +1,17 @@ +# Diagnosis — 2026-09-21 + +1. The controller keeps one report per machine, replaced on every report, by design: the question + is the machine's current state and a history would bury it. A host reports after every apply + and applies on its reconcile interval, so a permanent failure is a row that says "failed" at a + fresh time every few minutes, indistinguishable from a failure that just happened. +2. The four open questions, answered in turn. Escalate: yes, as a word in `status`, which is + where "what is wrong" is already read. The signal: consecutive identical reports, not a + duration — a machine shut for a week has had one attempt. Where: the controller, which sees + every node and can tell one stuck machine from a mesh-wide fault; the host does not know + whether a failure is permanent and must keep trying. A gating failure: unchanged by this, and + left open in the report. + +**Located in:** the controller's node report and `status`. The fix keeps, beside the last report, +when the current failure began and how many reports in a row have said it; three make the machine +stuck, said in `status` and in its JSON. Decided in +[ADR 0090](../../02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md).