Multiple fixes: status vocabulary aligned, 065/026/070/007 resolved, 066/069/049/046 located, 064 diagnosed #64

Merged
jschoubben merged 10 commits from feat/multiple-fixes into main 2026-09-21 17:23:35 +00:00
5 changed files with 97 additions and 5 deletions
Showing only changes of commit fcf34727fe - Show all commits
@@ -0,0 +1,64 @@
---
topic: the mesh
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0010-delivery.md
---
# 90. A failure that repeats is said to be stuck
## Context
Delivery is a comparison, not a one-shot ([ADR 0010](0010-delivery.md)): a node re-applies the
declaration it holds on a steady interval and reports each time. That is right for a failure that
goes away by itself — the overlay not up yet, a registry briefly unreachable — and it makes a
failure that will never go away look exactly the same. A resource nothing can ever apply is
attempted, fails, is reported, and is attempted again every few minutes, indefinitely; the mesh
keeps one report per machine, replaced, so each attempt arrives as "failed" at a fresh time.
Nothing distinguished "failed once, will succeed when its dependency arrives" from "failed
identically for ever", and nothing escalated the second
([issue 065](../04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md)).
## Considered Options
1. **The host gives up** after some number of attempts. Rejected: the host does not know whether
a failure is permanent — that a registry has not answered three times is not evidence it never
will — and a host that stops trying is a node that must be pushed to again by hand.
2. **A duration** since the failure was first seen. Rejected as the signal: a laptop shut for a
week has had one attempt, and a week is not evidence of anything.
3. **The controller counts identical reports**, and says when there are enough of them. Adopted.
## Decision
The mesh keeps, beside each machine's last report, when the current failure was first reported
and how many reports in a row have said it — the same outcome, the same refusal, the same failed
resources with the same words. A report that says something different starts the count again; a
clean apply clears it. **Three identical reports in a row make a machine stuck**: `status` says
so beside the failure, with the count and the time it began, and the machine-readable status
carries the same three facts. The host keeps retrying; being stuck is a statement about the
mesh's knowledge, not an instruction to the machine.
The controller counts rather than the host, because only it sees every node: one stuck machine
and a mesh-wide fault are different situations, and the host cannot tell them apart.
## Consequences
A resource that will never apply is visible from `status` after three reconcile intervals, to
anyone who looks, without being asked for. What got harder: nothing on the machine changes — a
gating failure that stops what follows still stops it, and this only makes the wait visible.
Whether a stuck machine should also be raised as an event, and what a gating failure should do,
stay open in the issue's own questions.
## How it is checked
An inventory test records the same failure three times and asserts the count and the unchanged
start; then a different failure, and asserts the count restarted; then a clean apply, and asserts
both cleared. The status command's own test asserts a stuck machine is said to be one.
## References
- [issue 065](../04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md)
- [ADR 0010](0010-delivery.md)
- [`03-DESIGN/01-to-be/10-delivery.md`](../03-DESIGN/01-to-be/10-delivery.md)
+1
View File
@@ -87,6 +87,7 @@ python3 00-META/checks/index.py fail if stale
- **0077** — [The parts are named controller, foundation, node — not control plane, substrate, master](0077-the-controller-and-the-foundation.md)
- **0083** — [One push leaves the mesh consistent](0083-one-push-leaves-the-mesh-consistent.md)
- **0088** — [The foundation filters before anything listens](0088-the-foundation-filters-before-anything-listens.md)
- **0090** — [A failure that repeats is said to be stuck](0090-a-failure-that-repeats-is-said-to-be-stuck.md)
### Its tiers, from the bottom up
+11 -1
View File
@@ -5,8 +5,9 @@ code:
- mesh-controller internal/builder
- mesh-controller cmd/mesh-controller (build, build --behind, push, status)
- mesh-controller internal/inventory/builds.go
updated: 2026-08-31
updated: 2026-09-21
decisions:
- 02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md
- 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md
- 02-DECISIONS/0010-delivery.md
- 02-DECISIONS/0009-modules-and-the-graph.md
@@ -216,3 +217,12 @@ the remedy is the same push:
**`status` says it and `push --behind` acts on it**, and both because the alternative is a flag that
knows something the person reading the status does not.
**A failure that repeats is said to be stuck**
([ADR 0090](../../02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md)). A machine
re-applies on its interval and reports each time, so a resource nothing can ever apply arrives as
the same failure over and over, at a fresh time each time. The mesh keeps, beside the last report,
when the current failure began and how many reports in a row have said it; three make the machine
stuck, and `status` says so beside the failure. The host keeps trying — stuck is what the mesh
knows, not what the machine is told. *How it is checked:* an inventory test counts three identical
reports, a different one, and a clean apply; the status test asserts the word appears.
@@ -1,9 +1,9 @@
---
status: open
status: resolved
opened: 2026-09-20
located-in: []
fixed-by:
amended-design:
located-in: [mesh-controller internal/inventory (node_report), mesh-controller cmd/mesh-controller (status)]
fixed-by: ADR 0090; mesh-controller feat/multiple-fixes (the report counts identical failures; status says stuck after three)
amended-design: 03-DESIGN/01-to-be/10-delivery.md
---
# 065 — A permanently failing resource is retried for ever with no escalation
@@ -0,0 +1,17 @@
# Diagnosis — 2026-09-21
1. The controller keeps one report per machine, replaced on every report, by design: the question
is the machine's current state and a history would bury it. A host reports after every apply
and applies on its reconcile interval, so a permanent failure is a row that says "failed" at a
fresh time every few minutes, indistinguishable from a failure that just happened.
2. The four open questions, answered in turn. Escalate: yes, as a word in `status`, which is
where "what is wrong" is already read. The signal: consecutive identical reports, not a
duration — a machine shut for a week has had one attempt. Where: the controller, which sees
every node and can tell one stuck machine from a mesh-wide fault; the host does not know
whether a failure is permanent and must keep trying. A gating failure: unchanged by this, and
left open in the report.
**Located in:** the controller's node report and `status`. The fix keeps, beside the last report,
when the current failure began and how many reports in a row have said it; three make the machine
stuck, said in `status` and in its JSON. Decided in
[ADR 0090](../../02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md).