ADR 0090: a failure that repeats is said to be stuck; issue 065 resolved

The controller kept one report per machine, replaced, so a resource nothing can
ever apply looked like a failure that had just happened, every few minutes, for
ever. It now counts identical reports and status says stuck after three.
This commit is contained in:
2026-09-21 17:43:02 +02:00
parent a5e351f4cb
commit fcf34727fe
5 changed files with 97 additions and 5 deletions
+11 -1
View File
@@ -5,8 +5,9 @@ code:
- mesh-controller internal/builder
- mesh-controller cmd/mesh-controller (build, build --behind, push, status)
- mesh-controller internal/inventory/builds.go
updated: 2026-08-31
updated: 2026-09-21
decisions:
- 02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md
- 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md
- 02-DECISIONS/0010-delivery.md
- 02-DECISIONS/0009-modules-and-the-graph.md
@@ -216,3 +217,12 @@ the remedy is the same push:
**`status` says it and `push --behind` acts on it**, and both because the alternative is a flag that
knows something the person reading the status does not.
**A failure that repeats is said to be stuck**
([ADR 0090](../../02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md)). A machine
re-applies on its interval and reports each time, so a resource nothing can ever apply arrives as
the same failure over and over, at a fresh time each time. The mesh keeps, beside the last report,
when the current failure began and how many reports in a row have said it; three make the machine
stuck, and `status` says so beside the failure. The host keeps trying — stuck is what the mesh
knows, not what the machine is told. *How it is checked:* an inventory test counts three identical
reports, a different one, and a clean apply; the status test asserts the word appears.