ADR 0090: a failure is the same by resource id, not by the host's words

This commit is contained in:
2026-09-21 19:23:04 +02:00
parent 159a583cf0
commit 6ee95a3f31
2 changed files with 7 additions and 6 deletions
@@ -34,8 +34,9 @@ identically for ever", and nothing escalated the second
The mesh keeps, beside each machine's last report, when the current failure was first reported
and how many reports in a row have said it — the same outcome, the same refusal, the same failed
resources with the same words. A report that says something different starts the count again; a
clean apply clears it. **Three identical reports in a row make a machine stuck**: `status` says
resources by id. Not by the host's words: an error carrying a duration or a counter would read as
new on every report, and the resource looping on it is exactly what this is for. A report that
says something different starts the count again; a clean apply clears it. **Three identical reports in a row make a machine stuck**: `status` says
so beside the failure, with the count and the time it began, and the machine-readable status
carries the same three facts. The host keeps retrying; being stuck is a statement about the
mesh's knowledge, not an instruction to the machine.
@@ -54,8 +55,8 @@ stay open in the issue's own questions.
## How it is checked
An inventory test records the same failure three times and asserts the count and the unchanged
start; then a different failure, and asserts the count restarted; then a clean apply, and asserts
both cleared. The status command's own test asserts a stuck machine is said to be one.
start; the same resource failing in other words, and asserts the count went on; a different
failure, and asserts it restarted; a clean apply, and asserts both cleared. The status command's own test asserts a stuck machine is said to be one.
## References
+2 -2
View File
@@ -222,7 +222,7 @@ knows something the person reading the status does not.
([ADR 0090](../../02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md)). A machine
re-applies on its interval and reports each time, so a resource nothing can ever apply arrives as
the same failure over and over, at a fresh time each time. The mesh keeps, beside the last report,
when the current failure began and how many reports in a row have said it; three make the machine
stuck, and `status` says so beside the failure. The host keeps trying — stuck is what the mesh
when the current failure began and how many reports in a row have said it — the same resources by
id, whatever the words; three make the machine stuck, and `status` says so beside the failure. The host keeps trying — stuck is what the mesh
knows, not what the machine is told. *How it is checked:* an inventory test counts three identical
reports, a different one, and a clean apply; the status test asserts the word appears.