Surfaced while reviewing the host apply loop for the controller/host split discussion. 065: a permanently-failing resource retries for ever with no escalation — reported, but never raises its hand as stuck. 066: a partly-applied declaration leaves a mixed state with no rollback, which is harmless for independent resources and unexamined for pairs that are only correct together. Both are questions HQ must answer, not incidents — status open, no fix proposed.
51 lines
2.9 KiB
Markdown
51 lines
2.9 KiB
Markdown
---
|
|
status: open
|
|
opened: 2026-09-20
|
|
located-in: []
|
|
fixed-by:
|
|
amended-design:
|
|
---
|
|
|
|
# 065 — A permanently failing resource is retried for ever with no escalation
|
|
|
|
## Symptom
|
|
|
|
A resource that can never apply — an image digest nothing will ever publish, a package that does
|
|
not exist, a file whose content is malformed for whatever reads it — is attempted, fails, is
|
|
reported, and is then attempted again on the host's reconcile loop every few minutes, indefinitely.
|
|
Nothing distinguishes "failed once, will succeed when its dependency arrives" from "failed and will
|
|
fail identically for ever", and nothing escalates the second case.
|
|
|
|
Observed by reading the apply loop, not from an incident: the host re-applies the declaration it
|
|
holds on a fixed interval (novox/hq ADR 0010 — delivery is a comparison, not a one-shot), which is
|
|
exactly right for a transient failure (the overlay is not up yet, the registry is briefly
|
|
unreachable) and makes a permanent failure silent in the way that matters. The failure is present
|
|
in the node's report and surfaced by `push --behind`, but only to someone who looks; the machine
|
|
itself does nothing to raise its hand as it loops on the same impossible resource.
|
|
|
|
## Why it matters beyond the instance
|
|
|
|
The retry loop exists so a machine converges without a person driving it, and that is the right
|
|
default. But "keep trying for ever" and "say nothing you are not asked" combine into the failure
|
|
shape this mesh exists to remove: a machine that reports success at the heartbeat level (it is up,
|
|
it is reconciling) while a resource inside it has never once applied. The distinction the host
|
|
draws so carefully within a single apply — reporting every failure, recording only what actually
|
|
took — is not carried across time: a resource that has failed every reconcile for a week reads the
|
|
same as one that failed its first attempt a minute ago.
|
|
|
|
A rule the whole design honours elsewhere is that a thing which cannot make progress must say so
|
|
rather than look busy. A reconcile that has re-attempted the same resource N times without it ever
|
|
succeeding is the host's equivalent, and nothing marks it.
|
|
|
|
## Open questions
|
|
|
|
- Should a resource that has failed every attempt for some duration or count be escalated — raised
|
|
distinctly to the control plane, marked in the node's state as "stuck" rather than merely
|
|
"waiting", so it is visible without someone running `push --behind`?
|
|
- Is the right signal a count of consecutive failures, an elapsed time since it last succeeded, or
|
|
the difference between "never once applied" and "applied then regressed"?
|
|
- Does escalation belong in the host (which sees the repetition) or the control plane (which sees
|
|
every node and could tell one stuck node from a mesh-wide fault)?
|
|
- Does anything change for a resource that is *gating* (a failed action/run-once stops what
|
|
follows), where the cost of a permanent failure is not one resource but everything after it?
|