Surfaced while reviewing the host apply loop for the controller/host split discussion. 065: a permanently-failing resource retries for ever with no escalation — reported, but never raises its hand as stuck. 066: a partly-applied declaration leaves a mixed state with no rollback, which is harmless for independent resources and unexamined for pairs that are only correct together. Both are questions HQ must answer, not incidents — status open, no fix proposed.
2.9 KiB
status, opened, located-in, fixed-by, amended-design
| status | opened | located-in | fixed-by | amended-design |
|---|---|---|---|---|
| open | 2026-09-20 |
065 — A permanently failing resource is retried for ever with no escalation
Symptom
A resource that can never apply — an image digest nothing will ever publish, a package that does not exist, a file whose content is malformed for whatever reads it — is attempted, fails, is reported, and is then attempted again on the host's reconcile loop every few minutes, indefinitely. Nothing distinguishes "failed once, will succeed when its dependency arrives" from "failed and will fail identically for ever", and nothing escalates the second case.
Observed by reading the apply loop, not from an incident: the host re-applies the declaration it
holds on a fixed interval (novox/hq ADR 0010 — delivery is a comparison, not a one-shot), which is
exactly right for a transient failure (the overlay is not up yet, the registry is briefly
unreachable) and makes a permanent failure silent in the way that matters. The failure is present
in the node's report and surfaced by push --behind, but only to someone who looks; the machine
itself does nothing to raise its hand as it loops on the same impossible resource.
Why it matters beyond the instance
The retry loop exists so a machine converges without a person driving it, and that is the right default. But "keep trying for ever" and "say nothing you are not asked" combine into the failure shape this mesh exists to remove: a machine that reports success at the heartbeat level (it is up, it is reconciling) while a resource inside it has never once applied. The distinction the host draws so carefully within a single apply — reporting every failure, recording only what actually took — is not carried across time: a resource that has failed every reconcile for a week reads the same as one that failed its first attempt a minute ago.
A rule the whole design honours elsewhere is that a thing which cannot make progress must say so rather than look busy. A reconcile that has re-attempted the same resource N times without it ever succeeding is the host's equivalent, and nothing marks it.
Open questions
- Should a resource that has failed every attempt for some duration or count be escalated — raised
distinctly to the control plane, marked in the node's state as "stuck" rather than merely
"waiting", so it is visible without someone running
push --behind? - Is the right signal a count of consecutive failures, an elapsed time since it last succeeded, or the difference between "never once applied" and "applied then regressed"?
- Does escalation belong in the host (which sees the repetition) or the control plane (which sees every node and could tell one stuck node from a mesh-wide fault)?
- Does anything change for a resource that is gating (a failed action/run-once stops what follows), where the cost of a permanent failure is not one resource but everything after it?