Files
hq/04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md
T
jschoubben a9bc0c1cc5 Open issues 065 and 066 — the apply failure model's two gaps
Surfaced while reviewing the host apply loop for the controller/host
split discussion. 065: a permanently-failing resource retries for ever
with no escalation — reported, but never raises its hand as stuck.
066: a partly-applied declaration leaves a mixed state with no rollback,
which is harmless for independent resources and unexamined for pairs
that are only correct together. Both are questions HQ must answer, not
incidents — status open, no fix proposed.
2026-09-20 14:28:04 +02:00

2.9 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
open 2026-09-20

065 — A permanently failing resource is retried for ever with no escalation

Symptom

A resource that can never apply — an image digest nothing will ever publish, a package that does not exist, a file whose content is malformed for whatever reads it — is attempted, fails, is reported, and is then attempted again on the host's reconcile loop every few minutes, indefinitely. Nothing distinguishes "failed once, will succeed when its dependency arrives" from "failed and will fail identically for ever", and nothing escalates the second case.

Observed by reading the apply loop, not from an incident: the host re-applies the declaration it holds on a fixed interval (novox/hq ADR 0010 — delivery is a comparison, not a one-shot), which is exactly right for a transient failure (the overlay is not up yet, the registry is briefly unreachable) and makes a permanent failure silent in the way that matters. The failure is present in the node's report and surfaced by push --behind, but only to someone who looks; the machine itself does nothing to raise its hand as it loops on the same impossible resource.

Why it matters beyond the instance

The retry loop exists so a machine converges without a person driving it, and that is the right default. But "keep trying for ever" and "say nothing you are not asked" combine into the failure shape this mesh exists to remove: a machine that reports success at the heartbeat level (it is up, it is reconciling) while a resource inside it has never once applied. The distinction the host draws so carefully within a single apply — reporting every failure, recording only what actually took — is not carried across time: a resource that has failed every reconcile for a week reads the same as one that failed its first attempt a minute ago.

A rule the whole design honours elsewhere is that a thing which cannot make progress must say so rather than look busy. A reconcile that has re-attempted the same resource N times without it ever succeeding is the host's equivalent, and nothing marks it.

Open questions

  • Should a resource that has failed every attempt for some duration or count be escalated — raised distinctly to the control plane, marked in the node's state as "stuck" rather than merely "waiting", so it is visible without someone running push --behind?
  • Is the right signal a count of consecutive failures, an elapsed time since it last succeeded, or the difference between "never once applied" and "applied then regressed"?
  • Does escalation belong in the host (which sees the repetition) or the control plane (which sees every node and could tell one stuck node from a mesh-wide fault)?
  • Does anything change for a resource that is gating (a failed action/run-once stops what follows), where the cost of a permanent failure is not one resource but everything after it?