Files
jschoubben fcf34727fe ADR 0090: a failure that repeats is said to be stuck; issue 065 resolved
The controller kept one report per machine, replaced, so a resource nothing can
ever apply looked like a failure that had just happened, every few minutes, for
ever. It now counts identical reports and status says stuck after three.
2026-09-21 17:43:02 +02:00

3.1 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
resolved 2026-09-20
mesh-controller internal/inventory (node_report)
mesh-controller cmd/mesh-controller (status)
ADR 0090; mesh-controller feat/multiple-fixes (the report counts identical failures; status says stuck after three) 03-DESIGN/01-to-be/10-delivery.md

065 — A permanently failing resource is retried for ever with no escalation

Symptom

A resource that can never apply — an image digest nothing will ever publish, a package that does not exist, a file whose content is malformed for whatever reads it — is attempted, fails, is reported, and is then attempted again on the host's reconcile loop every few minutes, indefinitely. Nothing distinguishes "failed once, will succeed when its dependency arrives" from "failed and will fail identically for ever", and nothing escalates the second case.

Observed by reading the apply loop, not from an incident: the host re-applies the declaration it holds on a fixed interval (novox/hq ADR 0010 — delivery is a comparison, not a one-shot), which is exactly right for a transient failure (the overlay is not up yet, the registry is briefly unreachable) and makes a permanent failure silent in the way that matters. The failure is present in the node's report and surfaced by push --behind, but only to someone who looks; the machine itself does nothing to raise its hand as it loops on the same impossible resource.

Why it matters beyond the instance

The retry loop exists so a machine converges without a person driving it, and that is the right default. But "keep trying for ever" and "say nothing you are not asked" combine into the failure shape this mesh exists to remove: a machine that reports success at the heartbeat level (it is up, it is reconciling) while a resource inside it has never once applied. The distinction the host draws so carefully within a single apply — reporting every failure, recording only what actually took — is not carried across time: a resource that has failed every reconcile for a week reads the same as one that failed its first attempt a minute ago.

A rule the whole design honours elsewhere is that a thing which cannot make progress must say so rather than look busy. A reconcile that has re-attempted the same resource N times without it ever succeeding is the host's equivalent, and nothing marks it.

Open questions

  • Should a resource that has failed every attempt for some duration or count be escalated — raised distinctly to the control plane, marked in the node's state as "stuck" rather than merely "waiting", so it is visible without someone running push --behind?
  • Is the right signal a count of consecutive failures, an elapsed time since it last succeeded, or the difference between "never once applied" and "applied then regressed"?
  • Does escalation belong in the host (which sees the repetition) or the control plane (which sees every node and could tell one stuck node from a mesh-wide fault)?
  • Does anything change for a resource that is gating (a failed action/run-once stops what follows), where the cost of a permanent failure is not one resource but everything after it?