From a9bc0c1cc55c0a239faa69bbab9cff1106ee3713 Mon Sep 17 00:00:00 2001 From: jochen Date: Sun, 20 Sep 2026 14:28:04 +0200 Subject: [PATCH] =?UTF-8?q?Open=20issues=20065=20and=20066=20=E2=80=94=20t?= =?UTF-8?q?he=20apply=20failure=20model's=20two=20gaps?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Surfaced while reviewing the host apply loop for the controller/host split discussion. 065: a permanently-failing resource retries for ever with no escalation — reported, but never raises its hand as stuck. 066: a partly-applied declaration leaves a mixed state with no rollback, which is harmless for independent resources and unexamined for pairs that are only correct together. Both are questions HQ must answer, not incidents — status open, no fix proposed. --- .../00-report.md | 50 +++++++++++++++++ .../00-report.md | 55 +++++++++++++++++++ 2 files changed, 105 insertions(+) create mode 100644 04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md create mode 100644 04-ISSUES/066-a-partly-applied-declaration-leaves-a-mixed-state-with-no-rollback/00-report.md diff --git a/04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md b/04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md new file mode 100644 index 0000000..6b77818 --- /dev/null +++ b/04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md @@ -0,0 +1,50 @@ +--- +status: open +opened: 2026-09-20 +located-in: [] +fixed-by: +amended-design: +--- + +# 065 — A permanently failing resource is retried for ever with no escalation + +## Symptom + +A resource that can never apply — an image digest nothing will ever publish, a package that does +not exist, a file whose content is malformed for whatever reads it — is attempted, fails, is +reported, and is then attempted again on the host's reconcile loop every few minutes, indefinitely. +Nothing distinguishes "failed once, will succeed when its dependency arrives" from "failed and will +fail identically for ever", and nothing escalates the second case. + +Observed by reading the apply loop, not from an incident: the host re-applies the declaration it +holds on a fixed interval (novox/hq ADR 0010 — delivery is a comparison, not a one-shot), which is +exactly right for a transient failure (the overlay is not up yet, the registry is briefly +unreachable) and makes a permanent failure silent in the way that matters. The failure is present +in the node's report and surfaced by `push --behind`, but only to someone who looks; the machine +itself does nothing to raise its hand as it loops on the same impossible resource. + +## Why it matters beyond the instance + +The retry loop exists so a machine converges without a person driving it, and that is the right +default. But "keep trying for ever" and "say nothing you are not asked" combine into the failure +shape this mesh exists to remove: a machine that reports success at the heartbeat level (it is up, +it is reconciling) while a resource inside it has never once applied. The distinction the host +draws so carefully within a single apply — reporting every failure, recording only what actually +took — is not carried across time: a resource that has failed every reconcile for a week reads the +same as one that failed its first attempt a minute ago. + +A rule the whole design honours elsewhere is that a thing which cannot make progress must say so +rather than look busy. A reconcile that has re-attempted the same resource N times without it ever +succeeding is the host's equivalent, and nothing marks it. + +## Open questions + +- Should a resource that has failed every attempt for some duration or count be escalated — raised + distinctly to the control plane, marked in the node's state as "stuck" rather than merely + "waiting", so it is visible without someone running `push --behind`? +- Is the right signal a count of consecutive failures, an elapsed time since it last succeeded, or + the difference between "never once applied" and "applied then regressed"? +- Does escalation belong in the host (which sees the repetition) or the control plane (which sees + every node and could tell one stuck node from a mesh-wide fault)? +- Does anything change for a resource that is *gating* (a failed action/run-once stops what + follows), where the cost of a permanent failure is not one resource but everything after it? diff --git a/04-ISSUES/066-a-partly-applied-declaration-leaves-a-mixed-state-with-no-rollback/00-report.md b/04-ISSUES/066-a-partly-applied-declaration-leaves-a-mixed-state-with-no-rollback/00-report.md new file mode 100644 index 0000000..76c65ab --- /dev/null +++ b/04-ISSUES/066-a-partly-applied-declaration-leaves-a-mixed-state-with-no-rollback/00-report.md @@ -0,0 +1,55 @@ +--- +status: open +opened: 2026-09-20 +located-in: [] +fixed-by: +amended-design: +--- + +# 066 — A partly applied declaration leaves a mixed state, with no rollback + +## Symptom + +When a host applies a declaration and some resources succeed while others fail, the machine is left +in a mixed state: the resources that applied stay applied, the ones that failed keep their previous +value or nothing, and there is no rollback to the state before the apply began. The apply is +deliberately not a transaction — everything is attempted, every failure is reported, and only what +actually took is recorded (novox/hq 04-ISSUES/011, which established that one broken resource must +not hold back the others). The consequence, unstated, is that "this push half-happened" is a state +the mesh can be in, and nothing names the machines that are in it as distinct from fully-converged +ones beyond the per-resource failure report. + +Observed by reading the apply loop, not from an incident. For most resources this is harmless and +correct: a file and a package on opposite sides of a declaration have nothing to do with each +other, so applying one and failing the other leaves each in a coherent state of its own. The +question is the resources where a half-state is not harmless — where two resources are only correct +together, and applying one without the other is a configuration the module never intended to exist. + +## Why it matters beyond the instance + +Gating (a failed action or run-once container stops what follows) covers the case where B must not +run until A has happened. It does not cover the case where A and B are ordinary state that must +change *together* — a file and the service that must be restarted for it, a firewall rule and the +container it admits, two files a program reads as a pair. A push that changes both and applies only +one leaves the machine in a state that is internally inconsistent, reports partial success, and is +held there by the reconcile loop re-attempting only the failed half. + +The design's answer so far is convergence: the next reconcile re-attempts the failed resource, and +once it succeeds the pair is whole again. That is sound when the failure is transient. It is the +window that is unexamined — how long a machine may sit half-applied, whether any pairing of +resources is dangerous enough that a half-state must never be observable, and whether the mesh +should be able to say "this machine is mid-change" rather than presenting a partial apply as an +ordinary converged-with-one-failure state. + +## Open questions + +- Are there resource pairings whose half-state is actively harmful (a live service reading one of + two files it needs, a route opened before its backend exists), and if so, does the declaration + need a way to say "these apply together or not at all"? +- Is the remedy per-apply atomicity for a declared group, or is it enough to make "this machine is + partway through a change" a first-class, visible state distinct from "converged, one resource + failed"? +- `restart-on` already couples a service to the files it reflects within one apply — is that the + seed of the grouping primitive, or a different concern? +- How would any of this be verified — a bed that fails one resource of a coupled pair and asserts + the machine is never observed serving the inconsistent combination?