Merge pull request 'Issues 065 and 066 — the apply failure model's two gaps' (#56) from issue/065-066-apply-failure-model into main

This commit was merged in pull request #56.
This commit is contained in:
2026-09-20 14:28:18 +02:00
2 changed files with 105 additions and 0 deletions
@@ -0,0 +1,50 @@
---
status: open
opened: 2026-09-20
located-in: []
fixed-by:
amended-design:
---
# 065 — A permanently failing resource is retried for ever with no escalation
## Symptom
A resource that can never apply — an image digest nothing will ever publish, a package that does
not exist, a file whose content is malformed for whatever reads it — is attempted, fails, is
reported, and is then attempted again on the host's reconcile loop every few minutes, indefinitely.
Nothing distinguishes "failed once, will succeed when its dependency arrives" from "failed and will
fail identically for ever", and nothing escalates the second case.
Observed by reading the apply loop, not from an incident: the host re-applies the declaration it
holds on a fixed interval (novox/hq ADR 0010 — delivery is a comparison, not a one-shot), which is
exactly right for a transient failure (the overlay is not up yet, the registry is briefly
unreachable) and makes a permanent failure silent in the way that matters. The failure is present
in the node's report and surfaced by `push --behind`, but only to someone who looks; the machine
itself does nothing to raise its hand as it loops on the same impossible resource.
## Why it matters beyond the instance
The retry loop exists so a machine converges without a person driving it, and that is the right
default. But "keep trying for ever" and "say nothing you are not asked" combine into the failure
shape this mesh exists to remove: a machine that reports success at the heartbeat level (it is up,
it is reconciling) while a resource inside it has never once applied. The distinction the host
draws so carefully within a single apply — reporting every failure, recording only what actually
took — is not carried across time: a resource that has failed every reconcile for a week reads the
same as one that failed its first attempt a minute ago.
A rule the whole design honours elsewhere is that a thing which cannot make progress must say so
rather than look busy. A reconcile that has re-attempted the same resource N times without it ever
succeeding is the host's equivalent, and nothing marks it.
## Open questions
- Should a resource that has failed every attempt for some duration or count be escalated — raised
distinctly to the control plane, marked in the node's state as "stuck" rather than merely
"waiting", so it is visible without someone running `push --behind`?
- Is the right signal a count of consecutive failures, an elapsed time since it last succeeded, or
the difference between "never once applied" and "applied then regressed"?
- Does escalation belong in the host (which sees the repetition) or the control plane (which sees
every node and could tell one stuck node from a mesh-wide fault)?
- Does anything change for a resource that is *gating* (a failed action/run-once stops what
follows), where the cost of a permanent failure is not one resource but everything after it?
@@ -0,0 +1,55 @@
---
status: open
opened: 2026-09-20
located-in: []
fixed-by:
amended-design:
---
# 066 — A partly applied declaration leaves a mixed state, with no rollback
## Symptom
When a host applies a declaration and some resources succeed while others fail, the machine is left
in a mixed state: the resources that applied stay applied, the ones that failed keep their previous
value or nothing, and there is no rollback to the state before the apply began. The apply is
deliberately not a transaction — everything is attempted, every failure is reported, and only what
actually took is recorded (novox/hq 04-ISSUES/011, which established that one broken resource must
not hold back the others). The consequence, unstated, is that "this push half-happened" is a state
the mesh can be in, and nothing names the machines that are in it as distinct from fully-converged
ones beyond the per-resource failure report.
Observed by reading the apply loop, not from an incident. For most resources this is harmless and
correct: a file and a package on opposite sides of a declaration have nothing to do with each
other, so applying one and failing the other leaves each in a coherent state of its own. The
question is the resources where a half-state is not harmless — where two resources are only correct
together, and applying one without the other is a configuration the module never intended to exist.
## Why it matters beyond the instance
Gating (a failed action or run-once container stops what follows) covers the case where B must not
run until A has happened. It does not cover the case where A and B are ordinary state that must
change *together* — a file and the service that must be restarted for it, a firewall rule and the
container it admits, two files a program reads as a pair. A push that changes both and applies only
one leaves the machine in a state that is internally inconsistent, reports partial success, and is
held there by the reconcile loop re-attempting only the failed half.
The design's answer so far is convergence: the next reconcile re-attempts the failed resource, and
once it succeeds the pair is whole again. That is sound when the failure is transient. It is the
window that is unexamined — how long a machine may sit half-applied, whether any pairing of
resources is dangerous enough that a half-state must never be observable, and whether the mesh
should be able to say "this machine is mid-change" rather than presenting a partial apply as an
ordinary converged-with-one-failure state.
## Open questions
- Are there resource pairings whose half-state is actively harmful (a live service reading one of
two files it needs, a route opened before its backend exists), and if so, does the declaration
need a way to say "these apply together or not at all"?
- Is the remedy per-apply atomicity for a declared group, or is it enough to make "this machine is
partway through a change" a first-class, visible state distinct from "converged, one resource
failed"?
- `restart-on` already couples a service to the files it reflects within one apply — is that the
seed of the grouping primitive, or a different concern?
- How would any of this be verified — a bed that fails one resource of a coupled pair and asserts
the machine is never observed serving the inconsistent combination?