Merge pull request 'Issues 065 and 066 — the apply failure model's two gaps' (#56) from issue/065-066-apply-failure-model into main
This commit was merged in pull request #56.
This commit is contained in:
+50
@@ -0,0 +1,50 @@
|
|||||||
|
---
|
||||||
|
status: open
|
||||||
|
opened: 2026-09-20
|
||||||
|
located-in: []
|
||||||
|
fixed-by:
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 065 — A permanently failing resource is retried for ever with no escalation
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
A resource that can never apply — an image digest nothing will ever publish, a package that does
|
||||||
|
not exist, a file whose content is malformed for whatever reads it — is attempted, fails, is
|
||||||
|
reported, and is then attempted again on the host's reconcile loop every few minutes, indefinitely.
|
||||||
|
Nothing distinguishes "failed once, will succeed when its dependency arrives" from "failed and will
|
||||||
|
fail identically for ever", and nothing escalates the second case.
|
||||||
|
|
||||||
|
Observed by reading the apply loop, not from an incident: the host re-applies the declaration it
|
||||||
|
holds on a fixed interval (novox/hq ADR 0010 — delivery is a comparison, not a one-shot), which is
|
||||||
|
exactly right for a transient failure (the overlay is not up yet, the registry is briefly
|
||||||
|
unreachable) and makes a permanent failure silent in the way that matters. The failure is present
|
||||||
|
in the node's report and surfaced by `push --behind`, but only to someone who looks; the machine
|
||||||
|
itself does nothing to raise its hand as it loops on the same impossible resource.
|
||||||
|
|
||||||
|
## Why it matters beyond the instance
|
||||||
|
|
||||||
|
The retry loop exists so a machine converges without a person driving it, and that is the right
|
||||||
|
default. But "keep trying for ever" and "say nothing you are not asked" combine into the failure
|
||||||
|
shape this mesh exists to remove: a machine that reports success at the heartbeat level (it is up,
|
||||||
|
it is reconciling) while a resource inside it has never once applied. The distinction the host
|
||||||
|
draws so carefully within a single apply — reporting every failure, recording only what actually
|
||||||
|
took — is not carried across time: a resource that has failed every reconcile for a week reads the
|
||||||
|
same as one that failed its first attempt a minute ago.
|
||||||
|
|
||||||
|
A rule the whole design honours elsewhere is that a thing which cannot make progress must say so
|
||||||
|
rather than look busy. A reconcile that has re-attempted the same resource N times without it ever
|
||||||
|
succeeding is the host's equivalent, and nothing marks it.
|
||||||
|
|
||||||
|
## Open questions
|
||||||
|
|
||||||
|
- Should a resource that has failed every attempt for some duration or count be escalated — raised
|
||||||
|
distinctly to the control plane, marked in the node's state as "stuck" rather than merely
|
||||||
|
"waiting", so it is visible without someone running `push --behind`?
|
||||||
|
- Is the right signal a count of consecutive failures, an elapsed time since it last succeeded, or
|
||||||
|
the difference between "never once applied" and "applied then regressed"?
|
||||||
|
- Does escalation belong in the host (which sees the repetition) or the control plane (which sees
|
||||||
|
every node and could tell one stuck node from a mesh-wide fault)?
|
||||||
|
- Does anything change for a resource that is *gating* (a failed action/run-once stops what
|
||||||
|
follows), where the cost of a permanent failure is not one resource but everything after it?
|
||||||
+55
@@ -0,0 +1,55 @@
|
|||||||
|
---
|
||||||
|
status: open
|
||||||
|
opened: 2026-09-20
|
||||||
|
located-in: []
|
||||||
|
fixed-by:
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 066 — A partly applied declaration leaves a mixed state, with no rollback
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
When a host applies a declaration and some resources succeed while others fail, the machine is left
|
||||||
|
in a mixed state: the resources that applied stay applied, the ones that failed keep their previous
|
||||||
|
value or nothing, and there is no rollback to the state before the apply began. The apply is
|
||||||
|
deliberately not a transaction — everything is attempted, every failure is reported, and only what
|
||||||
|
actually took is recorded (novox/hq 04-ISSUES/011, which established that one broken resource must
|
||||||
|
not hold back the others). The consequence, unstated, is that "this push half-happened" is a state
|
||||||
|
the mesh can be in, and nothing names the machines that are in it as distinct from fully-converged
|
||||||
|
ones beyond the per-resource failure report.
|
||||||
|
|
||||||
|
Observed by reading the apply loop, not from an incident. For most resources this is harmless and
|
||||||
|
correct: a file and a package on opposite sides of a declaration have nothing to do with each
|
||||||
|
other, so applying one and failing the other leaves each in a coherent state of its own. The
|
||||||
|
question is the resources where a half-state is not harmless — where two resources are only correct
|
||||||
|
together, and applying one without the other is a configuration the module never intended to exist.
|
||||||
|
|
||||||
|
## Why it matters beyond the instance
|
||||||
|
|
||||||
|
Gating (a failed action or run-once container stops what follows) covers the case where B must not
|
||||||
|
run until A has happened. It does not cover the case where A and B are ordinary state that must
|
||||||
|
change *together* — a file and the service that must be restarted for it, a firewall rule and the
|
||||||
|
container it admits, two files a program reads as a pair. A push that changes both and applies only
|
||||||
|
one leaves the machine in a state that is internally inconsistent, reports partial success, and is
|
||||||
|
held there by the reconcile loop re-attempting only the failed half.
|
||||||
|
|
||||||
|
The design's answer so far is convergence: the next reconcile re-attempts the failed resource, and
|
||||||
|
once it succeeds the pair is whole again. That is sound when the failure is transient. It is the
|
||||||
|
window that is unexamined — how long a machine may sit half-applied, whether any pairing of
|
||||||
|
resources is dangerous enough that a half-state must never be observable, and whether the mesh
|
||||||
|
should be able to say "this machine is mid-change" rather than presenting a partial apply as an
|
||||||
|
ordinary converged-with-one-failure state.
|
||||||
|
|
||||||
|
## Open questions
|
||||||
|
|
||||||
|
- Are there resource pairings whose half-state is actively harmful (a live service reading one of
|
||||||
|
two files it needs, a route opened before its backend exists), and if so, does the declaration
|
||||||
|
need a way to say "these apply together or not at all"?
|
||||||
|
- Is the remedy per-apply atomicity for a declared group, or is it enough to make "this machine is
|
||||||
|
partway through a change" a first-class, visible state distinct from "converged, one resource
|
||||||
|
failed"?
|
||||||
|
- `restart-on` already couples a service to the files it reflects within one apply — is that the
|
||||||
|
seed of the grouping primitive, or a different concern?
|
||||||
|
- How would any of this be verified — a bed that fails one resource of a coupled pair and asserts
|
||||||
|
the machine is never observed serving the inconsistent combination?
|
||||||
Reference in New Issue
Block a user