Merge pull request 'Issues 065 and 066 — the apply failure model's two gaps' (#56) from issue/065-066-apply-failure-model into main
This commit was merged in pull request #56.
This commit is contained in:
+50
@@ -0,0 +1,50 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-20
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 065 — A permanently failing resource is retried for ever with no escalation
|
||||
|
||||
## Symptom
|
||||
|
||||
A resource that can never apply — an image digest nothing will ever publish, a package that does
|
||||
not exist, a file whose content is malformed for whatever reads it — is attempted, fails, is
|
||||
reported, and is then attempted again on the host's reconcile loop every few minutes, indefinitely.
|
||||
Nothing distinguishes "failed once, will succeed when its dependency arrives" from "failed and will
|
||||
fail identically for ever", and nothing escalates the second case.
|
||||
|
||||
Observed by reading the apply loop, not from an incident: the host re-applies the declaration it
|
||||
holds on a fixed interval (novox/hq ADR 0010 — delivery is a comparison, not a one-shot), which is
|
||||
exactly right for a transient failure (the overlay is not up yet, the registry is briefly
|
||||
unreachable) and makes a permanent failure silent in the way that matters. The failure is present
|
||||
in the node's report and surfaced by `push --behind`, but only to someone who looks; the machine
|
||||
itself does nothing to raise its hand as it loops on the same impossible resource.
|
||||
|
||||
## Why it matters beyond the instance
|
||||
|
||||
The retry loop exists so a machine converges without a person driving it, and that is the right
|
||||
default. But "keep trying for ever" and "say nothing you are not asked" combine into the failure
|
||||
shape this mesh exists to remove: a machine that reports success at the heartbeat level (it is up,
|
||||
it is reconciling) while a resource inside it has never once applied. The distinction the host
|
||||
draws so carefully within a single apply — reporting every failure, recording only what actually
|
||||
took — is not carried across time: a resource that has failed every reconcile for a week reads the
|
||||
same as one that failed its first attempt a minute ago.
|
||||
|
||||
A rule the whole design honours elsewhere is that a thing which cannot make progress must say so
|
||||
rather than look busy. A reconcile that has re-attempted the same resource N times without it ever
|
||||
succeeding is the host's equivalent, and nothing marks it.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should a resource that has failed every attempt for some duration or count be escalated — raised
|
||||
distinctly to the control plane, marked in the node's state as "stuck" rather than merely
|
||||
"waiting", so it is visible without someone running `push --behind`?
|
||||
- Is the right signal a count of consecutive failures, an elapsed time since it last succeeded, or
|
||||
the difference between "never once applied" and "applied then regressed"?
|
||||
- Does escalation belong in the host (which sees the repetition) or the control plane (which sees
|
||||
every node and could tell one stuck node from a mesh-wide fault)?
|
||||
- Does anything change for a resource that is *gating* (a failed action/run-once stops what
|
||||
follows), where the cost of a permanent failure is not one resource but everything after it?
|
||||
+55
@@ -0,0 +1,55 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-20
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 066 — A partly applied declaration leaves a mixed state, with no rollback
|
||||
|
||||
## Symptom
|
||||
|
||||
When a host applies a declaration and some resources succeed while others fail, the machine is left
|
||||
in a mixed state: the resources that applied stay applied, the ones that failed keep their previous
|
||||
value or nothing, and there is no rollback to the state before the apply began. The apply is
|
||||
deliberately not a transaction — everything is attempted, every failure is reported, and only what
|
||||
actually took is recorded (novox/hq 04-ISSUES/011, which established that one broken resource must
|
||||
not hold back the others). The consequence, unstated, is that "this push half-happened" is a state
|
||||
the mesh can be in, and nothing names the machines that are in it as distinct from fully-converged
|
||||
ones beyond the per-resource failure report.
|
||||
|
||||
Observed by reading the apply loop, not from an incident. For most resources this is harmless and
|
||||
correct: a file and a package on opposite sides of a declaration have nothing to do with each
|
||||
other, so applying one and failing the other leaves each in a coherent state of its own. The
|
||||
question is the resources where a half-state is not harmless — where two resources are only correct
|
||||
together, and applying one without the other is a configuration the module never intended to exist.
|
||||
|
||||
## Why it matters beyond the instance
|
||||
|
||||
Gating (a failed action or run-once container stops what follows) covers the case where B must not
|
||||
run until A has happened. It does not cover the case where A and B are ordinary state that must
|
||||
change *together* — a file and the service that must be restarted for it, a firewall rule and the
|
||||
container it admits, two files a program reads as a pair. A push that changes both and applies only
|
||||
one leaves the machine in a state that is internally inconsistent, reports partial success, and is
|
||||
held there by the reconcile loop re-attempting only the failed half.
|
||||
|
||||
The design's answer so far is convergence: the next reconcile re-attempts the failed resource, and
|
||||
once it succeeds the pair is whole again. That is sound when the failure is transient. It is the
|
||||
window that is unexamined — how long a machine may sit half-applied, whether any pairing of
|
||||
resources is dangerous enough that a half-state must never be observable, and whether the mesh
|
||||
should be able to say "this machine is mid-change" rather than presenting a partial apply as an
|
||||
ordinary converged-with-one-failure state.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Are there resource pairings whose half-state is actively harmful (a live service reading one of
|
||||
two files it needs, a route opened before its backend exists), and if so, does the declaration
|
||||
need a way to say "these apply together or not at all"?
|
||||
- Is the remedy per-apply atomicity for a declared group, or is it enough to make "this machine is
|
||||
partway through a change" a first-class, visible state distinct from "converged, one resource
|
||||
failed"?
|
||||
- `restart-on` already couples a service to the files it reflects within one apply — is that the
|
||||
seed of the grouping primitive, or a different concern?
|
||||
- How would any of this be verified — a bed that fails one resource of a coupled pair and asserts
|
||||
the machine is never observed serving the inconsistent combination?
|
||||
Reference in New Issue
Block a user