56 lines
3.3 KiB
Markdown
56 lines
3.3 KiB
Markdown
---
|
|
status: located
|
|
opened: 2026-09-20
|
|
located-in: [mesh-host internal/apply, mesh-controller internal/inventory (node_report)]
|
|
fixed-by:
|
|
amended-design:
|
|
---
|
|
|
|
# 066 — A partly applied declaration leaves a mixed state, with no rollback
|
|
|
|
## Symptom
|
|
|
|
When a host applies a declaration and some resources succeed while others fail, the machine is left
|
|
in a mixed state: the resources that applied stay applied, the ones that failed keep their previous
|
|
value or nothing, and there is no rollback to the state before the apply began. The apply is
|
|
deliberately not a transaction — everything is attempted, every failure is reported, and only what
|
|
actually took is recorded (novox/hq 04-ISSUES/011, which established that one broken resource must
|
|
not hold back the others). The consequence, unstated, is that "this push half-happened" is a state
|
|
the mesh can be in, and nothing names the machines that are in it as distinct from fully-converged
|
|
ones beyond the per-resource failure report.
|
|
|
|
Observed by reading the apply loop, not from an incident. For most resources this is harmless and
|
|
correct: a file and a package on opposite sides of a declaration have nothing to do with each
|
|
other, so applying one and failing the other leaves each in a coherent state of its own. The
|
|
question is the resources where a half-state is not harmless — where two resources are only correct
|
|
together, and applying one without the other is a configuration the module never intended to exist.
|
|
|
|
## Why it matters beyond the instance
|
|
|
|
Gating (a failed action or run-once container stops what follows) covers the case where B must not
|
|
run until A has happened. It does not cover the case where A and B are ordinary state that must
|
|
change *together* — a file and the service that must be restarted for it, a firewall rule and the
|
|
container it admits, two files a program reads as a pair. A push that changes both and applies only
|
|
one leaves the machine in a state that is internally inconsistent, reports partial success, and is
|
|
held there by the reconcile loop re-attempting only the failed half.
|
|
|
|
The design's answer so far is convergence: the next reconcile re-attempts the failed resource, and
|
|
once it succeeds the pair is whole again. That is sound when the failure is transient. It is the
|
|
window that is unexamined — how long a machine may sit half-applied, whether any pairing of
|
|
resources is dangerous enough that a half-state must never be observable, and whether the mesh
|
|
should be able to say "this machine is mid-change" rather than presenting a partial apply as an
|
|
ordinary converged-with-one-failure state.
|
|
|
|
## Open questions
|
|
|
|
- Are there resource pairings whose half-state is actively harmful (a live service reading one of
|
|
two files it needs, a route opened before its backend exists), and if so, does the declaration
|
|
need a way to say "these apply together or not at all"?
|
|
- Is the remedy per-apply atomicity for a declared group, or is it enough to make "this machine is
|
|
partway through a change" a first-class, visible state distinct from "converged, one resource
|
|
failed"?
|
|
- `restart-on` already couples a service to the files it reflects within one apply — is that the
|
|
seed of the grouping primitive, or a different concern?
|
|
- How would any of this be verified — a bed that fails one resource of a coupled pair and asserts
|
|
the machine is never observed serving the inconsistent combination?
|