Files
hq/04-ISSUES/066-a-partly-applied-declaration-leaves-a-mixed-state-with-no-rollback/00-report.md
T

56 lines
3.3 KiB
Markdown

---
status: located
opened: 2026-09-20
located-in: [mesh-host internal/apply, mesh-controller internal/inventory (node_report)]
fixed-by:
amended-design:
---
# 066 — A partly applied declaration leaves a mixed state, with no rollback
## Symptom
When a host applies a declaration and some resources succeed while others fail, the machine is left
in a mixed state: the resources that applied stay applied, the ones that failed keep their previous
value or nothing, and there is no rollback to the state before the apply began. The apply is
deliberately not a transaction — everything is attempted, every failure is reported, and only what
actually took is recorded (novox/hq 04-ISSUES/011, which established that one broken resource must
not hold back the others). The consequence, unstated, is that "this push half-happened" is a state
the mesh can be in, and nothing names the machines that are in it as distinct from fully-converged
ones beyond the per-resource failure report.
Observed by reading the apply loop, not from an incident. For most resources this is harmless and
correct: a file and a package on opposite sides of a declaration have nothing to do with each
other, so applying one and failing the other leaves each in a coherent state of its own. The
question is the resources where a half-state is not harmless — where two resources are only correct
together, and applying one without the other is a configuration the module never intended to exist.
## Why it matters beyond the instance
Gating (a failed action or run-once container stops what follows) covers the case where B must not
run until A has happened. It does not cover the case where A and B are ordinary state that must
change *together* — a file and the service that must be restarted for it, a firewall rule and the
container it admits, two files a program reads as a pair. A push that changes both and applies only
one leaves the machine in a state that is internally inconsistent, reports partial success, and is
held there by the reconcile loop re-attempting only the failed half.
The design's answer so far is convergence: the next reconcile re-attempts the failed resource, and
once it succeeds the pair is whole again. That is sound when the failure is transient. It is the
window that is unexamined — how long a machine may sit half-applied, whether any pairing of
resources is dangerous enough that a half-state must never be observable, and whether the mesh
should be able to say "this machine is mid-change" rather than presenting a partial apply as an
ordinary converged-with-one-failure state.
## Open questions
- Are there resource pairings whose half-state is actively harmful (a live service reading one of
two files it needs, a route opened before its backend exists), and if so, does the declaration
need a way to say "these apply together or not at all"?
- Is the remedy per-apply atomicity for a declared group, or is it enough to make "this machine is
partway through a change" a first-class, visible state distinct from "converged, one resource
failed"?
- `restart-on` already couples a service to the files it reflects within one apply — is that the
seed of the grouping primitive, or a different concern?
- How would any of this be verified — a bed that fails one resource of a coupled pair and asserts
the machine is never observed serving the inconsistent combination?