Files
hq/04-ISSUES/066-a-partly-applied-declaration-leaves-a-mixed-state-with-no-rollback/00-report.md
T
jschoubben a9bc0c1cc5 Open issues 065 and 066 — the apply failure model's two gaps
Surfaced while reviewing the host apply loop for the controller/host
split discussion. 065: a permanently-failing resource retries for ever
with no escalation — reported, but never raises its hand as stuck.
066: a partly-applied declaration leaves a mixed state with no rollback,
which is harmless for independent resources and unexamined for pairs
that are only correct together. Both are questions HQ must answer, not
incidents — status open, no fix proposed.
2026-09-20 14:28:04 +02:00

3.3 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
open 2026-09-20

066 — A partly applied declaration leaves a mixed state, with no rollback

Symptom

When a host applies a declaration and some resources succeed while others fail, the machine is left in a mixed state: the resources that applied stay applied, the ones that failed keep their previous value or nothing, and there is no rollback to the state before the apply began. The apply is deliberately not a transaction — everything is attempted, every failure is reported, and only what actually took is recorded (novox/hq 04-ISSUES/011, which established that one broken resource must not hold back the others). The consequence, unstated, is that "this push half-happened" is a state the mesh can be in, and nothing names the machines that are in it as distinct from fully-converged ones beyond the per-resource failure report.

Observed by reading the apply loop, not from an incident. For most resources this is harmless and correct: a file and a package on opposite sides of a declaration have nothing to do with each other, so applying one and failing the other leaves each in a coherent state of its own. The question is the resources where a half-state is not harmless — where two resources are only correct together, and applying one without the other is a configuration the module never intended to exist.

Why it matters beyond the instance

Gating (a failed action or run-once container stops what follows) covers the case where B must not run until A has happened. It does not cover the case where A and B are ordinary state that must change together — a file and the service that must be restarted for it, a firewall rule and the container it admits, two files a program reads as a pair. A push that changes both and applies only one leaves the machine in a state that is internally inconsistent, reports partial success, and is held there by the reconcile loop re-attempting only the failed half.

The design's answer so far is convergence: the next reconcile re-attempts the failed resource, and once it succeeds the pair is whole again. That is sound when the failure is transient. It is the window that is unexamined — how long a machine may sit half-applied, whether any pairing of resources is dangerous enough that a half-state must never be observable, and whether the mesh should be able to say "this machine is mid-change" rather than presenting a partial apply as an ordinary converged-with-one-failure state.

Open questions

  • Are there resource pairings whose half-state is actively harmful (a live service reading one of two files it needs, a route opened before its backend exists), and if so, does the declaration need a way to say "these apply together or not at all"?
  • Is the remedy per-apply atomicity for a declared group, or is it enough to make "this machine is partway through a change" a first-class, visible state distinct from "converged, one resource failed"?
  • restart-on already couples a service to the files it reflects within one apply — is that the seed of the grouping primitive, or a different concern?
  • How would any of this be verified — a bed that fails one resource of a coupled pair and asserts the machine is never observed serving the inconsistent combination?