--- status: proposed date: 2026-08-27 deciders: jochen reconstructed: false extends: 0057-the-host-is-a-root-service-installed-as-a-package.md --- # 59. A host that cannot start is rolled back by the supervisor ## Context [`09-the-node-lifecycle.md`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md) left one thing unsolved: a host upgraded to a version that crashes on start. The supervisor restarts it, backs off, and the node is stuck on a binary that will not run. It was left because automatic recovery looked like *the host judging its own health*, which is the self-reference the rest of that document avoids. **That objection does not survive being asked properly.** A keepalive is not the host judging itself — it is something else judging the host. The question is only *what*, and that has one answer. ### Why the watchdog must be local The obvious candidate is the control plane, and it cannot be: - **Nothing dials a node.** [ADR 0039](0039-the-link-is-the-security-boundary.md) makes the link outbound and node-initiated, and a node has no listening control surface. The mesh has no way to reach in and act. - **The failure removes the reporting path.** A host that cannot start cannot link, so the mesh learns nothing to act on. So the watchdog is local, and the only local thing that is always present, already depended on, and is not the host is the **service manager**. ### The failure this actually prevents Sharper than "the node is down", and it is the reason to bother: > **A host that will not start looks exactly like a machine that was switched off.** [ADR 0036](0036-a-node-is-a-managed-machine.md) makes disconnection ordinary, and [`09`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md) deliberately puts no alarm on it — a laptop shut for three weeks is doing nothing wrong. **A bad upgrade is therefore invisible**: it presents as the one condition the design has decided not to be alarmed by. Worse, it arrives fleet-wide. A declaration naming a bad version reaches every node, and each one applies it, restarts, and stops talking. The mesh would report a fleet of quiet nodes and nothing else. ## Decision **The service manager rolls the host back to the last version that started.** Four parts, and each one is chosen so it works when the host does not: ### 1 — A version is confirmed by starting, not by seeming well On start, the host completes one full reconcile. If it does, it writes the running version to a plain file — `/var/lib/mesh-host/known-good` — and that is the whole of "confirmed". **Deliberately not health.** Not *the link is up*, because a disconnected node is ordinary and a laptop on a train would roll itself back. Not *everything applied cleanly*, because a resource that fails is the machine's problem and not the binary's. The claim is narrow and checkable: **it started, and it got through a reconcile.** ### 2 — The rollback is not the host binary The obvious mistake, and it would make the whole mechanism a no-op: `nox-mesh-host rollback` cannot be the recovery path for a `nox-mesh-host` that does not run. Rollback is a **small script shipped by the package**, which reads the `known-good` file and asks the package manager to install that version. It shares no code with the host and does not import it. ### 3 — The supervisor triggers it, after giving up ```ini [Service] Restart=on-failure StartLimitBurst=3 StartLimitIntervalSec=120 [Unit] OnFailure=nox-mesh-host-rollback.service ``` Three failures in two minutes is a binary that does not work, not a transient. The supervisor stops trying and runs the rollback unit, which downgrades and starts the host again. ### 4 — It rolls back once The rollback unit records that it fired. If the rolled-back version *also* fails to start, it does **not** fire again — the node stops, loudly, in a failed state. **Because a second rollback is a different diagnosis.** Once the previously-working binary also fails, the binary is not the problem: the machine is. Rolling back further would flap between two versions forever and bury the actual cause under a loop. ## Consequences - **The node keeps its own recovery**, which is the property the whole tier-0 design rests on: the host depends on nothing, and now its recovery depends on nothing either. - **The package cache must retain the previous version**, and that is a real requirement rather than an assumption — a package manager configured to clean its cache would delete the thing rollback needs. The package must pin that, and it is the sort of dependency that is discovered by the rollback failing. - **A fleet-wide bad upgrade becomes self-limiting.** Each node fails, rolls back, and comes back on the previous version — so the blast radius of a bad host release is one restart cycle per node rather than the entire fleet stopping. - **It catches "will not start" and nothing else.** A version that starts and is subtly wrong will not roll back, and should not: that is a bad release, which is delivery's problem ([ADR 0058](0058-delivery-ends-in-a-declaration.md)), not a supervision problem. - **It adds a unit and a script the host does not own**, which is [ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)'s line holding: the installation owns the host, and now it owns the host's recovery too. Consistent, and it means both arrive and are versioned together. - **A node in permanent failure is silent, and that is now the last gap.** After a second failure the node is stopped and cannot report it. What notices is the mesh seeing a node that has not been heard from — which `09` records as a fact with no threshold, and which this makes more important to look at than it was. - **The rollback path is exercised only when it is needed**, which is when nobody can afford it to be wrong. It wants a test that boots a lab node onto a deliberately broken host and asserts the previous version comes back. ## References - [ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md) — the restart-by-exiting this protects. - [ADR 0039](0039-the-link-is-the-security-boundary.md) — why the watchdog cannot be remote. - [ADR 0036](0036-a-node-is-a-managed-machine.md) — why the failure is invisible without this. - [`09-the-node-lifecycle.md`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md) — the gap this closes.