Files
hq/02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md
T
jschoubben 605c9fd441 Changes are pushed, not polled; and a stuck host rolls itself back
Two corrections and one new decision, all from Jochen catching things.

Pushed, not polled. I described updates as landing "on the next reconcile",
which reads as polling and is not the design. A declaration arrives as a
message on a link that is already open; the host applies it then. Polling over
an existing connection would be slower to land AND constant traffic to learn
nothing.

The timer is for drift and nothing else, and it cannot be replaced by an event
for a definitional reason: drift is change the mesh did not make -- somebody
edited a managed file, a distribution upgrade replaced a config -- so nothing
will ever publish a message about it. Only looking finds it.

Separated the heartbeat from the reconcile timer, which I had been conflating.
They point in opposite directions and answer different questions: the timer
looks at the machine and asks whether it still matches; the heartbeat reports
upward and is what makes silence mean something. A node with nothing to do
sends nothing, and without a heartbeat that is indistinguishable from a node
that stopped.

0059 -- a host that cannot start is rolled back by the service manager. I had
left this open on the grounds that recovery meant the host judging its own
health. That objection does not survive being asked properly: a keepalive is
something else judging the host. The watchdog must be local, because nothing
dials a node and a host that cannot start cannot report -- so it is the service
manager, which is already there.

The failure it prevents is sharper than "the node is down": a host that will
not start looks exactly like a machine somebody switched off, which is the one
condition this design has deliberately decided not to alarm on. So a bad
release reaches every node, each goes quiet, and the mesh reports a fleet of
sleeping laptops.

Confirmed means started and completed one reconcile -- deliberately not "the
link is up", or a laptop on a train would roll itself back. The rollback is a
script shipped by the package, not a host subcommand, because a binary that
will not start cannot be its own recovery. It rolls back once: a second failure
means the machine is the problem, not the binary.

Also refined the records checker, which produced a false positive: a proposed
record may extend another proposed one, because decisions are drafted in chains
and the alternative is marking things accepted to satisfy a check. An accepted
document resting on a proposed record still fails, and that was verified.

0057, 0058 and 0059 are all proposed.
2026-08-27 22:04:26 +02:00

6.3 KiB

status, date, deciders, reconstructed, extends
status date deciders reconstructed extends
proposed 2026-08-27 jochen false 0057-the-host-is-a-root-service-installed-as-a-package.md

59. A host that cannot start is rolled back by the supervisor

Context

09-the-node-lifecycle.md left one thing unsolved: a host upgraded to a version that crashes on start. The supervisor restarts it, backs off, and the node is stuck on a binary that will not run.

It was left because automatic recovery looked like the host judging its own health, which is the self-reference the rest of that document avoids.

That objection does not survive being asked properly. A keepalive is not the host judging itself — it is something else judging the host. The question is only what, and that has one answer.

Why the watchdog must be local

The obvious candidate is the control plane, and it cannot be:

  • Nothing dials a node. ADR 0039 makes the link outbound and node-initiated, and a node has no listening control surface. The mesh has no way to reach in and act.
  • The failure removes the reporting path. A host that cannot start cannot link, so the mesh learns nothing to act on.

So the watchdog is local, and the only local thing that is always present, already depended on, and is not the host is the service manager.

The failure this actually prevents

Sharper than "the node is down", and it is the reason to bother:

A host that will not start looks exactly like a machine that was switched off.

ADR 0036 makes disconnection ordinary, and 09 deliberately puts no alarm on it — a laptop shut for three weeks is doing nothing wrong. A bad upgrade is therefore invisible: it presents as the one condition the design has decided not to be alarmed by.

Worse, it arrives fleet-wide. A declaration naming a bad version reaches every node, and each one applies it, restarts, and stops talking. The mesh would report a fleet of quiet nodes and nothing else.

Decision

The service manager rolls the host back to the last version that started.

Four parts, and each one is chosen so it works when the host does not:

1 — A version is confirmed by starting, not by seeming well

On start, the host completes one full reconcile. If it does, it writes the running version to a plain file — /var/lib/mesh-host/known-good — and that is the whole of "confirmed".

Deliberately not health. Not the link is up, because a disconnected node is ordinary and a laptop on a train would roll itself back. Not everything applied cleanly, because a resource that fails is the machine's problem and not the binary's. The claim is narrow and checkable: it started, and it got through a reconcile.

2 — The rollback is not the host binary

The obvious mistake, and it would make the whole mechanism a no-op: nox-mesh-host rollback cannot be the recovery path for a nox-mesh-host that does not run.

Rollback is a small script shipped by the package, which reads the known-good file and asks the package manager to install that version. It shares no code with the host and does not import it.

3 — The supervisor triggers it, after giving up

[Service]
Restart=on-failure
StartLimitBurst=3
StartLimitIntervalSec=120

[Unit]
OnFailure=nox-mesh-host-rollback.service

Three failures in two minutes is a binary that does not work, not a transient. The supervisor stops trying and runs the rollback unit, which downgrades and starts the host again.

4 — It rolls back once

The rollback unit records that it fired. If the rolled-back version also fails to start, it does not fire again — the node stops, loudly, in a failed state.

Because a second rollback is a different diagnosis. Once the previously-working binary also fails, the binary is not the problem: the machine is. Rolling back further would flap between two versions forever and bury the actual cause under a loop.

Consequences

  • The node keeps its own recovery, which is the property the whole tier-0 design rests on: the host depends on nothing, and now its recovery depends on nothing either.
  • The package cache must retain the previous version, and that is a real requirement rather than an assumption — a package manager configured to clean its cache would delete the thing rollback needs. The package must pin that, and it is the sort of dependency that is discovered by the rollback failing.
  • A fleet-wide bad upgrade becomes self-limiting. Each node fails, rolls back, and comes back on the previous version — so the blast radius of a bad host release is one restart cycle per node rather than the entire fleet stopping.
  • It catches "will not start" and nothing else. A version that starts and is subtly wrong will not roll back, and should not: that is a bad release, which is delivery's problem (ADR 0058), not a supervision problem.
  • It adds a unit and a script the host does not own, which is ADR 0057's line holding: the installation owns the host, and now it owns the host's recovery too. Consistent, and it means both arrive and are versioned together.
  • A node in permanent failure is silent, and that is now the last gap. After a second failure the node is stopped and cannot report it. What notices is the mesh seeing a node that has not been heard from — which 09 records as a fact with no threshold, and which this makes more important to look at than it was.
  • The rollback path is exercised only when it is needed, which is when nobody can afford it to be wrong. It wants a test that boots a lab node onto a deliberately broken host and asserts the previous version comes back.

References