0060 -- the host is built per operating system. systemd and pacman are the Arch host's implementation, not abstractions the mesh has to grow. They are not independent choices: a machine has pacman because it is Arch, and the package manager, service manager and packaging format arrive together as one decision somebody made at install time. Rejected abstracting them, and the reason is correctness rather than effort. The service applier reads LoadState to tell "not installed" apart from "stopped", which is what stops it reporting absence as success. An interface spanning systemd and OpenRC degrades to what both express, and the lowest common denominator is exactly where that fault lives. Almost all of it is shared -- the vocabulary, store, apply loop, read-back discipline, refusal model, bundle and link are portable. Two appliers differ. And delivery was already per-OS, since a .pkg.tar.zst is an Arch artifact, so this is the seam that already existed. Android is the interesting case rather than Debian: no service manager, no package installation, usually no root. Such a host implements file, directory and action and refuses the rest -- the same refusal a host already gives an unknown type, with a different reason. Those three are the portable floor. The container runtime is deliberately left open: it is not an OS split, since Arch runs docker or podman. 0061 -- the init is asked for start-at-boot and restart-on-exit, and nothing else. Both are expressible in OpenRC, runit, s6 and an Android init.rc. Counting failed starts and rolling back moves into a launcher, because that is the one piece which must work when the host does not, and a script with a counter can be tested where OnFailure= can only be hoped for. Supersedes 0059, keeping its reasoning in full. The checker found all six places citing 0059 and refused the commit until they named the replacement.
190 lines
9.5 KiB
Markdown
190 lines
9.5 KiB
Markdown
---
|
|
status: superseded
|
|
superseded-by: 02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md
|
|
date: 2026-08-27
|
|
deciders: jochen
|
|
reconstructed: false
|
|
extends: 0057-the-host-is-a-root-service-installed-as-a-package.md
|
|
---
|
|
|
|
# 59. Two watchdogs: the mesh stages the rollout, the supervisor recovers the node
|
|
|
|
## Context
|
|
|
|
[`09-the-node-lifecycle.md`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md) left one thing
|
|
unsolved: a host upgraded to a version that crashes on start. The supervisor restarts it, backs
|
|
off, and the node is stuck on a binary that will not run.
|
|
|
|
It was left because automatic recovery looked like *the host judging its own health*, which is
|
|
the self-reference the rest of that document avoids.
|
|
|
|
**That objection does not survive being asked properly.** A keepalive is not the host judging
|
|
itself — it is something else judging the host.
|
|
|
|
### There are two watchdogs, and they cannot do each other's job
|
|
|
|
A first draft of this record concluded the watchdog must be local, and stopped there. That was
|
|
half an answer: it is true that recovery must be local, and false that the mesh has no part.
|
|
|
|
| | can see | can act |
|
|
|---|---|---|
|
|
| **the service manager**, on the node | that *this* process keeps dying | **yes** — restart it, replace it |
|
|
| **the mesh** | that *eleven of twelve nodes* went quiet after one declaration | **no** — nothing dials a node |
|
|
|
|
**Recovery must be local**, and that half stands:
|
|
|
|
- **Nothing dials a node.** [ADR 0039](0039-the-link-is-the-security-boundary.md) makes the link
|
|
outbound and node-initiated, with no listening control surface. The mesh has no way to reach
|
|
in and act.
|
|
- **The failure removes the reporting path.** A host that cannot start cannot link, so the mesh
|
|
learns nothing *from that node* to act on.
|
|
|
|
**But detection is the mesh's**, and it is the half a local watchdog structurally cannot do. A
|
|
node's supervisor sees one process failing and has no idea whether that is a broken machine or a
|
|
broken release. **Only something watching every node can tell those apart** — and telling them
|
|
apart is what decides whether the right response is *fix this machine* or *stop shipping this
|
|
version immediately*.
|
|
|
|
### The failure this actually prevents
|
|
|
|
Sharper than "the node is down", and it is the reason to bother:
|
|
|
|
> **A host that will not start looks exactly like a machine that was switched off.**
|
|
|
|
[ADR 0036](0036-a-node-is-a-managed-machine.md) makes disconnection ordinary, and
|
|
[`09`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md) deliberately puts no alarm on it — a
|
|
laptop shut for three weeks is doing nothing wrong. **A bad upgrade is therefore invisible**: it
|
|
presents as the one condition the design has decided not to be alarmed by.
|
|
|
|
Worse, it arrives fleet-wide. A declaration naming a bad version reaches every node, and each
|
|
one applies it, restarts, and stops talking. The mesh would report a fleet of quiet nodes and
|
|
nothing else.
|
|
|
|
## Decision
|
|
|
|
**The mesh stages the rollout and stops when nodes go quiet. The service manager recovers the
|
|
node it is on.** Prevention and recovery, and neither substitutes for the other.
|
|
|
|
### The mesh stages a host rollout
|
|
|
|
A host version does not reach every node at once. The delivery context updates a few nodes'
|
|
declarations, **waits for those nodes to heartbeat on the new version**, and only then continues.
|
|
|
|
```
|
|
update 2 nodes ─► heard from both, running the new version ─► continue
|
|
└► silence past the window ─► STOP. Report.
|
|
```
|
|
|
|
**Silence is the signal, and it is available because of the heartbeat**
|
|
([ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)). A node that upgraded
|
|
and cannot start stops reporting; that is indistinguishable from a switched-off machine *for one
|
|
node*, and completely distinguishable across a batch that was all told the same thing at the
|
|
same time.
|
|
|
|
**This is what keeps a bad release from becoming a fleet outage.** Local rollback repairs a node
|
|
after the fact; staging means most nodes never receive the bad version at all. A stopped rollout
|
|
is two broken nodes and a report, rather than every node quiet at once.
|
|
|
|
**It does not replace local recovery**, for two reasons. The canary nodes still break, and
|
|
somebody has to be able to fix them. And a node that was offline during the staged rollout gets
|
|
the declaration when it reconnects, with no batch around it and nothing watching — so it must be
|
|
able to recover alone.
|
|
|
|
### On the node: the service manager rolls back
|
|
|
|
Four parts, and each one is chosen so it works when the host does not:
|
|
|
|
### 1 — A version is confirmed by starting, not by seeming well
|
|
|
|
On start, the host completes one full reconcile. If it does, it writes the running version to a
|
|
plain file — `/var/lib/mesh-host/known-good` — and that is the whole of "confirmed".
|
|
|
|
**Deliberately not health.** Not *the link is up*, because a disconnected node is ordinary and a
|
|
laptop on a train would roll itself back. Not *everything applied cleanly*, because a resource
|
|
that fails is the machine's problem and not the binary's. The claim is narrow and checkable: **it
|
|
started, and it got through a reconcile.**
|
|
|
|
### 2 — The rollback is not the host binary
|
|
|
|
The obvious mistake, and it would make the whole mechanism a no-op: `nox-mesh-host rollback`
|
|
cannot be the recovery path for a `nox-mesh-host` that does not run.
|
|
|
|
Rollback is a **small script shipped by the package**, which reads the `known-good` file and
|
|
asks the package manager to install that version. It shares no code with the host and does not
|
|
import it.
|
|
|
|
### 3 — The supervisor triggers it, after giving up
|
|
|
|
```ini
|
|
[Service]
|
|
Restart=always
|
|
StartLimitBurst=3
|
|
StartLimitIntervalSec=120
|
|
|
|
[Unit]
|
|
OnFailure=nox-mesh-host-rollback.service
|
|
```
|
|
|
|
**`Restart=always`, not `on-failure`**, and the difference is load-bearing rather than a
|
|
preference. [ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md) has the host
|
|
restart onto a new binary by **exiting cleanly** — and `on-failure` does not restart a process
|
|
that exited zero. An earlier draft of this record specified `on-failure` and would have left
|
|
every upgraded node stopped, having successfully upgraded. Caught by reading the two records
|
|
against each other rather than by either alone.
|
|
|
|
Three failures in two minutes is a binary that does not work, not a transient. The supervisor
|
|
stops trying, the unit enters a failed state, and `OnFailure` runs the rollback unit — which
|
|
downgrades and starts the host again.
|
|
|
|
### 4 — With nothing to roll back to, it does not try
|
|
|
|
A machine whose host has *never* completed a reconcile has no `known-good`. The rollback unit
|
|
finds nothing, does nothing, and says so.
|
|
|
|
That is the right outcome: there is no previous version, so the node was never working, and the
|
|
failure belongs to the installation rather than to an upgrade. Attempting a rollback here would
|
|
mean guessing at a version, which is how a recovery mechanism becomes a second fault.
|
|
|
|
### 5 — It rolls back once
|
|
|
|
The rollback unit records that it fired. If the rolled-back version *also* fails to start, it
|
|
does **not** fire again — the node stops, loudly, in a failed state.
|
|
|
|
**Because a second rollback is a different diagnosis.** Once the previously-working binary also
|
|
fails, the binary is not the problem: the machine is. Rolling back further would flap between
|
|
two versions forever and bury the actual cause under a loop.
|
|
|
|
## Consequences
|
|
|
|
- **The node keeps its own recovery**, which is the property the whole tier-0 design rests on:
|
|
the host depends on nothing, and now its recovery depends on nothing either.
|
|
- **The package cache must retain the previous version**, and that is a real requirement rather
|
|
than an assumption — a package manager configured to clean its cache would delete the thing
|
|
rollback needs. The package must pin that, and it is the sort of dependency that is discovered
|
|
by the rollback failing.
|
|
- **A fleet-wide bad upgrade becomes self-limiting.** Each node fails, rolls back, and comes
|
|
back on the previous version — so the blast radius of a bad host release is one restart cycle
|
|
per node rather than the entire fleet stopping.
|
|
- **It catches "will not start" and nothing else.** A version that starts and is subtly wrong
|
|
will not roll back, and should not: that is a bad release, which is delivery's problem
|
|
([ADR 0058](0058-delivery-ends-in-a-declaration.md)), not a supervision problem.
|
|
- **It adds a unit and a script the host does not own**, which is
|
|
[ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)'s line holding: the
|
|
installation owns the host, and now it owns the host's recovery too. Consistent, and it means
|
|
both arrive and are versioned together.
|
|
- **A node in permanent failure is silent, and that is now the last gap.** After a second
|
|
failure the node is stopped and cannot report it. What notices is the mesh seeing a node that
|
|
has not been heard from — which `09` records as a fact with no threshold, and which this makes
|
|
more important to look at than it was.
|
|
- **The rollback path is exercised only when it is needed**, which is when nobody can afford it
|
|
to be wrong. It wants a test that boots a lab node onto a deliberately broken host and asserts
|
|
the previous version comes back.
|
|
|
|
## References
|
|
|
|
- [ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md) — the restart-by-exiting
|
|
this protects.
|
|
- [ADR 0039](0039-the-link-is-the-security-boundary.md) — why the watchdog cannot be remote.
|
|
- [ADR 0036](0036-a-node-is-a-managed-machine.md) — why the failure is invisible without this.
|
|
- [`09-the-node-lifecycle.md`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md) — the gap this closes.
|