Per-OS hosts, and an init asked for only start and restart

0060 -- the host is built per operating system. systemd and pacman are the Arch
host's implementation, not abstractions the mesh has to grow. They are not
independent choices: a machine has pacman because it is Arch, and the package
manager, service manager and packaging format arrive together as one decision
somebody made at install time.

Rejected abstracting them, and the reason is correctness rather than effort.
The service applier reads LoadState to tell "not installed" apart from
"stopped", which is what stops it reporting absence as success. An interface
spanning systemd and OpenRC degrades to what both express, and the lowest
common denominator is exactly where that fault lives.

Almost all of it is shared -- the vocabulary, store, apply loop, read-back
discipline, refusal model, bundle and link are portable. Two appliers differ.
And delivery was already per-OS, since a .pkg.tar.zst is an Arch artifact, so
this is the seam that already existed.

Android is the interesting case rather than Debian: no service manager, no
package installation, usually no root. Such a host implements file, directory
and action and refuses the rest -- the same refusal a host already gives an
unknown type, with a different reason. Those three are the portable floor.

The container runtime is deliberately left open: it is not an OS split, since
Arch runs docker or podman.

0061 -- the init is asked for start-at-boot and restart-on-exit, and nothing
else. Both are expressible in OpenRC, runit, s6 and an Android init.rc.
Counting failed starts and rolling back moves into a launcher, because that is
the one piece which must work when the host does not, and a script with a
counter can be tested where OnFailure= can only be hoped for. Supersedes 0059,
keeping its reasoning in full.

The checker found all six places citing 0059 and refused the commit until they
named the replacement.
This commit is contained in:
2026-08-27 23:46:11 +02:00
parent dcc4b8339c
commit e1ad39b500
5 changed files with 267 additions and 22 deletions
+30 -12
View File
@@ -12,7 +12,7 @@ decisions:
- 02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md
- 02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md
- 02-DECISIONS/0058-delivery-ends-in-a-declaration.md
- 02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md
- 02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md
---
# The node lifecycle
@@ -233,8 +233,8 @@ So the two periodic things do different jobs and should not be conflated:
without a heartbeat that is indistinguishable from a node that stopped. With one, *last heard
from* is a fact beside every node — which is what
[how long disconnected](#how-long-disconnected-and-who-is-told) reports and what
[ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md) exists because
a stuck node cannot send.
[ADR 0061](../../02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md) exists
because a stuck node cannot send.
**Rebooting mid-apply is safe by construction.** The store records each resource *after* it
worked ([ADR 0035](../../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md)), so a host that
@@ -394,11 +394,29 @@ own apply completes. A node must therefore report the version it is **running**,
installed — otherwise the mesh believes an upgrade landed at step 1.
**A version that crashes on start rolls itself back**
([ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md)). The service
manager gives up after three failures in two minutes and runs a rollback script — shipped by the
package rather than being a host subcommand, because a binary that will not start cannot be its
own recovery. It reinstalls the version recorded in `known-good`, which the host wrote the last
time it completed a reconcile.
([ADR 0061](../../02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md)).
What the init starts is not the host but a **launcher**, and the launcher is where the policy
lives:
```
init ──► nox-mesh-host-launch ──► nox-mesh-host
├─ halted? say so and stop; a person has to look
├─ count this start attempt
├─ too many, not yet rolled back? roll back, then start
├─ too many, already rolled back? halt — the machine is the problem
└─ otherwise start the host
```
It reinstalls the version recorded in `known-good`, which the host wrote the last time it
completed a reconcile — and the host clears the attempt counter at the same moment, for the same
reason.
**The launcher rather than the init's own features**, because this is the one thing that must
work on a machine where nothing else does. A shell script with a counter can be run against a
stub package manager and asserted; `OnFailure=` in a unit file can only be read and hoped for.
It also means the init is asked for nothing but *start* and *restart*, which every init can
do.
**It rolls back once.** If the previous version also fails, the node stops in a failed state
rather than flapping between two binaries. A second failure is a different diagnosis: the
@@ -520,10 +538,10 @@ the same command against a mesh that is one machine old.
## Still open
- ~~**Automatic rollback of a bad host version.**~~ **Resolved** by
[ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md): the service
manager gives up after three failures and runs a rollback script — shipped by the package, not
the host binary, because a binary that will not start cannot recover itself. It rolls back
once; a second failure means the machine is the problem, not the binary.
[ADR 0061](../../02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md): a launcher
counts failed starts and rolls back — shipped by the package, not the host binary, because a
binary that will not start cannot recover itself. It rolls back once; a second failure means
the machine is the problem, not the binary.
- **How a previous declaration is retained and chosen**, which is what rollback of anything else
would use ([ADR 0058](../../02-DECISIONS/0058-delivery-ends-in-a-declaration.md)).
- **A node returning after months** applies a very large jump in one go. Correct, and untested.