diff --git a/02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md b/02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md index a957f12..8c3c892 100644 --- a/02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md +++ b/02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md @@ -1,5 +1,6 @@ --- -status: accepted +status: superseded +superseded-by: 02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md date: 2026-08-27 deciders: jochen reconstructed: false diff --git a/02-DECISIONS/0060-the-host-is-built-per-operating-system.md b/02-DECISIONS/0060-the-host-is-built-per-operating-system.md new file mode 100644 index 0000000..83c79ee --- /dev/null +++ b/02-DECISIONS/0060-the-host-is-built-per-operating-system.md @@ -0,0 +1,112 @@ +--- +status: accepted +date: 2026-08-28 +deciders: jochen +reconstructed: false +extends: 0057-the-host-is-a-root-service-installed-as-a-package.md +--- + +# 60. The host is built per operating system + +## Context + +Three of the host's six shapes need something from the machine: `service` needs a service +manager, `package` needs a package manager, `container` needs a container runtime. The other +three — `file`, `directory`, `action` — need only a filesystem and the ability to run something. + +The host names those capabilities generically and implements them specifically. The detector +reports `container-runtime`, `package-manager`, `service-manager`; the appliers call `docker`, +`pacman` and `systemctl`. **So the design says capability and the code says Arch**, and nothing +records which of those is the intent. + +The question that surfaced it: what happens on a machine that has podman, or one that does not +run systemd? And underneath it, a live contradiction — +[research 012](../01-RESEARCH/012-the-minimum-viable-node/00-overview.md) decides that on +conflict during adoption *the machine's configuration is kept*, so a machine with podman keeps +podman, and then the container applier calls `docker` and fails. + +## Considered options + +1. **Abstract each capability behind an interface.** One host, adapters per service manager and + package manager. Rejected, and the reason is correctness rather than effort: the service + applier reads `LoadState` to tell *not installed* apart from *stopped*, which is what stops it + reporting absence as success. An interface spanning systemd and OpenRC degrades to what both + can express, and **the lowest common denominator is exactly where that fault lives**. +2. **Support one operating system and say so.** Honest, and it makes every other machine + permanently out of scope rather than not-yet. +3. **A host per operating system.** Chosen. + +## Decision + +**The host is built for an operating system family, and `systemd` and `pacman` are the Arch +host's implementation rather than abstractions the mesh has to grow.** + +``` +mesh-host-arch-x86_64 pacman · systemctl · docker +mesh-host-debian-x86_64 apt · systemctl · docker (when there is a machine) +mesh-host-alpine-x86_64 apk · rc-service · podman (when there is a machine) +``` + +**These are not independent choices and treating them as such was the error.** A machine has +pacman *because* it is Arch. The package manager, the service manager and the packaging format +arrive together, as one decision somebody made when they installed the operating system. + +### Almost all of it is shared + +Not a rewrite per operating system. The declaration vocabulary, the store, the apply loop, the +read-back discipline, the refusal model, the bundle and the link are all portable. **What differs +is two appliers**, and the rest is compiled around them. + +### The control plane names the package, because the host does not decide + +A container runtime is `docker` on Arch and `docker.io` on Debian. Mapping *this node needs a +container runtime* to a package name is **deciding**, which +[ADR 0037](0037-the-host-applies-it-does-not-decide.md) puts outside the host. + +It needs no new mechanism: the profile already reports what the machine is, so the declaration a +node receives is already tailored to that node. The host receives a package name and installs it. + +### A host that cannot implement a shape refuses it + +The interesting case is not Debian, it is **Android** — no service manager it will lend us, no +package installation, usually no root. Such a host implements `file`, `directory` and `action`, +and nothing else. + +That needs no new mechanism either. A host already refuses a type it does not know; *this host +does not implement `package`* is the same refusal with a different reason, and the profile +reports which shapes it implements so the control plane never sends one it cannot do. + +**`file`, `directory` and `action` are the portable floor.** They work anywhere there is a +filesystem and a way to run something, which makes a partial host a real thing rather than a +broken one. + +## Consequences + +- **Each implementation stays as sharp as its operating system allows.** The `LoadState` + distinction survives because the Arch host knows it is systemd. Nothing is degraded to fit an + interface spanning systems we do not run. +- **Delivery already worked this way**, which is the strongest sign this is the right seam: the + host ships as a package from the mesh's own repository + ([ADR 0058](0058-delivery-ends-in-a-declaration.md)), and a `.pkg.tar.zst` is an Arch artifact. + A per-OS binary is consistent with what was already decided rather than an addition to it. +- **The container runtime is left unresolved on purpose**, because it is not an OS split: Arch + runs docker or podman. Per-OS hosts answer the service and package managers and leave the + runtime a genuine choice *within* a host. That is a separate decision. +- **A second operating system is now additive rather than a redesign** — write two appliers, ship + a package. And it will be designed against a real machine rather than a guess, which is the + point of not building the abstraction now. +- **The profile has to report the operating system**, and today it reports capabilities without + saying which system they belong to. Small, and needed before the control plane can tailor a + package name. +- **Nothing states the machine requirements yet.** A machine missing a capability fails at apply + time rather than being refused up front, even though the host already detects it. That is a + gap this record makes visible and does not close. + +## References + +- [ADR 0037](0037-the-host-applies-it-does-not-decide.md) — why the host is handed a package name. +- [ADR 0041](0041-the-host-depends-on-nothing.md) — one static binary, now per system as well as + per architecture. +- [ADR 0058](0058-delivery-ends-in-a-declaration.md) — delivery, which was already per-OS. +- [Research 012](../01-RESEARCH/012-the-minimum-viable-node/00-overview.md) — adoption keeping + the machine's configuration, which hardcoding a runtime contradicts. diff --git a/02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md b/02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md new file mode 100644 index 0000000..dcd2578 --- /dev/null +++ b/02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md @@ -0,0 +1,109 @@ +--- +status: accepted +date: 2026-08-28 +deciders: jochen +reconstructed: false +supersedes: 0059-a-host-that-cannot-start-rolls-itself-back.md +extends: 0060-the-host-is-built-per-operating-system.md +--- + +# 61. The host asks an init for start and restart, and nothing else + +## Context + +[ADR 0059](0059-a-host-that-cannot-start-rolls-itself-back.md) built the node's recovery out of +systemd's own features: `StartLimitBurst` to decide a binary is broken, `OnFailure` to run a +rollback unit. It works, and it makes recovery — the thing that matters most when a node is +stuck — the most systemd-specific part of the whole host. + +[ADR 0060](0060-the-host-is-built-per-operating-system.md) makes the host per operating system, +which raises the obvious question: how much of an init does the host actually need? + +Counted honestly, three things, and only one of them is special: + +| | any init? | +|---|---| +| start at boot | **yes** | +| restart it when it exits | **yes** | +| give up after N failures and run something else | **no** — that is systemd's `StartLimitBurst` and `OnFailure` | + +So the recovery mechanism is the only reason the host needs *this* init rather than *an* init. +And it is the part that must work on a machine where the host does not, which makes "it is +expressed in unit-file syntax" a poor place for it: unit syntax is not something we can test, and +the one time it runs is the one time nobody can afford it to be wrong. + +**The substrate does not need systemd either**, which is what makes this worth doing rather than +merely tidy. The bootstrap is package → container → action → container, and every mesh workload +is a container the runtime restarts. Nothing in it declares a `service`. + +## Decision + +**An init is asked for two things: start this at boot, and start it again if it exits.** Both are +expressible in systemd, OpenRC, runit, s6 and an Android `init.rc`. + +**Everything else moves into a launcher**, which is what the init actually starts: + +``` +init ──► nox-mesh-host-launch ──► nox-mesh-host + │ + ├─ halted? say so and stop. a person has to look + ├─ count this start attempt + ├─ too many, and not yet rolled back? roll back, then start + ├─ too many, and already rolled back? halt — the machine is the problem + └─ otherwise start the host + +host, on a completed reconcile ──► clears the counter, records known-good +``` + +**This keeps everything ADR 0059 decided and changes only where it lives.** Two watchdogs still, +and neither substitutes for the other: the mesh stages a host rollout and stops when nodes go +quiet; the node recovers itself. Recovery is still local, because nothing dials a node and a host +that cannot start cannot report. It rolls back once, because a second failure of a +previously-working binary is a different diagnosis. The rollback still shares no code with the +host, because a binary that will not start cannot be its own recovery. + +### What is gained by moving it + +- **It becomes testable.** A shell script with a counter can be run against a stub package + manager and asserted, which is how the rollback script is already tested. `OnFailure=` can be + read and hoped for. +- **The host becomes runnable under any init**, which is what makes an Alpine or Android host + possible later rather than blocked on porting the recovery. +- **The give-up policy stops being configuration and becomes code we own.** Three attempts is a + decision, and it should live where decisions are read and tested. + +### What it costs + +- **One more process in the chain**, and it runs before the host on every start. +- **The counter is the whole mechanism, and it is the part to get right.** Never cleared, and + the node rolls back on a healthy boot; cleared too eagerly, and it never rolls back at all. + It is cleared by the host on a **completed reconcile** — the same event that records + known-good, for the same reason. +- **The launcher is a thing that can itself be broken**, and nothing recovers it. That is one + turtle down from where we were, not zero: the alternative was unit syntax, which also cannot + recover itself and additionally cannot be tested. + +## Consequences + +- **`Restart=always` stays load-bearing and stays subtle.** The host restarts onto a new binary + by exiting cleanly ([ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)), so + whatever supervises must restart on a zero exit. This caught out an earlier draft of 0059, + which specified `on-failure` and would have left every upgraded node stopped. +- **The unit file becomes trivial**, which is the point: start, restart, a state directory. + Nothing in it encodes policy, so porting it is transcription rather than design. +- **A halted node is silent**, unchanged from 0059 and still the last gap. What notices is the + mesh seeing a node it has not heard from. +- **The rollback script grows into a launcher** rather than being replaced. Its tested behaviour — + roll back once, refuse to guess with no known-good, fail loudly when the package is not + cached — carries over and is where the new counter logic joins it. +- **`service` survives as a shape**, and this record does not remove it. Almost nothing in the + design declares one, but adoption takes over machines already in use whose units somebody + chose, and saying the mesh may never manage those is a larger decision than this one. + +## References + +- [ADR 0059](0059-a-host-that-cannot-start-rolls-itself-back.md) — superseded; its reasoning + about two watchdogs, rolling back once, and recovery being local is kept in full. +- [ADR 0060](0060-the-host-is-built-per-operating-system.md) — why this question was asked. +- [ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md) — restart-by-exiting, + which constrains what the init must do. diff --git a/03-DESIGN/01-to-be/05-the-node-host.md b/03-DESIGN/01-to-be/05-the-node-host.md index c31cba0..78dd62d 100644 --- a/03-DESIGN/01-to-be/05-the-node-host.md +++ b/03-DESIGN/01-to-be/05-the-node-host.md @@ -13,7 +13,7 @@ decisions: - 02-DECISIONS/0041-the-host-depends-on-nothing.md - 02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md - 02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md - - 02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md + - 02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md --- # The node host @@ -162,23 +162,28 @@ After=network-online.target Wants=network-online.target [Service] -Type=notify -ExecStart=/usr/bin/nox-mesh-host run +ExecStart=/usr/lib/nox-mesh-host/launch Restart=always RestartSec=5s -StartLimitBurst=3 -StartLimitIntervalSec=120 StateDirectory=mesh-host [Install] WantedBy=multi-user.target ``` +**Two lines of policy, and that is deliberate** +([ADR 0061](../../02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md)). The init is +asked to *start this at boot* and *start it again if it exits*, and nothing else. Both are +expressible in OpenRC, runit, s6 and an Android `init.rc`, so porting this file is transcription +rather than design. + **`Restart=always` and not `on-failure`**: the host restarts onto a new binary by exiting -*cleanly*, and `on-failure` would not restart it — an upgraded node would be left stopped, -having successfully upgraded. The start limit is what -[ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md) uses to decide -a binary is broken rather than unlucky. +*cleanly*, so a supervisor that only restarts on failure would leave every upgraded node stopped, +having successfully upgraded. + +**What the init does not do is decide when to give up.** Counting failed starts and rolling back +lives in the launcher, where it can be tested — `OnFailure=` in a unit file can only be read and +hoped for, and it is the one thing that has to work on a machine where nothing else does. **The package owns this file. The host never does.** It manages `service` resources, and its own unit is a service — the temptation is obvious and it ends with a host stopping itself half way diff --git a/03-DESIGN/01-to-be/09-the-node-lifecycle.md b/03-DESIGN/01-to-be/09-the-node-lifecycle.md index 1fc32ab..e5d1fbe 100644 --- a/03-DESIGN/01-to-be/09-the-node-lifecycle.md +++ b/03-DESIGN/01-to-be/09-the-node-lifecycle.md @@ -12,7 +12,7 @@ decisions: - 02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md - 02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md - 02-DECISIONS/0058-delivery-ends-in-a-declaration.md - - 02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md + - 02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md --- # The node lifecycle @@ -233,8 +233,8 @@ So the two periodic things do different jobs and should not be conflated: without a heartbeat that is indistinguishable from a node that stopped. With one, *last heard from* is a fact beside every node — which is what [how long disconnected](#how-long-disconnected-and-who-is-told) reports and what -[ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md) exists because -a stuck node cannot send. +[ADR 0061](../../02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md) exists +because a stuck node cannot send. **Rebooting mid-apply is safe by construction.** The store records each resource *after* it worked ([ADR 0035](../../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md)), so a host that @@ -394,11 +394,29 @@ own apply completes. A node must therefore report the version it is **running**, installed — otherwise the mesh believes an upgrade landed at step 1. **A version that crashes on start rolls itself back** -([ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md)). The service -manager gives up after three failures in two minutes and runs a rollback script — shipped by the -package rather than being a host subcommand, because a binary that will not start cannot be its -own recovery. It reinstalls the version recorded in `known-good`, which the host wrote the last -time it completed a reconcile. +([ADR 0061](../../02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md)). + +What the init starts is not the host but a **launcher**, and the launcher is where the policy +lives: + +``` +init ──► nox-mesh-host-launch ──► nox-mesh-host + ├─ halted? say so and stop; a person has to look + ├─ count this start attempt + ├─ too many, not yet rolled back? roll back, then start + ├─ too many, already rolled back? halt — the machine is the problem + └─ otherwise start the host +``` + +It reinstalls the version recorded in `known-good`, which the host wrote the last time it +completed a reconcile — and the host clears the attempt counter at the same moment, for the same +reason. + +**The launcher rather than the init's own features**, because this is the one thing that must +work on a machine where nothing else does. A shell script with a counter can be run against a +stub package manager and asserted; `OnFailure=` in a unit file can only be read and hoped for. +It also means the init is asked for nothing but *start* and *restart*, which every init can +do. **It rolls back once.** If the previous version also fails, the node stops in a failed state rather than flapping between two binaries. A second failure is a different diagnosis: the @@ -520,10 +538,10 @@ the same command against a mesh that is one machine old. ## Still open - ~~**Automatic rollback of a bad host version.**~~ **Resolved** by - [ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md): the service - manager gives up after three failures and runs a rollback script — shipped by the package, not - the host binary, because a binary that will not start cannot recover itself. It rolls back - once; a second failure means the machine is the problem, not the binary. + [ADR 0061](../../02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md): a launcher + counts failed starts and rolls back — shipped by the package, not the host binary, because a + binary that will not start cannot recover itself. It rolls back once; a second failure means + the machine is the problem, not the binary. - **How a previous declaration is retained and chosen**, which is what rollback of anything else would use ([ADR 0058](../../02-DECISIONS/0058-delivery-ends-in-a-declaration.md)). - **A node returning after months** applies a very large jump in one go. Correct, and untested.