diff --git a/00-META/checks/records.py b/00-META/checks/records.py index 09bfd7e..4730d40 100644 --- a/00-META/checks/records.py +++ b/00-META/checks/records.py @@ -153,6 +153,12 @@ def check_rests_on(failures, records): continue status = records[number]["front"].get("status") if status != "accepted": + # A proposed record may extend another proposed one. Decisions are drafted in + # chains -- 0059 extends 0057 while both await review -- and refusing that would + # mean either drafting out of order or marking records accepted to satisfy a + # check, which is the failure this repository already made once. + if frontmatter(read(path)).get("status") == "proposed": + continue # An as-is document describes what runs, and what runs was built under # whatever was decided at the time. ADR 0056: "as-is describing a superseded # decision is exactly what as-is is for." diff --git a/02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md b/02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md index ee676e1..e42ab8c 100644 --- a/02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md +++ b/02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md @@ -100,24 +100,43 @@ Three conditions, and they are the whole safety argument: This makes a fleet-wide host upgrade an ordinary declaration, which the first draft gave up. -### It reconciles on four triggers +### Changes are pushed. The timer is for drift, and only for drift -| Trigger | Why | -|---|---| -| **start** | the machine may have changed while nothing was running | -| **a declaration arrives** | the ordinary path | -| **a timer** | drift. Something other than the host changed the machine — a person, a package upgrade replacing a config file | -| **reconnect** | it may have missed declarations while disconnected ([ADR 0036](0036-a-node-is-a-managed-machine.md)) | +**The host does not poll for work.** A new declaration arrives as a message on the link, and the +host applies it then ([ADR 0001](0001-nodes-communicate-over-a-broker.md)). Polling for updates +over a connection that already exists would be strictly worse in both directions: slower to +land, and constant traffic to learn nothing. -**The timer is what makes the store's claim true.** Without it a machine that drifted stays -drifted until somebody changes a declaration, and `owned` reports what the host *applied* rather -than what is *there* — which is +Four triggers, and only one of them is a clock: + +| Trigger | Kind | Why | +|---|---|---| +| **a declaration arrives** | **pushed** | the ordinary path — this is how changes land | +| **start** | event | the machine may have changed while nothing was running | +| **reconnect** | event | declarations may have been missed while disconnected | +| **a timer** | periodic | **drift**, and nothing else | + +**The timer cannot be replaced by an event, and the reason is definitional.** Drift is change the +*mesh did not make* — somebody edited a managed file, a distribution upgrade replaced a config, +a container was stopped by hand. **Nothing will ever send a message about it**, because the thing +that did it is not part of the mesh. A local periodic check is the only way to see it at all. + +Without it, `owned` reports what the host *applied* rather than what is *there*, which is [ADR 0035](0035-a-picture-is-read-from-what-runs.md) violated by omission. -**Ten minutes**, configurable. Short enough that drift is bounded by something a person would -notice anyway, long enough that a fleet is not doing constant work. The reconcile is cheap: it -asks the package database, the service manager and the container runtime about resources the -host already knows it owns. +**Ten minutes**, configurable. The check is cheap: it asks the package database, the service +manager and the container runtime about resources the host already knows it owns. + +### It reports upward on a heartbeat + +Separate from reconciling, and easy to conflate with it: the node tells the mesh what it is — +its running version, what it holds, what it last applied — on link, after every apply, and +periodically while idle. + +**The heartbeat is what makes silence mean something.** Without it, the mesh cannot distinguish +a node that is fine and has had nothing to do from one that stopped. With it, *last heard from* +is a fact per node, and [ADR 0059](0059-a-host-that-cannot-start-rolls-itself-back.md) is what +handles the case where the node cannot report at all. ## Consequences diff --git a/02-DECISIONS/0058-delivery-ends-in-a-declaration.md b/02-DECISIONS/0058-delivery-ends-in-a-declaration.md index 2e89c38..75ad3e7 100644 --- a/02-DECISIONS/0058-delivery-ends-in-a-declaration.md +++ b/02-DECISIONS/0058-delivery-ends-in-a-declaration.md @@ -90,8 +90,13 @@ The host is tier 0, which makes it tempting to treat as special. It is not: 2. publish packages it and puts it in **the mesh's own package repository** — which is a directory of files behind the object store and the proxy, so it needs no new machinery; 3. deploy updates each node's declaration to name the new version; -4. each host applies `package: nox-mesh-host` on its next reconcile, exactly as it applies any - other package. +4. **the new declaration is pushed to each node**, and the host applies + `package: nox-mesh-host` on arrival, exactly as it applies any other package. + +Step 4 is a push and not a poll. The link is already open +([ADR 0001](0001-nodes-communicate-over-a-broker.md)), so a node learns of a change when it is +made — a node that is offline learns on reconnect, which is what makes *outstanding* a real +category rather than a euphemism for lost. **No new resource type is needed**, which is the test of whether this is uniform or a special case wearing a uniform. The repository is reachable because a `file` resource put its address in diff --git a/02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md b/02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md new file mode 100644 index 0000000..d92a9c8 --- /dev/null +++ b/02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md @@ -0,0 +1,133 @@ +--- +status: proposed +date: 2026-08-27 +deciders: jochen +reconstructed: false +extends: 0057-the-host-is-a-root-service-installed-as-a-package.md +--- + +# 59. A host that cannot start is rolled back by the supervisor + +## Context + +[`09-the-node-lifecycle.md`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md) left one thing +unsolved: a host upgraded to a version that crashes on start. The supervisor restarts it, backs +off, and the node is stuck on a binary that will not run. + +It was left because automatic recovery looked like *the host judging its own health*, which is +the self-reference the rest of that document avoids. + +**That objection does not survive being asked properly.** A keepalive is not the host judging +itself — it is something else judging the host. The question is only *what*, and that has one +answer. + +### Why the watchdog must be local + +The obvious candidate is the control plane, and it cannot be: + +- **Nothing dials a node.** [ADR 0039](0039-the-link-is-the-security-boundary.md) makes the link + outbound and node-initiated, and a node has no listening control surface. The mesh has no way + to reach in and act. +- **The failure removes the reporting path.** A host that cannot start cannot link, so the mesh + learns nothing to act on. + +So the watchdog is local, and the only local thing that is always present, already depended on, +and is not the host is the **service manager**. + +### The failure this actually prevents + +Sharper than "the node is down", and it is the reason to bother: + +> **A host that will not start looks exactly like a machine that was switched off.** + +[ADR 0036](0036-a-node-is-a-managed-machine.md) makes disconnection ordinary, and +[`09`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md) deliberately puts no alarm on it — a +laptop shut for three weeks is doing nothing wrong. **A bad upgrade is therefore invisible**: it +presents as the one condition the design has decided not to be alarmed by. + +Worse, it arrives fleet-wide. A declaration naming a bad version reaches every node, and each +one applies it, restarts, and stops talking. The mesh would report a fleet of quiet nodes and +nothing else. + +## Decision + +**The service manager rolls the host back to the last version that started.** + +Four parts, and each one is chosen so it works when the host does not: + +### 1 — A version is confirmed by starting, not by seeming well + +On start, the host completes one full reconcile. If it does, it writes the running version to a +plain file — `/var/lib/mesh-host/known-good` — and that is the whole of "confirmed". + +**Deliberately not health.** Not *the link is up*, because a disconnected node is ordinary and a +laptop on a train would roll itself back. Not *everything applied cleanly*, because a resource +that fails is the machine's problem and not the binary's. The claim is narrow and checkable: **it +started, and it got through a reconcile.** + +### 2 — The rollback is not the host binary + +The obvious mistake, and it would make the whole mechanism a no-op: `nox-mesh-host rollback` +cannot be the recovery path for a `nox-mesh-host` that does not run. + +Rollback is a **small script shipped by the package**, which reads the `known-good` file and +asks the package manager to install that version. It shares no code with the host and does not +import it. + +### 3 — The supervisor triggers it, after giving up + +```ini +[Service] +Restart=on-failure +StartLimitBurst=3 +StartLimitIntervalSec=120 + +[Unit] +OnFailure=nox-mesh-host-rollback.service +``` + +Three failures in two minutes is a binary that does not work, not a transient. The supervisor +stops trying and runs the rollback unit, which downgrades and starts the host again. + +### 4 — It rolls back once + +The rollback unit records that it fired. If the rolled-back version *also* fails to start, it +does **not** fire again — the node stops, loudly, in a failed state. + +**Because a second rollback is a different diagnosis.** Once the previously-working binary also +fails, the binary is not the problem: the machine is. Rolling back further would flap between +two versions forever and bury the actual cause under a loop. + +## Consequences + +- **The node keeps its own recovery**, which is the property the whole tier-0 design rests on: + the host depends on nothing, and now its recovery depends on nothing either. +- **The package cache must retain the previous version**, and that is a real requirement rather + than an assumption — a package manager configured to clean its cache would delete the thing + rollback needs. The package must pin that, and it is the sort of dependency that is discovered + by the rollback failing. +- **A fleet-wide bad upgrade becomes self-limiting.** Each node fails, rolls back, and comes + back on the previous version — so the blast radius of a bad host release is one restart cycle + per node rather than the entire fleet stopping. +- **It catches "will not start" and nothing else.** A version that starts and is subtly wrong + will not roll back, and should not: that is a bad release, which is delivery's problem + ([ADR 0058](0058-delivery-ends-in-a-declaration.md)), not a supervision problem. +- **It adds a unit and a script the host does not own**, which is + [ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)'s line holding: the + installation owns the host, and now it owns the host's recovery too. Consistent, and it means + both arrive and are versioned together. +- **A node in permanent failure is silent, and that is now the last gap.** After a second + failure the node is stopped and cannot report it. What notices is the mesh seeing a node that + has not been heard from — which `09` records as a fact with no threshold, and which this makes + more important to look at than it was. +- **The rollback path is exercised only when it is needed**, which is when nobody can afford it + to be wrong. It wants a test that boots a lab node onto a deliberately broken host and asserts + the previous version comes back. + +## References + +- [ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md) — the restart-by-exiting + this protects. +- [ADR 0039](0039-the-link-is-the-security-boundary.md) — why the watchdog cannot be remote. +- [ADR 0036](0036-a-node-is-a-managed-machine.md) — why the failure is invisible without this. +- [`09-the-node-lifecycle.md`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md) — the gap this closes. diff --git a/03-DESIGN/01-to-be/09-the-node-lifecycle.md b/03-DESIGN/01-to-be/09-the-node-lifecycle.md index 1bd3cc7..f8eaaac 100644 --- a/03-DESIGN/01-to-be/09-the-node-lifecycle.md +++ b/03-DESIGN/01-to-be/09-the-node-lifecycle.md @@ -158,23 +158,43 @@ outcome **derived** from the worst line rather than stated alongside it. ## enrolled: what running actually looks like -Four reconcile triggers -([ADR 0057](../../02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md)): +**Changes are pushed, not polled.** A declaration arrives as a message on the link and the host +applies it then. The link is already open and outbound +([ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md), +[ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md)) — asking it +repeatedly whether anything has changed would be slower to land *and* constant traffic to learn +nothing. -| | | -|---|---| -| **start** | the machine may have changed while nothing was running | -| **a declaration arrives** | the ordinary path | -| **every ten minutes** | drift — something other than the host changed the machine | -| **reconnect** | declarations may have been missed | +| Trigger | Kind | | +|---|---|---| +| **a declaration arrives** | **pushed** | the ordinary path — this is how a change lands | +| **start** | event | the machine may have changed while nothing was running | +| **reconnect** | event | declarations may have been missed | +| **every ten minutes** | periodic | **drift, and only drift** | -**Rebooting mid-apply is safe, and it is safe by construction.** The store records each resource -*after* it worked ([ADR 0035](../../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md)), so -a host that dies half way through comes back, finds the completed ones already matching, and -applies the rest. The rule that exists to stop the host lying about what it did also makes it -crash-safe. +**Why the timer cannot be an event.** Drift is change the *mesh did not make* — somebody edited +a managed file, a distribution upgrade replaced a config, a container was stopped by hand. +Nothing will ever publish a message about it, because whatever did it is not part of the mesh. +Only looking finds it. ---- +So the two periodic things do different jobs and should not be conflated: + +| | direction | answers | +|---|---|---| +| **reconcile timer** | local, looks at the machine | *does this machine still match what it was told?* | +| **heartbeat** | upward, reports to the mesh | *is this node still here, and what is it running?* | + +**The heartbeat is what makes silence mean something.** A node with nothing to do sends nothing; +without a heartbeat that is indistinguishable from a node that stopped. With one, *last heard +from* is a fact beside every node — which is what +[how long disconnected](#how-long-disconnected-and-who-is-told) reports and what +[ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md) exists because +a stuck node cannot send. + +**Rebooting mid-apply is safe by construction.** The store records each resource *after* it +worked ([ADR 0035](../../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md)), so a host that +dies half way through comes back, finds the completed ones already matching, and applies the +rest. The rule that exists to stop the host lying about what it did also makes it crash-safe. ## Updating what the node holds @@ -290,7 +310,8 @@ push to mesh-host ├─ publish packaged, into the mesh's own package repository └─ deploy each node's declaration now names the new version │ - └─ every host applies it on its next reconcile + └─ pushed to each node; the host applies it on arrival + (a node that is offline gets it on reconnect) ``` **Compared with today.** The current pipeline's third silo runs *once per node* and sends each @@ -327,11 +348,21 @@ part-way through an apply. It stops by finishing. own apply completes. A node must therefore report the version it is **running**, not the one installed — otherwise the mesh believes an upgrade landed at step 1. -**What is not solved: a new version that crashes on start.** systemd will restart it, back off, -and the node is stuck on a binary that does not run. The previous package is still in the -package manager's cache, so a person at the machine can downgrade — but there is no automatic -rollback, and designing one means the host judging its own health, which is the kind of -self-reference the rest of this document avoids. **Named, not solved.** +**A version that crashes on start rolls itself back** +([ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md)). The service +manager gives up after three failures in two minutes and runs a rollback script — shipped by the +package rather than being a host subcommand, because a binary that will not start cannot be its +own recovery. It reinstalls the version recorded in `known-good`, which the host wrote the last +time it completed a reconcile. + +**It rolls back once.** If the previous version also fails, the node stops in a failed state +rather than flapping between two binaries. A second failure is a different diagnosis: the +machine is the problem, not the binary. + +**Why this matters more than it looks.** A host that will not start cannot link, and a node that +is not linking looks exactly like a machine somebody switched off — which is the one condition +this design has deliberately decided not to alarm on. Without rollback, a bad release reaches +every node, each one goes quiet, and the mesh reports a fleet of sleeping laptops. --- @@ -443,8 +474,11 @@ the same command against a mesh that is one machine old. ## Still open -- **Automatic rollback of a bad host version.** Above. A node can be left on a binary that will - not start, and recovering needs a person at the machine. +- ~~**Automatic rollback of a bad host version.**~~ **Resolved** by + [ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md): the service + manager gives up after three failures and runs a rollback script — shipped by the package, not + the host binary, because a binary that will not start cannot recover itself. It rolls back + once; a second failure means the machine is the problem, not the binary. - **How a previous declaration is retained and chosen**, which is what rollback of anything else would use ([ADR 0058](../../02-DECISIONS/0058-delivery-ends-in-a-declaration.md)). - **A node returning after months** applies a very large jump in one go. Correct, and untested.