Changes are pushed, not polled; and a stuck host rolls itself back
Two corrections and one new decision, all from Jochen catching things. Pushed, not polled. I described updates as landing "on the next reconcile", which reads as polling and is not the design. A declaration arrives as a message on a link that is already open; the host applies it then. Polling over an existing connection would be slower to land AND constant traffic to learn nothing. The timer is for drift and nothing else, and it cannot be replaced by an event for a definitional reason: drift is change the mesh did not make -- somebody edited a managed file, a distribution upgrade replaced a config -- so nothing will ever publish a message about it. Only looking finds it. Separated the heartbeat from the reconcile timer, which I had been conflating. They point in opposite directions and answer different questions: the timer looks at the machine and asks whether it still matches; the heartbeat reports upward and is what makes silence mean something. A node with nothing to do sends nothing, and without a heartbeat that is indistinguishable from a node that stopped. 0059 -- a host that cannot start is rolled back by the service manager. I had left this open on the grounds that recovery meant the host judging its own health. That objection does not survive being asked properly: a keepalive is something else judging the host. The watchdog must be local, because nothing dials a node and a host that cannot start cannot report -- so it is the service manager, which is already there. The failure it prevents is sharper than "the node is down": a host that will not start looks exactly like a machine somebody switched off, which is the one condition this design has deliberately decided not to alarm on. So a bad release reaches every node, each goes quiet, and the mesh reports a fleet of sleeping laptops. Confirmed means started and completed one reconcile -- deliberately not "the link is up", or a laptop on a train would roll itself back. The rollback is a script shipped by the package, not a host subcommand, because a binary that will not start cannot be its own recovery. It rolls back once: a second failure means the machine is the problem, not the binary. Also refined the records checker, which produced a false positive: a proposed record may extend another proposed one, because decisions are drafted in chains and the alternative is marking things accepted to satisfy a check. An accepted document resting on a proposed record still fails, and that was verified. 0057, 0058 and 0059 are all proposed.
This commit is contained in:
@@ -153,6 +153,12 @@ def check_rests_on(failures, records):
|
|||||||
continue
|
continue
|
||||||
status = records[number]["front"].get("status")
|
status = records[number]["front"].get("status")
|
||||||
if status != "accepted":
|
if status != "accepted":
|
||||||
|
# A proposed record may extend another proposed one. Decisions are drafted in
|
||||||
|
# chains -- 0059 extends 0057 while both await review -- and refusing that would
|
||||||
|
# mean either drafting out of order or marking records accepted to satisfy a
|
||||||
|
# check, which is the failure this repository already made once.
|
||||||
|
if frontmatter(read(path)).get("status") == "proposed":
|
||||||
|
continue
|
||||||
# An as-is document describes what runs, and what runs was built under
|
# An as-is document describes what runs, and what runs was built under
|
||||||
# whatever was decided at the time. ADR 0056: "as-is describing a superseded
|
# whatever was decided at the time. ADR 0056: "as-is describing a superseded
|
||||||
# decision is exactly what as-is is for."
|
# decision is exactly what as-is is for."
|
||||||
|
|||||||
@@ -100,24 +100,43 @@ Three conditions, and they are the whole safety argument:
|
|||||||
|
|
||||||
This makes a fleet-wide host upgrade an ordinary declaration, which the first draft gave up.
|
This makes a fleet-wide host upgrade an ordinary declaration, which the first draft gave up.
|
||||||
|
|
||||||
### It reconciles on four triggers
|
### Changes are pushed. The timer is for drift, and only for drift
|
||||||
|
|
||||||
| Trigger | Why |
|
**The host does not poll for work.** A new declaration arrives as a message on the link, and the
|
||||||
|---|---|
|
host applies it then ([ADR 0001](0001-nodes-communicate-over-a-broker.md)). Polling for updates
|
||||||
| **start** | the machine may have changed while nothing was running |
|
over a connection that already exists would be strictly worse in both directions: slower to
|
||||||
| **a declaration arrives** | the ordinary path |
|
land, and constant traffic to learn nothing.
|
||||||
| **a timer** | drift. Something other than the host changed the machine — a person, a package upgrade replacing a config file |
|
|
||||||
| **reconnect** | it may have missed declarations while disconnected ([ADR 0036](0036-a-node-is-a-managed-machine.md)) |
|
|
||||||
|
|
||||||
**The timer is what makes the store's claim true.** Without it a machine that drifted stays
|
Four triggers, and only one of them is a clock:
|
||||||
drifted until somebody changes a declaration, and `owned` reports what the host *applied* rather
|
|
||||||
than what is *there* — which is
|
| Trigger | Kind | Why |
|
||||||
|
|---|---|---|
|
||||||
|
| **a declaration arrives** | **pushed** | the ordinary path — this is how changes land |
|
||||||
|
| **start** | event | the machine may have changed while nothing was running |
|
||||||
|
| **reconnect** | event | declarations may have been missed while disconnected |
|
||||||
|
| **a timer** | periodic | **drift**, and nothing else |
|
||||||
|
|
||||||
|
**The timer cannot be replaced by an event, and the reason is definitional.** Drift is change the
|
||||||
|
*mesh did not make* — somebody edited a managed file, a distribution upgrade replaced a config,
|
||||||
|
a container was stopped by hand. **Nothing will ever send a message about it**, because the thing
|
||||||
|
that did it is not part of the mesh. A local periodic check is the only way to see it at all.
|
||||||
|
|
||||||
|
Without it, `owned` reports what the host *applied* rather than what is *there*, which is
|
||||||
[ADR 0035](0035-a-picture-is-read-from-what-runs.md) violated by omission.
|
[ADR 0035](0035-a-picture-is-read-from-what-runs.md) violated by omission.
|
||||||
|
|
||||||
**Ten minutes**, configurable. Short enough that drift is bounded by something a person would
|
**Ten minutes**, configurable. The check is cheap: it asks the package database, the service
|
||||||
notice anyway, long enough that a fleet is not doing constant work. The reconcile is cheap: it
|
manager and the container runtime about resources the host already knows it owns.
|
||||||
asks the package database, the service manager and the container runtime about resources the
|
|
||||||
host already knows it owns.
|
### It reports upward on a heartbeat
|
||||||
|
|
||||||
|
Separate from reconciling, and easy to conflate with it: the node tells the mesh what it is —
|
||||||
|
its running version, what it holds, what it last applied — on link, after every apply, and
|
||||||
|
periodically while idle.
|
||||||
|
|
||||||
|
**The heartbeat is what makes silence mean something.** Without it, the mesh cannot distinguish
|
||||||
|
a node that is fine and has had nothing to do from one that stopped. With it, *last heard from*
|
||||||
|
is a fact per node, and [ADR 0059](0059-a-host-that-cannot-start-rolls-itself-back.md) is what
|
||||||
|
handles the case where the node cannot report at all.
|
||||||
|
|
||||||
## Consequences
|
## Consequences
|
||||||
|
|
||||||
|
|||||||
@@ -90,8 +90,13 @@ The host is tier 0, which makes it tempting to treat as special. It is not:
|
|||||||
2. publish packages it and puts it in **the mesh's own package repository** — which is a
|
2. publish packages it and puts it in **the mesh's own package repository** — which is a
|
||||||
directory of files behind the object store and the proxy, so it needs no new machinery;
|
directory of files behind the object store and the proxy, so it needs no new machinery;
|
||||||
3. deploy updates each node's declaration to name the new version;
|
3. deploy updates each node's declaration to name the new version;
|
||||||
4. each host applies `package: nox-mesh-host` on its next reconcile, exactly as it applies any
|
4. **the new declaration is pushed to each node**, and the host applies
|
||||||
other package.
|
`package: nox-mesh-host` on arrival, exactly as it applies any other package.
|
||||||
|
|
||||||
|
Step 4 is a push and not a poll. The link is already open
|
||||||
|
([ADR 0001](0001-nodes-communicate-over-a-broker.md)), so a node learns of a change when it is
|
||||||
|
made — a node that is offline learns on reconnect, which is what makes *outstanding* a real
|
||||||
|
category rather than a euphemism for lost.
|
||||||
|
|
||||||
**No new resource type is needed**, which is the test of whether this is uniform or a special
|
**No new resource type is needed**, which is the test of whether this is uniform or a special
|
||||||
case wearing a uniform. The repository is reachable because a `file` resource put its address in
|
case wearing a uniform. The repository is reachable because a `file` resource put its address in
|
||||||
|
|||||||
@@ -0,0 +1,133 @@
|
|||||||
|
---
|
||||||
|
status: proposed
|
||||||
|
date: 2026-08-27
|
||||||
|
deciders: jochen
|
||||||
|
reconstructed: false
|
||||||
|
extends: 0057-the-host-is-a-root-service-installed-as-a-package.md
|
||||||
|
---
|
||||||
|
|
||||||
|
# 59. A host that cannot start is rolled back by the supervisor
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
[`09-the-node-lifecycle.md`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md) left one thing
|
||||||
|
unsolved: a host upgraded to a version that crashes on start. The supervisor restarts it, backs
|
||||||
|
off, and the node is stuck on a binary that will not run.
|
||||||
|
|
||||||
|
It was left because automatic recovery looked like *the host judging its own health*, which is
|
||||||
|
the self-reference the rest of that document avoids.
|
||||||
|
|
||||||
|
**That objection does not survive being asked properly.** A keepalive is not the host judging
|
||||||
|
itself — it is something else judging the host. The question is only *what*, and that has one
|
||||||
|
answer.
|
||||||
|
|
||||||
|
### Why the watchdog must be local
|
||||||
|
|
||||||
|
The obvious candidate is the control plane, and it cannot be:
|
||||||
|
|
||||||
|
- **Nothing dials a node.** [ADR 0039](0039-the-link-is-the-security-boundary.md) makes the link
|
||||||
|
outbound and node-initiated, and a node has no listening control surface. The mesh has no way
|
||||||
|
to reach in and act.
|
||||||
|
- **The failure removes the reporting path.** A host that cannot start cannot link, so the mesh
|
||||||
|
learns nothing to act on.
|
||||||
|
|
||||||
|
So the watchdog is local, and the only local thing that is always present, already depended on,
|
||||||
|
and is not the host is the **service manager**.
|
||||||
|
|
||||||
|
### The failure this actually prevents
|
||||||
|
|
||||||
|
Sharper than "the node is down", and it is the reason to bother:
|
||||||
|
|
||||||
|
> **A host that will not start looks exactly like a machine that was switched off.**
|
||||||
|
|
||||||
|
[ADR 0036](0036-a-node-is-a-managed-machine.md) makes disconnection ordinary, and
|
||||||
|
[`09`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md) deliberately puts no alarm on it — a
|
||||||
|
laptop shut for three weeks is doing nothing wrong. **A bad upgrade is therefore invisible**: it
|
||||||
|
presents as the one condition the design has decided not to be alarmed by.
|
||||||
|
|
||||||
|
Worse, it arrives fleet-wide. A declaration naming a bad version reaches every node, and each
|
||||||
|
one applies it, restarts, and stops talking. The mesh would report a fleet of quiet nodes and
|
||||||
|
nothing else.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
**The service manager rolls the host back to the last version that started.**
|
||||||
|
|
||||||
|
Four parts, and each one is chosen so it works when the host does not:
|
||||||
|
|
||||||
|
### 1 — A version is confirmed by starting, not by seeming well
|
||||||
|
|
||||||
|
On start, the host completes one full reconcile. If it does, it writes the running version to a
|
||||||
|
plain file — `/var/lib/mesh-host/known-good` — and that is the whole of "confirmed".
|
||||||
|
|
||||||
|
**Deliberately not health.** Not *the link is up*, because a disconnected node is ordinary and a
|
||||||
|
laptop on a train would roll itself back. Not *everything applied cleanly*, because a resource
|
||||||
|
that fails is the machine's problem and not the binary's. The claim is narrow and checkable: **it
|
||||||
|
started, and it got through a reconcile.**
|
||||||
|
|
||||||
|
### 2 — The rollback is not the host binary
|
||||||
|
|
||||||
|
The obvious mistake, and it would make the whole mechanism a no-op: `nox-mesh-host rollback`
|
||||||
|
cannot be the recovery path for a `nox-mesh-host` that does not run.
|
||||||
|
|
||||||
|
Rollback is a **small script shipped by the package**, which reads the `known-good` file and
|
||||||
|
asks the package manager to install that version. It shares no code with the host and does not
|
||||||
|
import it.
|
||||||
|
|
||||||
|
### 3 — The supervisor triggers it, after giving up
|
||||||
|
|
||||||
|
```ini
|
||||||
|
[Service]
|
||||||
|
Restart=on-failure
|
||||||
|
StartLimitBurst=3
|
||||||
|
StartLimitIntervalSec=120
|
||||||
|
|
||||||
|
[Unit]
|
||||||
|
OnFailure=nox-mesh-host-rollback.service
|
||||||
|
```
|
||||||
|
|
||||||
|
Three failures in two minutes is a binary that does not work, not a transient. The supervisor
|
||||||
|
stops trying and runs the rollback unit, which downgrades and starts the host again.
|
||||||
|
|
||||||
|
### 4 — It rolls back once
|
||||||
|
|
||||||
|
The rollback unit records that it fired. If the rolled-back version *also* fails to start, it
|
||||||
|
does **not** fire again — the node stops, loudly, in a failed state.
|
||||||
|
|
||||||
|
**Because a second rollback is a different diagnosis.** Once the previously-working binary also
|
||||||
|
fails, the binary is not the problem: the machine is. Rolling back further would flap between
|
||||||
|
two versions forever and bury the actual cause under a loop.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
- **The node keeps its own recovery**, which is the property the whole tier-0 design rests on:
|
||||||
|
the host depends on nothing, and now its recovery depends on nothing either.
|
||||||
|
- **The package cache must retain the previous version**, and that is a real requirement rather
|
||||||
|
than an assumption — a package manager configured to clean its cache would delete the thing
|
||||||
|
rollback needs. The package must pin that, and it is the sort of dependency that is discovered
|
||||||
|
by the rollback failing.
|
||||||
|
- **A fleet-wide bad upgrade becomes self-limiting.** Each node fails, rolls back, and comes
|
||||||
|
back on the previous version — so the blast radius of a bad host release is one restart cycle
|
||||||
|
per node rather than the entire fleet stopping.
|
||||||
|
- **It catches "will not start" and nothing else.** A version that starts and is subtly wrong
|
||||||
|
will not roll back, and should not: that is a bad release, which is delivery's problem
|
||||||
|
([ADR 0058](0058-delivery-ends-in-a-declaration.md)), not a supervision problem.
|
||||||
|
- **It adds a unit and a script the host does not own**, which is
|
||||||
|
[ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)'s line holding: the
|
||||||
|
installation owns the host, and now it owns the host's recovery too. Consistent, and it means
|
||||||
|
both arrive and are versioned together.
|
||||||
|
- **A node in permanent failure is silent, and that is now the last gap.** After a second
|
||||||
|
failure the node is stopped and cannot report it. What notices is the mesh seeing a node that
|
||||||
|
has not been heard from — which `09` records as a fact with no threshold, and which this makes
|
||||||
|
more important to look at than it was.
|
||||||
|
- **The rollback path is exercised only when it is needed**, which is when nobody can afford it
|
||||||
|
to be wrong. It wants a test that boots a lab node onto a deliberately broken host and asserts
|
||||||
|
the previous version comes back.
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- [ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md) — the restart-by-exiting
|
||||||
|
this protects.
|
||||||
|
- [ADR 0039](0039-the-link-is-the-security-boundary.md) — why the watchdog cannot be remote.
|
||||||
|
- [ADR 0036](0036-a-node-is-a-managed-machine.md) — why the failure is invisible without this.
|
||||||
|
- [`09-the-node-lifecycle.md`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md) — the gap this closes.
|
||||||
@@ -158,23 +158,43 @@ outcome **derived** from the worst line rather than stated alongside it.
|
|||||||
|
|
||||||
## enrolled: what running actually looks like
|
## enrolled: what running actually looks like
|
||||||
|
|
||||||
Four reconcile triggers
|
**Changes are pushed, not polled.** A declaration arrives as a message on the link and the host
|
||||||
([ADR 0057](../../02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md)):
|
applies it then. The link is already open and outbound
|
||||||
|
([ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md),
|
||||||
|
[ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md)) — asking it
|
||||||
|
repeatedly whether anything has changed would be slower to land *and* constant traffic to learn
|
||||||
|
nothing.
|
||||||
|
|
||||||
| | |
|
| Trigger | Kind | |
|
||||||
|---|---|
|
|---|---|---|
|
||||||
| **start** | the machine may have changed while nothing was running |
|
| **a declaration arrives** | **pushed** | the ordinary path — this is how a change lands |
|
||||||
| **a declaration arrives** | the ordinary path |
|
| **start** | event | the machine may have changed while nothing was running |
|
||||||
| **every ten minutes** | drift — something other than the host changed the machine |
|
| **reconnect** | event | declarations may have been missed |
|
||||||
| **reconnect** | declarations may have been missed |
|
| **every ten minutes** | periodic | **drift, and only drift** |
|
||||||
|
|
||||||
**Rebooting mid-apply is safe, and it is safe by construction.** The store records each resource
|
**Why the timer cannot be an event.** Drift is change the *mesh did not make* — somebody edited
|
||||||
*after* it worked ([ADR 0035](../../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md)), so
|
a managed file, a distribution upgrade replaced a config, a container was stopped by hand.
|
||||||
a host that dies half way through comes back, finds the completed ones already matching, and
|
Nothing will ever publish a message about it, because whatever did it is not part of the mesh.
|
||||||
applies the rest. The rule that exists to stop the host lying about what it did also makes it
|
Only looking finds it.
|
||||||
crash-safe.
|
|
||||||
|
|
||||||
---
|
So the two periodic things do different jobs and should not be conflated:
|
||||||
|
|
||||||
|
| | direction | answers |
|
||||||
|
|---|---|---|
|
||||||
|
| **reconcile timer** | local, looks at the machine | *does this machine still match what it was told?* |
|
||||||
|
| **heartbeat** | upward, reports to the mesh | *is this node still here, and what is it running?* |
|
||||||
|
|
||||||
|
**The heartbeat is what makes silence mean something.** A node with nothing to do sends nothing;
|
||||||
|
without a heartbeat that is indistinguishable from a node that stopped. With one, *last heard
|
||||||
|
from* is a fact beside every node — which is what
|
||||||
|
[how long disconnected](#how-long-disconnected-and-who-is-told) reports and what
|
||||||
|
[ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md) exists because
|
||||||
|
a stuck node cannot send.
|
||||||
|
|
||||||
|
**Rebooting mid-apply is safe by construction.** The store records each resource *after* it
|
||||||
|
worked ([ADR 0035](../../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md)), so a host that
|
||||||
|
dies half way through comes back, finds the completed ones already matching, and applies the
|
||||||
|
rest. The rule that exists to stop the host lying about what it did also makes it crash-safe.
|
||||||
|
|
||||||
## Updating what the node holds
|
## Updating what the node holds
|
||||||
|
|
||||||
@@ -290,7 +310,8 @@ push to mesh-host
|
|||||||
├─ publish packaged, into the mesh's own package repository
|
├─ publish packaged, into the mesh's own package repository
|
||||||
└─ deploy each node's declaration now names the new version
|
└─ deploy each node's declaration now names the new version
|
||||||
│
|
│
|
||||||
└─ every host applies it on its next reconcile
|
└─ pushed to each node; the host applies it on arrival
|
||||||
|
(a node that is offline gets it on reconnect)
|
||||||
```
|
```
|
||||||
|
|
||||||
**Compared with today.** The current pipeline's third silo runs *once per node* and sends each
|
**Compared with today.** The current pipeline's third silo runs *once per node* and sends each
|
||||||
@@ -327,11 +348,21 @@ part-way through an apply. It stops by finishing.
|
|||||||
own apply completes. A node must therefore report the version it is **running**, not the one
|
own apply completes. A node must therefore report the version it is **running**, not the one
|
||||||
installed — otherwise the mesh believes an upgrade landed at step 1.
|
installed — otherwise the mesh believes an upgrade landed at step 1.
|
||||||
|
|
||||||
**What is not solved: a new version that crashes on start.** systemd will restart it, back off,
|
**A version that crashes on start rolls itself back**
|
||||||
and the node is stuck on a binary that does not run. The previous package is still in the
|
([ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md)). The service
|
||||||
package manager's cache, so a person at the machine can downgrade — but there is no automatic
|
manager gives up after three failures in two minutes and runs a rollback script — shipped by the
|
||||||
rollback, and designing one means the host judging its own health, which is the kind of
|
package rather than being a host subcommand, because a binary that will not start cannot be its
|
||||||
self-reference the rest of this document avoids. **Named, not solved.**
|
own recovery. It reinstalls the version recorded in `known-good`, which the host wrote the last
|
||||||
|
time it completed a reconcile.
|
||||||
|
|
||||||
|
**It rolls back once.** If the previous version also fails, the node stops in a failed state
|
||||||
|
rather than flapping between two binaries. A second failure is a different diagnosis: the
|
||||||
|
machine is the problem, not the binary.
|
||||||
|
|
||||||
|
**Why this matters more than it looks.** A host that will not start cannot link, and a node that
|
||||||
|
is not linking looks exactly like a machine somebody switched off — which is the one condition
|
||||||
|
this design has deliberately decided not to alarm on. Without rollback, a bad release reaches
|
||||||
|
every node, each one goes quiet, and the mesh reports a fleet of sleeping laptops.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -443,8 +474,11 @@ the same command against a mesh that is one machine old.
|
|||||||
|
|
||||||
## Still open
|
## Still open
|
||||||
|
|
||||||
- **Automatic rollback of a bad host version.** Above. A node can be left on a binary that will
|
- ~~**Automatic rollback of a bad host version.**~~ **Resolved** by
|
||||||
not start, and recovering needs a person at the machine.
|
[ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md): the service
|
||||||
|
manager gives up after three failures and runs a rollback script — shipped by the package, not
|
||||||
|
the host binary, because a binary that will not start cannot recover itself. It rolls back
|
||||||
|
once; a second failure means the machine is the problem, not the binary.
|
||||||
- **How a previous declaration is retained and chosen**, which is what rollback of anything else
|
- **How a previous declaration is retained and chosen**, which is what rollback of anything else
|
||||||
would use ([ADR 0058](../../02-DECISIONS/0058-delivery-ends-in-a-declaration.md)).
|
would use ([ADR 0058](../../02-DECISIONS/0058-delivery-ends-in-a-declaration.md)).
|
||||||
- **A node returning after months** applies a very large jump in one go. Correct, and untested.
|
- **A node returning after months** applies a very large jump in one go. Correct, and untested.
|
||||||
|
|||||||
Reference in New Issue
Block a user