Changes are pushed, not polled; and a stuck host rolls itself back
Two corrections and one new decision, all from Jochen catching things. Pushed, not polled. I described updates as landing "on the next reconcile", which reads as polling and is not the design. A declaration arrives as a message on a link that is already open; the host applies it then. Polling over an existing connection would be slower to land AND constant traffic to learn nothing. The timer is for drift and nothing else, and it cannot be replaced by an event for a definitional reason: drift is change the mesh did not make -- somebody edited a managed file, a distribution upgrade replaced a config -- so nothing will ever publish a message about it. Only looking finds it. Separated the heartbeat from the reconcile timer, which I had been conflating. They point in opposite directions and answer different questions: the timer looks at the machine and asks whether it still matches; the heartbeat reports upward and is what makes silence mean something. A node with nothing to do sends nothing, and without a heartbeat that is indistinguishable from a node that stopped. 0059 -- a host that cannot start is rolled back by the service manager. I had left this open on the grounds that recovery meant the host judging its own health. That objection does not survive being asked properly: a keepalive is something else judging the host. The watchdog must be local, because nothing dials a node and a host that cannot start cannot report -- so it is the service manager, which is already there. The failure it prevents is sharper than "the node is down": a host that will not start looks exactly like a machine somebody switched off, which is the one condition this design has deliberately decided not to alarm on. So a bad release reaches every node, each goes quiet, and the mesh reports a fleet of sleeping laptops. Confirmed means started and completed one reconcile -- deliberately not "the link is up", or a laptop on a train would roll itself back. The rollback is a script shipped by the package, not a host subcommand, because a binary that will not start cannot be its own recovery. It rolls back once: a second failure means the machine is the problem, not the binary. Also refined the records checker, which produced a false positive: a proposed record may extend another proposed one, because decisions are drafted in chains and the alternative is marking things accepted to satisfy a check. An accepted document resting on a proposed record still fails, and that was verified. 0057, 0058 and 0059 are all proposed.
This commit is contained in:
@@ -100,24 +100,43 @@ Three conditions, and they are the whole safety argument:
|
||||
|
||||
This makes a fleet-wide host upgrade an ordinary declaration, which the first draft gave up.
|
||||
|
||||
### It reconciles on four triggers
|
||||
### Changes are pushed. The timer is for drift, and only for drift
|
||||
|
||||
| Trigger | Why |
|
||||
|---|---|
|
||||
| **start** | the machine may have changed while nothing was running |
|
||||
| **a declaration arrives** | the ordinary path |
|
||||
| **a timer** | drift. Something other than the host changed the machine — a person, a package upgrade replacing a config file |
|
||||
| **reconnect** | it may have missed declarations while disconnected ([ADR 0036](0036-a-node-is-a-managed-machine.md)) |
|
||||
**The host does not poll for work.** A new declaration arrives as a message on the link, and the
|
||||
host applies it then ([ADR 0001](0001-nodes-communicate-over-a-broker.md)). Polling for updates
|
||||
over a connection that already exists would be strictly worse in both directions: slower to
|
||||
land, and constant traffic to learn nothing.
|
||||
|
||||
**The timer is what makes the store's claim true.** Without it a machine that drifted stays
|
||||
drifted until somebody changes a declaration, and `owned` reports what the host *applied* rather
|
||||
than what is *there* — which is
|
||||
Four triggers, and only one of them is a clock:
|
||||
|
||||
| Trigger | Kind | Why |
|
||||
|---|---|---|
|
||||
| **a declaration arrives** | **pushed** | the ordinary path — this is how changes land |
|
||||
| **start** | event | the machine may have changed while nothing was running |
|
||||
| **reconnect** | event | declarations may have been missed while disconnected |
|
||||
| **a timer** | periodic | **drift**, and nothing else |
|
||||
|
||||
**The timer cannot be replaced by an event, and the reason is definitional.** Drift is change the
|
||||
*mesh did not make* — somebody edited a managed file, a distribution upgrade replaced a config,
|
||||
a container was stopped by hand. **Nothing will ever send a message about it**, because the thing
|
||||
that did it is not part of the mesh. A local periodic check is the only way to see it at all.
|
||||
|
||||
Without it, `owned` reports what the host *applied* rather than what is *there*, which is
|
||||
[ADR 0035](0035-a-picture-is-read-from-what-runs.md) violated by omission.
|
||||
|
||||
**Ten minutes**, configurable. Short enough that drift is bounded by something a person would
|
||||
notice anyway, long enough that a fleet is not doing constant work. The reconcile is cheap: it
|
||||
asks the package database, the service manager and the container runtime about resources the
|
||||
host already knows it owns.
|
||||
**Ten minutes**, configurable. The check is cheap: it asks the package database, the service
|
||||
manager and the container runtime about resources the host already knows it owns.
|
||||
|
||||
### It reports upward on a heartbeat
|
||||
|
||||
Separate from reconciling, and easy to conflate with it: the node tells the mesh what it is —
|
||||
its running version, what it holds, what it last applied — on link, after every apply, and
|
||||
periodically while idle.
|
||||
|
||||
**The heartbeat is what makes silence mean something.** Without it, the mesh cannot distinguish
|
||||
a node that is fine and has had nothing to do from one that stopped. With it, *last heard from*
|
||||
is a fact per node, and [ADR 0059](0059-a-host-that-cannot-start-rolls-itself-back.md) is what
|
||||
handles the case where the node cannot report at all.
|
||||
|
||||
## Consequences
|
||||
|
||||
|
||||
@@ -90,8 +90,13 @@ The host is tier 0, which makes it tempting to treat as special. It is not:
|
||||
2. publish packages it and puts it in **the mesh's own package repository** — which is a
|
||||
directory of files behind the object store and the proxy, so it needs no new machinery;
|
||||
3. deploy updates each node's declaration to name the new version;
|
||||
4. each host applies `package: nox-mesh-host` on its next reconcile, exactly as it applies any
|
||||
other package.
|
||||
4. **the new declaration is pushed to each node**, and the host applies
|
||||
`package: nox-mesh-host` on arrival, exactly as it applies any other package.
|
||||
|
||||
Step 4 is a push and not a poll. The link is already open
|
||||
([ADR 0001](0001-nodes-communicate-over-a-broker.md)), so a node learns of a change when it is
|
||||
made — a node that is offline learns on reconnect, which is what makes *outstanding* a real
|
||||
category rather than a euphemism for lost.
|
||||
|
||||
**No new resource type is needed**, which is the test of whether this is uniform or a special
|
||||
case wearing a uniform. The repository is reachable because a `file` resource put its address in
|
||||
|
||||
@@ -0,0 +1,133 @@
|
||||
---
|
||||
status: proposed
|
||||
date: 2026-08-27
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 0057-the-host-is-a-root-service-installed-as-a-package.md
|
||||
---
|
||||
|
||||
# 59. A host that cannot start is rolled back by the supervisor
|
||||
|
||||
## Context
|
||||
|
||||
[`09-the-node-lifecycle.md`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md) left one thing
|
||||
unsolved: a host upgraded to a version that crashes on start. The supervisor restarts it, backs
|
||||
off, and the node is stuck on a binary that will not run.
|
||||
|
||||
It was left because automatic recovery looked like *the host judging its own health*, which is
|
||||
the self-reference the rest of that document avoids.
|
||||
|
||||
**That objection does not survive being asked properly.** A keepalive is not the host judging
|
||||
itself — it is something else judging the host. The question is only *what*, and that has one
|
||||
answer.
|
||||
|
||||
### Why the watchdog must be local
|
||||
|
||||
The obvious candidate is the control plane, and it cannot be:
|
||||
|
||||
- **Nothing dials a node.** [ADR 0039](0039-the-link-is-the-security-boundary.md) makes the link
|
||||
outbound and node-initiated, and a node has no listening control surface. The mesh has no way
|
||||
to reach in and act.
|
||||
- **The failure removes the reporting path.** A host that cannot start cannot link, so the mesh
|
||||
learns nothing to act on.
|
||||
|
||||
So the watchdog is local, and the only local thing that is always present, already depended on,
|
||||
and is not the host is the **service manager**.
|
||||
|
||||
### The failure this actually prevents
|
||||
|
||||
Sharper than "the node is down", and it is the reason to bother:
|
||||
|
||||
> **A host that will not start looks exactly like a machine that was switched off.**
|
||||
|
||||
[ADR 0036](0036-a-node-is-a-managed-machine.md) makes disconnection ordinary, and
|
||||
[`09`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md) deliberately puts no alarm on it — a
|
||||
laptop shut for three weeks is doing nothing wrong. **A bad upgrade is therefore invisible**: it
|
||||
presents as the one condition the design has decided not to be alarmed by.
|
||||
|
||||
Worse, it arrives fleet-wide. A declaration naming a bad version reaches every node, and each
|
||||
one applies it, restarts, and stops talking. The mesh would report a fleet of quiet nodes and
|
||||
nothing else.
|
||||
|
||||
## Decision
|
||||
|
||||
**The service manager rolls the host back to the last version that started.**
|
||||
|
||||
Four parts, and each one is chosen so it works when the host does not:
|
||||
|
||||
### 1 — A version is confirmed by starting, not by seeming well
|
||||
|
||||
On start, the host completes one full reconcile. If it does, it writes the running version to a
|
||||
plain file — `/var/lib/mesh-host/known-good` — and that is the whole of "confirmed".
|
||||
|
||||
**Deliberately not health.** Not *the link is up*, because a disconnected node is ordinary and a
|
||||
laptop on a train would roll itself back. Not *everything applied cleanly*, because a resource
|
||||
that fails is the machine's problem and not the binary's. The claim is narrow and checkable: **it
|
||||
started, and it got through a reconcile.**
|
||||
|
||||
### 2 — The rollback is not the host binary
|
||||
|
||||
The obvious mistake, and it would make the whole mechanism a no-op: `nox-mesh-host rollback`
|
||||
cannot be the recovery path for a `nox-mesh-host` that does not run.
|
||||
|
||||
Rollback is a **small script shipped by the package**, which reads the `known-good` file and
|
||||
asks the package manager to install that version. It shares no code with the host and does not
|
||||
import it.
|
||||
|
||||
### 3 — The supervisor triggers it, after giving up
|
||||
|
||||
```ini
|
||||
[Service]
|
||||
Restart=on-failure
|
||||
StartLimitBurst=3
|
||||
StartLimitIntervalSec=120
|
||||
|
||||
[Unit]
|
||||
OnFailure=nox-mesh-host-rollback.service
|
||||
```
|
||||
|
||||
Three failures in two minutes is a binary that does not work, not a transient. The supervisor
|
||||
stops trying and runs the rollback unit, which downgrades and starts the host again.
|
||||
|
||||
### 4 — It rolls back once
|
||||
|
||||
The rollback unit records that it fired. If the rolled-back version *also* fails to start, it
|
||||
does **not** fire again — the node stops, loudly, in a failed state.
|
||||
|
||||
**Because a second rollback is a different diagnosis.** Once the previously-working binary also
|
||||
fails, the binary is not the problem: the machine is. Rolling back further would flap between
|
||||
two versions forever and bury the actual cause under a loop.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **The node keeps its own recovery**, which is the property the whole tier-0 design rests on:
|
||||
the host depends on nothing, and now its recovery depends on nothing either.
|
||||
- **The package cache must retain the previous version**, and that is a real requirement rather
|
||||
than an assumption — a package manager configured to clean its cache would delete the thing
|
||||
rollback needs. The package must pin that, and it is the sort of dependency that is discovered
|
||||
by the rollback failing.
|
||||
- **A fleet-wide bad upgrade becomes self-limiting.** Each node fails, rolls back, and comes
|
||||
back on the previous version — so the blast radius of a bad host release is one restart cycle
|
||||
per node rather than the entire fleet stopping.
|
||||
- **It catches "will not start" and nothing else.** A version that starts and is subtly wrong
|
||||
will not roll back, and should not: that is a bad release, which is delivery's problem
|
||||
([ADR 0058](0058-delivery-ends-in-a-declaration.md)), not a supervision problem.
|
||||
- **It adds a unit and a script the host does not own**, which is
|
||||
[ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)'s line holding: the
|
||||
installation owns the host, and now it owns the host's recovery too. Consistent, and it means
|
||||
both arrive and are versioned together.
|
||||
- **A node in permanent failure is silent, and that is now the last gap.** After a second
|
||||
failure the node is stopped and cannot report it. What notices is the mesh seeing a node that
|
||||
has not been heard from — which `09` records as a fact with no threshold, and which this makes
|
||||
more important to look at than it was.
|
||||
- **The rollback path is exercised only when it is needed**, which is when nobody can afford it
|
||||
to be wrong. It wants a test that boots a lab node onto a deliberately broken host and asserts
|
||||
the previous version comes back.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md) — the restart-by-exiting
|
||||
this protects.
|
||||
- [ADR 0039](0039-the-link-is-the-security-boundary.md) — why the watchdog cannot be remote.
|
||||
- [ADR 0036](0036-a-node-is-a-managed-machine.md) — why the failure is invisible without this.
|
||||
- [`09-the-node-lifecycle.md`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md) — the gap this closes.
|
||||
Reference in New Issue
Block a user