Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24
@@ -153,6 +153,12 @@ def check_rests_on(failures, records):
|
||||
continue
|
||||
status = records[number]["front"].get("status")
|
||||
if status != "accepted":
|
||||
# A proposed record may extend another proposed one. Decisions are drafted in
|
||||
# chains -- 0059 extends 0057 while both await review -- and refusing that would
|
||||
# mean either drafting out of order or marking records accepted to satisfy a
|
||||
# check, which is the failure this repository already made once.
|
||||
if frontmatter(read(path)).get("status") == "proposed":
|
||||
continue
|
||||
# An as-is document describes what runs, and what runs was built under
|
||||
# whatever was decided at the time. ADR 0056: "as-is describing a superseded
|
||||
# decision is exactly what as-is is for."
|
||||
|
||||
@@ -100,24 +100,43 @@ Three conditions, and they are the whole safety argument:
|
||||
|
||||
This makes a fleet-wide host upgrade an ordinary declaration, which the first draft gave up.
|
||||
|
||||
### It reconciles on four triggers
|
||||
### Changes are pushed. The timer is for drift, and only for drift
|
||||
|
||||
| Trigger | Why |
|
||||
|---|---|
|
||||
| **start** | the machine may have changed while nothing was running |
|
||||
| **a declaration arrives** | the ordinary path |
|
||||
| **a timer** | drift. Something other than the host changed the machine — a person, a package upgrade replacing a config file |
|
||||
| **reconnect** | it may have missed declarations while disconnected ([ADR 0036](0036-a-node-is-a-managed-machine.md)) |
|
||||
**The host does not poll for work.** A new declaration arrives as a message on the link, and the
|
||||
host applies it then ([ADR 0001](0001-nodes-communicate-over-a-broker.md)). Polling for updates
|
||||
over a connection that already exists would be strictly worse in both directions: slower to
|
||||
land, and constant traffic to learn nothing.
|
||||
|
||||
**The timer is what makes the store's claim true.** Without it a machine that drifted stays
|
||||
drifted until somebody changes a declaration, and `owned` reports what the host *applied* rather
|
||||
than what is *there* — which is
|
||||
Four triggers, and only one of them is a clock:
|
||||
|
||||
| Trigger | Kind | Why |
|
||||
|---|---|---|
|
||||
| **a declaration arrives** | **pushed** | the ordinary path — this is how changes land |
|
||||
| **start** | event | the machine may have changed while nothing was running |
|
||||
| **reconnect** | event | declarations may have been missed while disconnected |
|
||||
| **a timer** | periodic | **drift**, and nothing else |
|
||||
|
||||
**The timer cannot be replaced by an event, and the reason is definitional.** Drift is change the
|
||||
*mesh did not make* — somebody edited a managed file, a distribution upgrade replaced a config,
|
||||
a container was stopped by hand. **Nothing will ever send a message about it**, because the thing
|
||||
that did it is not part of the mesh. A local periodic check is the only way to see it at all.
|
||||
|
||||
Without it, `owned` reports what the host *applied* rather than what is *there*, which is
|
||||
[ADR 0035](0035-a-picture-is-read-from-what-runs.md) violated by omission.
|
||||
|
||||
**Ten minutes**, configurable. Short enough that drift is bounded by something a person would
|
||||
notice anyway, long enough that a fleet is not doing constant work. The reconcile is cheap: it
|
||||
asks the package database, the service manager and the container runtime about resources the
|
||||
host already knows it owns.
|
||||
**Ten minutes**, configurable. The check is cheap: it asks the package database, the service
|
||||
manager and the container runtime about resources the host already knows it owns.
|
||||
|
||||
### It reports upward on a heartbeat
|
||||
|
||||
Separate from reconciling, and easy to conflate with it: the node tells the mesh what it is —
|
||||
its running version, what it holds, what it last applied — on link, after every apply, and
|
||||
periodically while idle.
|
||||
|
||||
**The heartbeat is what makes silence mean something.** Without it, the mesh cannot distinguish
|
||||
a node that is fine and has had nothing to do from one that stopped. With it, *last heard from*
|
||||
is a fact per node, and [ADR 0059](0059-a-host-that-cannot-start-rolls-itself-back.md) is what
|
||||
handles the case where the node cannot report at all.
|
||||
|
||||
## Consequences
|
||||
|
||||
|
||||
@@ -90,8 +90,13 @@ The host is tier 0, which makes it tempting to treat as special. It is not:
|
||||
2. publish packages it and puts it in **the mesh's own package repository** — which is a
|
||||
directory of files behind the object store and the proxy, so it needs no new machinery;
|
||||
3. deploy updates each node's declaration to name the new version;
|
||||
4. each host applies `package: nox-mesh-host` on its next reconcile, exactly as it applies any
|
||||
other package.
|
||||
4. **the new declaration is pushed to each node**, and the host applies
|
||||
`package: nox-mesh-host` on arrival, exactly as it applies any other package.
|
||||
|
||||
Step 4 is a push and not a poll. The link is already open
|
||||
([ADR 0001](0001-nodes-communicate-over-a-broker.md)), so a node learns of a change when it is
|
||||
made — a node that is offline learns on reconnect, which is what makes *outstanding* a real
|
||||
category rather than a euphemism for lost.
|
||||
|
||||
**No new resource type is needed**, which is the test of whether this is uniform or a special
|
||||
case wearing a uniform. The repository is reachable because a `file` resource put its address in
|
||||
|
||||
@@ -0,0 +1,133 @@
|
||||
---
|
||||
status: proposed
|
||||
date: 2026-08-27
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 0057-the-host-is-a-root-service-installed-as-a-package.md
|
||||
---
|
||||
|
||||
# 59. A host that cannot start is rolled back by the supervisor
|
||||
|
||||
## Context
|
||||
|
||||
[`09-the-node-lifecycle.md`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md) left one thing
|
||||
unsolved: a host upgraded to a version that crashes on start. The supervisor restarts it, backs
|
||||
off, and the node is stuck on a binary that will not run.
|
||||
|
||||
It was left because automatic recovery looked like *the host judging its own health*, which is
|
||||
the self-reference the rest of that document avoids.
|
||||
|
||||
**That objection does not survive being asked properly.** A keepalive is not the host judging
|
||||
itself — it is something else judging the host. The question is only *what*, and that has one
|
||||
answer.
|
||||
|
||||
### Why the watchdog must be local
|
||||
|
||||
The obvious candidate is the control plane, and it cannot be:
|
||||
|
||||
- **Nothing dials a node.** [ADR 0039](0039-the-link-is-the-security-boundary.md) makes the link
|
||||
outbound and node-initiated, and a node has no listening control surface. The mesh has no way
|
||||
to reach in and act.
|
||||
- **The failure removes the reporting path.** A host that cannot start cannot link, so the mesh
|
||||
learns nothing to act on.
|
||||
|
||||
So the watchdog is local, and the only local thing that is always present, already depended on,
|
||||
and is not the host is the **service manager**.
|
||||
|
||||
### The failure this actually prevents
|
||||
|
||||
Sharper than "the node is down", and it is the reason to bother:
|
||||
|
||||
> **A host that will not start looks exactly like a machine that was switched off.**
|
||||
|
||||
[ADR 0036](0036-a-node-is-a-managed-machine.md) makes disconnection ordinary, and
|
||||
[`09`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md) deliberately puts no alarm on it — a
|
||||
laptop shut for three weeks is doing nothing wrong. **A bad upgrade is therefore invisible**: it
|
||||
presents as the one condition the design has decided not to be alarmed by.
|
||||
|
||||
Worse, it arrives fleet-wide. A declaration naming a bad version reaches every node, and each
|
||||
one applies it, restarts, and stops talking. The mesh would report a fleet of quiet nodes and
|
||||
nothing else.
|
||||
|
||||
## Decision
|
||||
|
||||
**The service manager rolls the host back to the last version that started.**
|
||||
|
||||
Four parts, and each one is chosen so it works when the host does not:
|
||||
|
||||
### 1 — A version is confirmed by starting, not by seeming well
|
||||
|
||||
On start, the host completes one full reconcile. If it does, it writes the running version to a
|
||||
plain file — `/var/lib/mesh-host/known-good` — and that is the whole of "confirmed".
|
||||
|
||||
**Deliberately not health.** Not *the link is up*, because a disconnected node is ordinary and a
|
||||
laptop on a train would roll itself back. Not *everything applied cleanly*, because a resource
|
||||
that fails is the machine's problem and not the binary's. The claim is narrow and checkable: **it
|
||||
started, and it got through a reconcile.**
|
||||
|
||||
### 2 — The rollback is not the host binary
|
||||
|
||||
The obvious mistake, and it would make the whole mechanism a no-op: `nox-mesh-host rollback`
|
||||
cannot be the recovery path for a `nox-mesh-host` that does not run.
|
||||
|
||||
Rollback is a **small script shipped by the package**, which reads the `known-good` file and
|
||||
asks the package manager to install that version. It shares no code with the host and does not
|
||||
import it.
|
||||
|
||||
### 3 — The supervisor triggers it, after giving up
|
||||
|
||||
```ini
|
||||
[Service]
|
||||
Restart=on-failure
|
||||
StartLimitBurst=3
|
||||
StartLimitIntervalSec=120
|
||||
|
||||
[Unit]
|
||||
OnFailure=nox-mesh-host-rollback.service
|
||||
```
|
||||
|
||||
Three failures in two minutes is a binary that does not work, not a transient. The supervisor
|
||||
stops trying and runs the rollback unit, which downgrades and starts the host again.
|
||||
|
||||
### 4 — It rolls back once
|
||||
|
||||
The rollback unit records that it fired. If the rolled-back version *also* fails to start, it
|
||||
does **not** fire again — the node stops, loudly, in a failed state.
|
||||
|
||||
**Because a second rollback is a different diagnosis.** Once the previously-working binary also
|
||||
fails, the binary is not the problem: the machine is. Rolling back further would flap between
|
||||
two versions forever and bury the actual cause under a loop.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **The node keeps its own recovery**, which is the property the whole tier-0 design rests on:
|
||||
the host depends on nothing, and now its recovery depends on nothing either.
|
||||
- **The package cache must retain the previous version**, and that is a real requirement rather
|
||||
than an assumption — a package manager configured to clean its cache would delete the thing
|
||||
rollback needs. The package must pin that, and it is the sort of dependency that is discovered
|
||||
by the rollback failing.
|
||||
- **A fleet-wide bad upgrade becomes self-limiting.** Each node fails, rolls back, and comes
|
||||
back on the previous version — so the blast radius of a bad host release is one restart cycle
|
||||
per node rather than the entire fleet stopping.
|
||||
- **It catches "will not start" and nothing else.** A version that starts and is subtly wrong
|
||||
will not roll back, and should not: that is a bad release, which is delivery's problem
|
||||
([ADR 0058](0058-delivery-ends-in-a-declaration.md)), not a supervision problem.
|
||||
- **It adds a unit and a script the host does not own**, which is
|
||||
[ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)'s line holding: the
|
||||
installation owns the host, and now it owns the host's recovery too. Consistent, and it means
|
||||
both arrive and are versioned together.
|
||||
- **A node in permanent failure is silent, and that is now the last gap.** After a second
|
||||
failure the node is stopped and cannot report it. What notices is the mesh seeing a node that
|
||||
has not been heard from — which `09` records as a fact with no threshold, and which this makes
|
||||
more important to look at than it was.
|
||||
- **The rollback path is exercised only when it is needed**, which is when nobody can afford it
|
||||
to be wrong. It wants a test that boots a lab node onto a deliberately broken host and asserts
|
||||
the previous version comes back.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md) — the restart-by-exiting
|
||||
this protects.
|
||||
- [ADR 0039](0039-the-link-is-the-security-boundary.md) — why the watchdog cannot be remote.
|
||||
- [ADR 0036](0036-a-node-is-a-managed-machine.md) — why the failure is invisible without this.
|
||||
- [`09-the-node-lifecycle.md`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md) — the gap this closes.
|
||||
@@ -158,23 +158,43 @@ outcome **derived** from the worst line rather than stated alongside it.
|
||||
|
||||
## enrolled: what running actually looks like
|
||||
|
||||
Four reconcile triggers
|
||||
([ADR 0057](../../02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md)):
|
||||
**Changes are pushed, not polled.** A declaration arrives as a message on the link and the host
|
||||
applies it then. The link is already open and outbound
|
||||
([ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md),
|
||||
[ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md)) — asking it
|
||||
repeatedly whether anything has changed would be slower to land *and* constant traffic to learn
|
||||
nothing.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| **start** | the machine may have changed while nothing was running |
|
||||
| **a declaration arrives** | the ordinary path |
|
||||
| **every ten minutes** | drift — something other than the host changed the machine |
|
||||
| **reconnect** | declarations may have been missed |
|
||||
| Trigger | Kind | |
|
||||
|---|---|---|
|
||||
| **a declaration arrives** | **pushed** | the ordinary path — this is how a change lands |
|
||||
| **start** | event | the machine may have changed while nothing was running |
|
||||
| **reconnect** | event | declarations may have been missed |
|
||||
| **every ten minutes** | periodic | **drift, and only drift** |
|
||||
|
||||
**Rebooting mid-apply is safe, and it is safe by construction.** The store records each resource
|
||||
*after* it worked ([ADR 0035](../../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md)), so
|
||||
a host that dies half way through comes back, finds the completed ones already matching, and
|
||||
applies the rest. The rule that exists to stop the host lying about what it did also makes it
|
||||
crash-safe.
|
||||
**Why the timer cannot be an event.** Drift is change the *mesh did not make* — somebody edited
|
||||
a managed file, a distribution upgrade replaced a config, a container was stopped by hand.
|
||||
Nothing will ever publish a message about it, because whatever did it is not part of the mesh.
|
||||
Only looking finds it.
|
||||
|
||||
---
|
||||
So the two periodic things do different jobs and should not be conflated:
|
||||
|
||||
| | direction | answers |
|
||||
|---|---|---|
|
||||
| **reconcile timer** | local, looks at the machine | *does this machine still match what it was told?* |
|
||||
| **heartbeat** | upward, reports to the mesh | *is this node still here, and what is it running?* |
|
||||
|
||||
**The heartbeat is what makes silence mean something.** A node with nothing to do sends nothing;
|
||||
without a heartbeat that is indistinguishable from a node that stopped. With one, *last heard
|
||||
from* is a fact beside every node — which is what
|
||||
[how long disconnected](#how-long-disconnected-and-who-is-told) reports and what
|
||||
[ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md) exists because
|
||||
a stuck node cannot send.
|
||||
|
||||
**Rebooting mid-apply is safe by construction.** The store records each resource *after* it
|
||||
worked ([ADR 0035](../../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md)), so a host that
|
||||
dies half way through comes back, finds the completed ones already matching, and applies the
|
||||
rest. The rule that exists to stop the host lying about what it did also makes it crash-safe.
|
||||
|
||||
## Updating what the node holds
|
||||
|
||||
@@ -290,7 +310,8 @@ push to mesh-host
|
||||
├─ publish packaged, into the mesh's own package repository
|
||||
└─ deploy each node's declaration now names the new version
|
||||
│
|
||||
└─ every host applies it on its next reconcile
|
||||
└─ pushed to each node; the host applies it on arrival
|
||||
(a node that is offline gets it on reconnect)
|
||||
```
|
||||
|
||||
**Compared with today.** The current pipeline's third silo runs *once per node* and sends each
|
||||
@@ -327,11 +348,21 @@ part-way through an apply. It stops by finishing.
|
||||
own apply completes. A node must therefore report the version it is **running**, not the one
|
||||
installed — otherwise the mesh believes an upgrade landed at step 1.
|
||||
|
||||
**What is not solved: a new version that crashes on start.** systemd will restart it, back off,
|
||||
and the node is stuck on a binary that does not run. The previous package is still in the
|
||||
package manager's cache, so a person at the machine can downgrade — but there is no automatic
|
||||
rollback, and designing one means the host judging its own health, which is the kind of
|
||||
self-reference the rest of this document avoids. **Named, not solved.**
|
||||
**A version that crashes on start rolls itself back**
|
||||
([ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md)). The service
|
||||
manager gives up after three failures in two minutes and runs a rollback script — shipped by the
|
||||
package rather than being a host subcommand, because a binary that will not start cannot be its
|
||||
own recovery. It reinstalls the version recorded in `known-good`, which the host wrote the last
|
||||
time it completed a reconcile.
|
||||
|
||||
**It rolls back once.** If the previous version also fails, the node stops in a failed state
|
||||
rather than flapping between two binaries. A second failure is a different diagnosis: the
|
||||
machine is the problem, not the binary.
|
||||
|
||||
**Why this matters more than it looks.** A host that will not start cannot link, and a node that
|
||||
is not linking looks exactly like a machine somebody switched off — which is the one condition
|
||||
this design has deliberately decided not to alarm on. Without rollback, a bad release reaches
|
||||
every node, each one goes quiet, and the mesh reports a fleet of sleeping laptops.
|
||||
|
||||
---
|
||||
|
||||
@@ -443,8 +474,11 @@ the same command against a mesh that is one machine old.
|
||||
|
||||
## Still open
|
||||
|
||||
- **Automatic rollback of a bad host version.** Above. A node can be left on a binary that will
|
||||
not start, and recovering needs a person at the machine.
|
||||
- ~~**Automatic rollback of a bad host version.**~~ **Resolved** by
|
||||
[ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md): the service
|
||||
manager gives up after three failures and runs a rollback script — shipped by the package, not
|
||||
the host binary, because a binary that will not start cannot recover itself. It rolls back
|
||||
once; a second failure means the machine is the problem, not the binary.
|
||||
- **How a previous declaration is retained and chosen**, which is what rollback of anything else
|
||||
would use ([ADR 0058](../../02-DECISIONS/0058-delivery-ends-in-a-declaration.md)).
|
||||
- **A node returning after months** applies a very large jump in one go. Correct, and untested.
|
||||
|
||||
Reference in New Issue
Block a user