Changes are pushed, not polled; and a stuck host rolls itself back

Two corrections and one new decision, all from Jochen catching things.

Pushed, not polled. I described updates as landing "on the next reconcile",
which reads as polling and is not the design. A declaration arrives as a
message on a link that is already open; the host applies it then. Polling over
an existing connection would be slower to land AND constant traffic to learn
nothing.

The timer is for drift and nothing else, and it cannot be replaced by an event
for a definitional reason: drift is change the mesh did not make -- somebody
edited a managed file, a distribution upgrade replaced a config -- so nothing
will ever publish a message about it. Only looking finds it.

Separated the heartbeat from the reconcile timer, which I had been conflating.
They point in opposite directions and answer different questions: the timer
looks at the machine and asks whether it still matches; the heartbeat reports
upward and is what makes silence mean something. A node with nothing to do
sends nothing, and without a heartbeat that is indistinguishable from a node
that stopped.

0059 -- a host that cannot start is rolled back by the service manager. I had
left this open on the grounds that recovery meant the host judging its own
health. That objection does not survive being asked properly: a keepalive is
something else judging the host. The watchdog must be local, because nothing
dials a node and a host that cannot start cannot report -- so it is the service
manager, which is already there.

The failure it prevents is sharper than "the node is down": a host that will
not start looks exactly like a machine somebody switched off, which is the one
condition this design has deliberately decided not to alarm on. So a bad
release reaches every node, each goes quiet, and the mesh reports a fleet of
sleeping laptops.

Confirmed means started and completed one reconcile -- deliberately not "the
link is up", or a laptop on a train would roll itself back. The rollback is a
script shipped by the package, not a host subcommand, because a binary that
will not start cannot be its own recovery. It rolls back once: a second failure
means the machine is the problem, not the binary.

Also refined the records checker, which produced a false positive: a proposed
record may extend another proposed one, because decisions are drafted in chains
and the alternative is marking things accepted to satisfy a check. An accepted
document resting on a proposed record still fails, and that was verified.

0057, 0058 and 0059 are all proposed.
This commit is contained in:
2026-08-27 22:04:26 +02:00
parent aeea2a9f9a
commit 605c9fd441
5 changed files with 235 additions and 38 deletions
+56 -22
View File
@@ -158,23 +158,43 @@ outcome **derived** from the worst line rather than stated alongside it.
## enrolled: what running actually looks like
Four reconcile triggers
([ADR 0057](../../02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md)):
**Changes are pushed, not polled.** A declaration arrives as a message on the link and the host
applies it then. The link is already open and outbound
([ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md),
[ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md)) — asking it
repeatedly whether anything has changed would be slower to land *and* constant traffic to learn
nothing.
| | |
|---|---|
| **start** | the machine may have changed while nothing was running |
| **a declaration arrives** | the ordinary path |
| **every ten minutes** | drift — something other than the host changed the machine |
| **reconnect** | declarations may have been missed |
| Trigger | Kind | |
|---|---|---|
| **a declaration arrives** | **pushed** | the ordinary path — this is how a change lands |
| **start** | event | the machine may have changed while nothing was running |
| **reconnect** | event | declarations may have been missed |
| **every ten minutes** | periodic | **drift, and only drift** |
**Rebooting mid-apply is safe, and it is safe by construction.** The store records each resource
*after* it worked ([ADR 0035](../../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md)), so
a host that dies half way through comes back, finds the completed ones already matching, and
applies the rest. The rule that exists to stop the host lying about what it did also makes it
crash-safe.
**Why the timer cannot be an event.** Drift is change the *mesh did not make* — somebody edited
a managed file, a distribution upgrade replaced a config, a container was stopped by hand.
Nothing will ever publish a message about it, because whatever did it is not part of the mesh.
Only looking finds it.
---
So the two periodic things do different jobs and should not be conflated:
| | direction | answers |
|---|---|---|
| **reconcile timer** | local, looks at the machine | *does this machine still match what it was told?* |
| **heartbeat** | upward, reports to the mesh | *is this node still here, and what is it running?* |
**The heartbeat is what makes silence mean something.** A node with nothing to do sends nothing;
without a heartbeat that is indistinguishable from a node that stopped. With one, *last heard
from* is a fact beside every node — which is what
[how long disconnected](#how-long-disconnected-and-who-is-told) reports and what
[ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md) exists because
a stuck node cannot send.
**Rebooting mid-apply is safe by construction.** The store records each resource *after* it
worked ([ADR 0035](../../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md)), so a host that
dies half way through comes back, finds the completed ones already matching, and applies the
rest. The rule that exists to stop the host lying about what it did also makes it crash-safe.
## Updating what the node holds
@@ -290,7 +310,8 @@ push to mesh-host
├─ publish packaged, into the mesh's own package repository
└─ deploy each node's declaration now names the new version
│
└─ every host applies it on its next reconcile
└─ pushed to each node; the host applies it on arrival
(a node that is offline gets it on reconnect)
```
**Compared with today.** The current pipeline's third silo runs *once per node* and sends each
@@ -327,11 +348,21 @@ part-way through an apply. It stops by finishing.
own apply completes. A node must therefore report the version it is **running**, not the one
installed — otherwise the mesh believes an upgrade landed at step 1.
**What is not solved: a new version that crashes on start.** systemd will restart it, back off,
and the node is stuck on a binary that does not run. The previous package is still in the
package manager's cache, so a person at the machine can downgrade — but there is no automatic
rollback, and designing one means the host judging its own health, which is the kind of
self-reference the rest of this document avoids. **Named, not solved.**
**A version that crashes on start rolls itself back**
([ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md)). The service
manager gives up after three failures in two minutes and runs a rollback script — shipped by the
package rather than being a host subcommand, because a binary that will not start cannot be its
own recovery. It reinstalls the version recorded in `known-good`, which the host wrote the last
time it completed a reconcile.
**It rolls back once.** If the previous version also fails, the node stops in a failed state
rather than flapping between two binaries. A second failure is a different diagnosis: the
machine is the problem, not the binary.
**Why this matters more than it looks.** A host that will not start cannot link, and a node that
is not linking looks exactly like a machine somebody switched off — which is the one condition
this design has deliberately decided not to alarm on. Without rollback, a bad release reaches
every node, each one goes quiet, and the mesh reports a fleet of sleeping laptops.
---
@@ -443,8 +474,11 @@ the same command against a mesh that is one machine old.
## Still open
- **Automatic rollback of a bad host version.** Above. A node can be left on a binary that will
not start, and recovering needs a person at the machine.
- ~~**Automatic rollback of a bad host version.**~~ **Resolved** by
[ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md): the service
manager gives up after three failures and runs a rollback script — shipped by the package, not
the host binary, because a binary that will not start cannot recover itself. It rolls back
once; a second failure means the machine is the problem, not the binary.
- **How a previous declaration is retained and chosen**, which is what rollback of anything else
would use ([ADR 0058](../../02-DECISIONS/0058-delivery-ends-in-a-declaration.md)).
- **A node returning after months** applies a very large jump in one go. Correct, and untested.