Changes are pushed, not polled; and a stuck host rolls itself back

Two corrections and one new decision, all from Jochen catching things.

Pushed, not polled. I described updates as landing "on the next reconcile",
which reads as polling and is not the design. A declaration arrives as a
message on a link that is already open; the host applies it then. Polling over
an existing connection would be slower to land AND constant traffic to learn
nothing.

The timer is for drift and nothing else, and it cannot be replaced by an event
for a definitional reason: drift is change the mesh did not make -- somebody
edited a managed file, a distribution upgrade replaced a config -- so nothing
will ever publish a message about it. Only looking finds it.

Separated the heartbeat from the reconcile timer, which I had been conflating.
They point in opposite directions and answer different questions: the timer
looks at the machine and asks whether it still matches; the heartbeat reports
upward and is what makes silence mean something. A node with nothing to do
sends nothing, and without a heartbeat that is indistinguishable from a node
that stopped.

0059 -- a host that cannot start is rolled back by the service manager. I had
left this open on the grounds that recovery meant the host judging its own
health. That objection does not survive being asked properly: a keepalive is
something else judging the host. The watchdog must be local, because nothing
dials a node and a host that cannot start cannot report -- so it is the service
manager, which is already there.

The failure it prevents is sharper than "the node is down": a host that will
not start looks exactly like a machine somebody switched off, which is the one
condition this design has deliberately decided not to alarm on. So a bad
release reaches every node, each goes quiet, and the mesh reports a fleet of
sleeping laptops.

Confirmed means started and completed one reconcile -- deliberately not "the
link is up", or a laptop on a train would roll itself back. The rollback is a
script shipped by the package, not a host subcommand, because a binary that
will not start cannot be its own recovery. It rolls back once: a second failure
means the machine is the problem, not the binary.

Also refined the records checker, which produced a false positive: a proposed
record may extend another proposed one, because decisions are drafted in chains
and the alternative is marking things accepted to satisfy a check. An accepted
document resting on a proposed record still fails, and that was verified.

0057, 0058 and 0059 are all proposed.
This commit is contained in:
2026-08-27 22:04:26 +02:00
parent aeea2a9f9a
commit 605c9fd441
5 changed files with 235 additions and 38 deletions
@@ -90,8 +90,13 @@ The host is tier 0, which makes it tempting to treat as special. It is not:
2. publish packages it and puts it in **the mesh's own package repository** — which is a
directory of files behind the object store and the proxy, so it needs no new machinery;
3. deploy updates each node's declaration to name the new version;
4. each host applies `package: nox-mesh-host` on its next reconcile, exactly as it applies any
other package.
4. **the new declaration is pushed to each node**, and the host applies
`package: nox-mesh-host` on arrival, exactly as it applies any other package.
Step 4 is a push and not a poll. The link is already open
([ADR 0001](0001-nodes-communicate-over-a-broker.md)), so a node learns of a change when it is
made — a node that is offline learns on reconnect, which is what makes *outstanding* a real
category rather than a euphemism for lost.
**No new resource type is needed**, which is the test of whether this is uniform or a special
case wearing a uniform. The repository is reachable because a `file` resource put its address in