Commit Graph
3 Commits
Author SHA1 Message Date
jschoubben 605c9fd441 Changes are pushed, not polled; and a stuck host rolls itself back
Two corrections and one new decision, all from Jochen catching things.

Pushed, not polled. I described updates as landing "on the next reconcile",
which reads as polling and is not the design. A declaration arrives as a
message on a link that is already open; the host applies it then. Polling over
an existing connection would be slower to land AND constant traffic to learn
nothing.

The timer is for drift and nothing else, and it cannot be replaced by an event
for a definitional reason: drift is change the mesh did not make -- somebody
edited a managed file, a distribution upgrade replaced a config -- so nothing
will ever publish a message about it. Only looking finds it.

Separated the heartbeat from the reconcile timer, which I had been conflating.
They point in opposite directions and answer different questions: the timer
looks at the machine and asks whether it still matches; the heartbeat reports
upward and is what makes silence mean something. A node with nothing to do
sends nothing, and without a heartbeat that is indistinguishable from a node
that stopped.

0059 -- a host that cannot start is rolled back by the service manager. I had
left this open on the grounds that recovery meant the host judging its own
health. That objection does not survive being asked properly: a keepalive is
something else judging the host. The watchdog must be local, because nothing
dials a node and a host that cannot start cannot report -- so it is the service
manager, which is already there.

The failure it prevents is sharper than "the node is down": a host that will
not start looks exactly like a machine somebody switched off, which is the one
condition this design has deliberately decided not to alarm on. So a bad
release reaches every node, each goes quiet, and the mesh reports a fleet of
sleeping laptops.

Confirmed means started and completed one reconcile -- deliberately not "the
link is up", or a laptop on a train would roll itself back. The rollback is a
script shipped by the package, not a host subcommand, because a binary that
will not start cannot be its own recovery. It rolls back once: a second failure
means the machine is the problem, not the binary.

Also refined the records checker, which produced a false positive: a proposed
record may extend another proposed one, because decisions are drafted in chains
and the alternative is marking things accepted to satisfy a check. An accepted
document resting on a proposed record still fails, and that was verified.

0057, 0058 and 0059 are all proposed.
2026-08-27 22:04:26 +02:00
jschoubben 2204b01909 Design the node lifecycle end to end
The host was described as a component and never as something that runs for
years on a machine somebody else also uses. 09 covers every state a machine can
be in and every transition between them.

Four states: unmanaged, hosted, enrolled, disconnected. Only the last two are
nodes, and they are the same node in two situations. `hosted` -- the host
installed but never told which mesh it belongs to -- had no name before and is
where a machine sits between the two adoption commands.

Things that were unclear and now are not:

The first node walks the same path in an unusual order: reconcile from the
bundle, the control plane it just raised issues a token, enrol against it. Its
specialness lasts two commands. A side effect worth having -- enrolment is
exercised on node one, rather than being written and first used on node two.

Enrolment reports profile and inventory BEFORE the control plane decides
anything. The profile is the input to that decision, not a diagnostic; the
control plane cannot decide what a machine should run without knowing what it
can run.

Rebooting mid-apply is safe by construction. The store records each resource
after it worked, so a host that dies half way through comes back and applies
the rest. The rule that stops the host lying about what it did also makes it
crash-safe.

Retiring splits in two. Graceful is a final empty declaration. A node that is
gone will reconcile its last declaration forever -- the honest consequence of
making disconnection ordinary. The answer is not to make the host expire but
that the node holds nothing that outlives revocation: every grant is a per-node
credential revoked at the provider. A lost node keeps running and stops being
able to reach anything. Said plainly rather than implying the mesh can switch a
machine off, which it cannot and should not.

Losing the store is quiet and permanent, so it gets its own section. The host
re-enrols and re-applies fine; what does not come back is removal, because
resources it no longer has a record of become unowned and sit there
indefinitely.

Also corrects 0057, which said the mesh must not upgrade the host at all. That
conflated two acts. Replacing the binary is safe -- Unix keeps the running
inode. Stopping the unit is not. So the host may apply a package naming itself,
and restarts by finishing its apply and exiting cleanly, letting the supervisor
start it on the new binary. It never asks the service manager to restart it.
That makes a fleet-wide host upgrade an ordinary declaration, which the first
draft gave up on.

0057 remains proposed.
2026-08-27 21:16:44 +02:00
jschoubben 3ab11c96ef Say what the host process is: a root service, installed as a package
The design described what the host does and never what it is at runtime. The
words daemon, long-running, interval, poll and heartbeat appeared nowhere in it
or in the relevant decisions. What exists is a command that runs and exits;
what the design needs is a process holding a link. Nobody had written down that
those differ, so several questions had no answer.

0057 settles them. It runs on every node -- the host is what makes a machine
managed, so a machine without one is not a node. Root, because no useful subset
of the job is unprivileged. A systemd unit, because something must survive a
reboot to hold the link.

It never manages its own unit. The temptation is obvious and it ends with a
host stopping itself half way through an apply, leaving a machine with nothing
running to fix it. The installation owns the host; the host owns everything
else.

Installed as a package, with a tarball as the floor. The package carries the
unit file, the state directory and an upgrade path, which a bare binary does
not. But the mesh's package repository is hosted on the mesh, so any route that
needs the mesh to install the thing that joins the mesh is a circle -- the
tarball is the path that must never acquire a dependency.

Reconciles on start, on a declaration, on a timer and on reconnect. The timer
is the one easy to leave out, and without it `owned` reports what the host
applied rather than what is there -- ADR 0035 violated by omission.

The records checker caught this commit on its first attempt: 05 listed 0057 in
its frontmatter while 0057 is still proposed, and a to-be document may not rest
on an unaccepted record. The section now says so in the body instead.
2026-08-27 21:06:00 +02:00