Design the node lifecycle end to end

The host was described as a component and never as something that runs for
years on a machine somebody else also uses. 09 covers every state a machine can
be in and every transition between them.

Four states: unmanaged, hosted, enrolled, disconnected. Only the last two are
nodes, and they are the same node in two situations. `hosted` -- the host
installed but never told which mesh it belongs to -- had no name before and is
where a machine sits between the two adoption commands.

Things that were unclear and now are not:

The first node walks the same path in an unusual order: reconcile from the
bundle, the control plane it just raised issues a token, enrol against it. Its
specialness lasts two commands. A side effect worth having -- enrolment is
exercised on node one, rather than being written and first used on node two.

Enrolment reports profile and inventory BEFORE the control plane decides
anything. The profile is the input to that decision, not a diagnostic; the
control plane cannot decide what a machine should run without knowing what it
can run.

Rebooting mid-apply is safe by construction. The store records each resource
after it worked, so a host that dies half way through comes back and applies
the rest. The rule that stops the host lying about what it did also makes it
crash-safe.

Retiring splits in two. Graceful is a final empty declaration. A node that is
gone will reconcile its last declaration forever -- the honest consequence of
making disconnection ordinary. The answer is not to make the host expire but
that the node holds nothing that outlives revocation: every grant is a per-node
credential revoked at the provider. A lost node keeps running and stops being
able to reach anything. Said plainly rather than implying the mesh can switch a
machine off, which it cannot and should not.

Losing the store is quiet and permanent, so it gets its own section. The host
re-enrols and re-applies fine; what does not come back is removal, because
resources it no longer has a record of become unowned and sit there
indefinitely.

Also corrects 0057, which said the mesh must not upgrade the host at all. That
conflated two acts. Replacing the binary is safe -- Unix keeps the running
inode. Stopping the unit is not. So the host may apply a package naming itself,
and restarts by finishing its apply and exiting cleanly, letting the supervisor
start it on the new binary. It never asks the service manager to restart it.
That makes a fleet-wide host upgrade an ordinary declaration, which the first
draft gave up on.

0057 remains proposed.
This commit is contained in:
2026-08-27 21:16:44 +02:00
parent 3ab11c96ef
commit 2204b01909
3 changed files with 344 additions and 8 deletions
@@ -72,14 +72,33 @@ and neither can a node whose mesh is down — which is exactly when somebody is
**Any path that requires the mesh to install the thing that joins the mesh is a circle**, so the
tarball is the path that is never allowed to acquire a dependency.
### The mesh does not upgrade the host
### The host may replace its own binary; it may not stop its own unit
A running process replacing its own binary and restarting mid-apply is the self-management
problem wearing a different hat. **Upgrading the host is an act on the machine**, by the package
manager, not a declaration the host applies to itself.
The first draft of this record said the mesh must not upgrade the host at all. That was too
broad, and it conflated two different acts.
Recorded as a limit rather than a plan: it means a fleet-wide host upgrade is not currently a
mesh operation, and that will be felt.
**Replacing the binary is safe.** Unix keeps the running executable's inode open, so a package
upgrade writes a new file and the running process continues on the old one, undisturbed.
**Stopping the unit is what is unsafe** — that is the host killing itself part-way through an
apply, leaving a machine with nothing running to finish or fix it.
So the host may apply a `package` naming itself. What it must never do is ask the service
manager to restart it.
**The restart happens by exiting, not by asking.** When the host notices its own executable has
been replaced, it finishes the apply it is in, reports what it did, and **exits cleanly**. The
supervisor's `Restart=always` starts it again, on the new binary. Nothing stops the host; the
host stops, having finished.
Three conditions, and they are the whole safety argument:
- **after** the apply completes and its outcomes are recorded — never mid-way;
- **only** when the executable actually changed, which Linux reports plainly: a replaced
`/proc/self/exe` reads as the old path marked deleted;
- **exit zero**, so a restart is what a supervisor does next rather than a failure it backs off
from.
This makes a fleet-wide host upgrade an ordinary declaration, which the first draft gave up.
### It reconciles on four triggers
@@ -115,8 +134,14 @@ host already knows it owns.
- **The timer makes drift visible and also makes it loud.** A resource the host cannot apply
will now fail every ten minutes rather than once. That is correct and it needs somewhere to go
other than a log nobody reads — which is `observability`'s, and it does not exist yet.
- **Host upgrades are outside the mesh**, so a fleet-wide upgrade is currently manual. Nothing
here solves it and it should not be solved by giving the host a self-upgrade path.
- **A host upgrade is an ordinary declaration**, which is worth the care it needs: the exit
path is the only place the host deliberately stops, and a bug there is a node that restarts in
a loop or never comes back. It wants a test that the host does **not** exit when its binary is
unchanged, as much as one that it does when it changed.
- **A version-skewed fleet is now normal and needs saying.** Nodes restart onto the new binary
at whatever moment their apply finishes, so "the fleet is upgraded" is a range rather than an
instant. What a node reports as its version must be the **running** one, not the installed
one, or the mesh will believe an upgrade landed before it took effect.
## References