Commit Graph
2 Commits
Author SHA1 Message Date
jschoubben aeea2a9f9a Resolve the host lifecycle's open items, and say how the host is delivered
The upgrade question turned out to be a delivery question, so 0058 answers
both.

Today's third silo runs once per node and sends each one a command to install
and start. That is where the as-is records a package install that 404ed from
every mirror while the job went green, an image pull failure that did not fail
the deploy, and a verify stage that was built and never scheduled because it
was missing from a list.

The shape underneath all of those is that the thing reporting success was not
the thing doing the work. Meanwhile ADR 0037 has given every node a component
that applies state, reads back and reports -- so two mechanisms now change a
node and only one checks its work.

0058: a pipeline ends when the declaration is updated. Deploy stops sending
commands to nodes and becomes one write. The host applies it on its next
reconcile, and the host cannot report success it did not verify. The verify
stage disappears as a stage, which is the point -- verification stops being a
step that can be left off a list.

A pipeline result now means "the declaration is updated, and here is which
nodes have applied it". It does not wait for every node, because a node may be
legitimately switched off for a week. Outstanding is reported separately from
failed, since conflating them is how the old system produced a stall with no
error anywhere.

The host is delivered by exactly this path and needs no new resource type: a
`file` writes the package manager's config pointing at the mesh's repository, a
`package` names the version. Added a step I had missed -- before exiting for a
restart, the host runs the new binary once. A package can install something
that does not execute here, and that turns "the node never came back" into "the
apply failed and said why".

Six open items resolved: re-enrolment is decided when the token is issued and
revokes the previous identity; the mesh keeps a recovery copy of what each node
reports it owns, which un-strands the orphans; last-contact is reported with no
threshold, because a laptop off for three weeks is doing nothing wrong;
adoption always completes but a failed line makes a node ineligible for
assignment; a briefing is a structured document whose outcome is computed from
its lines; and the token is printed once and carried by hand, which is the
property that makes it worth anything.

Still open and named: automatic rollback of a host version that will not start.

0057 and 0058 are both proposed.
2026-08-27 21:53:38 +02:00
jschoubben 2204b01909 Design the node lifecycle end to end
The host was described as a component and never as something that runs for
years on a machine somebody else also uses. 09 covers every state a machine can
be in and every transition between them.

Four states: unmanaged, hosted, enrolled, disconnected. Only the last two are
nodes, and they are the same node in two situations. `hosted` -- the host
installed but never told which mesh it belongs to -- had no name before and is
where a machine sits between the two adoption commands.

Things that were unclear and now are not:

The first node walks the same path in an unusual order: reconcile from the
bundle, the control plane it just raised issues a token, enrol against it. Its
specialness lasts two commands. A side effect worth having -- enrolment is
exercised on node one, rather than being written and first used on node two.

Enrolment reports profile and inventory BEFORE the control plane decides
anything. The profile is the input to that decision, not a diagnostic; the
control plane cannot decide what a machine should run without knowing what it
can run.

Rebooting mid-apply is safe by construction. The store records each resource
after it worked, so a host that dies half way through comes back and applies
the rest. The rule that stops the host lying about what it did also makes it
crash-safe.

Retiring splits in two. Graceful is a final empty declaration. A node that is
gone will reconcile its last declaration forever -- the honest consequence of
making disconnection ordinary. The answer is not to make the host expire but
that the node holds nothing that outlives revocation: every grant is a per-node
credential revoked at the provider. A lost node keeps running and stops being
able to reach anything. Said plainly rather than implying the mesh can switch a
machine off, which it cannot and should not.

Losing the store is quiet and permanent, so it gets its own section. The host
re-enrols and re-applies fine; what does not come back is removal, because
resources it no longer has a record of become unowned and sit there
indefinitely.

Also corrects 0057, which said the mesh must not upgrade the host at all. That
conflated two acts. Replacing the binary is safe -- Unix keeps the running
inode. Stopping the unit is not. So the host may apply a package naming itself,
and restarts by finishing its apply and exiting cleanly, letting the supervisor
start it on the new binary. It never asks the service manager to restart it.
That makes a fleet-wide host upgrade an ordinary declaration, which the first
draft gave up on.

0057 remains proposed.
2026-08-27 21:16:44 +02:00