Files
hq/02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md
T
jschoubben 19997d56c3 Approve 0057-0059, with four corrections from review
Not approved as drafted -- four things came out of checking them against each
other, and one was a bug that would have broken every upgrade.

The bug: 0059 specified Restart=on-failure while 0057 has the host restart onto
a new binary by exiting CLEANLY. on-failure does not restart a process that
exited zero, so every upgraded node would have been left stopped, having
successfully upgraded. Found by reading the two records against each other
rather than by either alone. Now Restart=always in all three places that
mention it.

The host cannot run in a container, and the reason is decisive rather than
stylistic: step 0 of the substrate bootstrap installs the container runtime, so
a host inside a container would need the thing it exists to install. It would
also break 0041 -- copy it onto a machine and run it stops being true when the
machine must already have a runtime. Everything above tier 0 is a container;
the host is not. That split is the tier boundary, not an inconsistency.

systemd is named rather than abstracted. An init is not a dependency in 0041's
sense: 0041 is about what must be installed before the host works, and an init
is not installed, it is what the machine already is. The unit file is the only
systemd-specific artefact and it belongs to the package, so a machine with a
different supervisor ships a different package.

The mesh is a watchdog, and my first draft was half an answer. Recovery must be
local -- nothing dials a node, and a host that cannot start cannot report. But
detection is the mesh's, and a local supervisor structurally cannot do it: it
sees one process failing and cannot tell a broken machine from a broken
release. Only something watching every node can, and that distinction decides
whether the response is "fix this machine" or "stop shipping this version". So
a host rollout is staged -- a few nodes, wait for heartbeats, continue or stop
on silence. Local rollback still needed, because the canary nodes break and
because a node offline during the rollout gets the declaration later with no
batch around it.

The first declaration is the overlay and nothing else. Forced, because a node's
address and peers are assigned rather than chosen. But also the way back in: a
node reachable over the overlay can be fixed by hand if a later declaration
breaks it, and a large first declaration risks a node that is broken and
unreachable at once.

Also stated plainly, because it reads as a contradiction: nodes reach each
other over the overlay and every node consumes from the broker; what 0039
forbids is an inbound CONTROL surface, not reachability.

And in 06: no node holds a credential to any control-plane store, for reads or
writes. Four ADRs already say this separately and none of them said it in one
place. Nodes state over the broker; the owning context writes. With a note that
most high-frequency writes are observability's, not the registry's -- routing
logs into the registry would be the shared-schema mistake arriving through a
door marked performance.
2026-08-27 22:12:12 +02:00

11 KiB

status, date, deciders, reconstructed, extends
status date deciders reconstructed extends
accepted 2026-08-27 jochen false 0041-the-host-depends-on-nothing.md

57. The host is a root service, installed as a package, that never manages itself

Context

05-the-node-host.md describes what the host does and never says what it is at runtime. Searched: the words daemon, long-running, interval, poll and heartbeat appear in none of it, nor in ADR 0038 or ADR 0039.

What exists today is a command that runs and exits — mesh-host apply FILE. What the design requires is a process holding an outbound link to the control plane. Nobody wrote down that these are different things, so several questions have no answer: does it reconcile on a timer, what happens when the link drops, and who installs the unit that starts it — given that the host is the thing that installs units.

Decision

It runs on every node, and that is what a node is

ADR 0036 defines a node as a managed machine. The host is what makes it managed, so a machine without one is not a node with a missing component; it is not a node. There is no partial mode, no agentless node, and no second way in.

It is a root service

Root, because there is no useful subset of its job that is not privileged: it writes under /etc, installs packages, manages units, and runs containers. A host that dropped privilege could apply almost nothing, and the almost is where the confusion would live.

A service rather than a command, because it holds the link, and something must survive a reboot to hold it. The command-line entry points remain — they are how a person inspects and rescues a machine — but the ordinary case is a unit that is always up.

It cannot run in a container, and the reason is the bootstrap

Worth stating because everything else the mesh runs is a container, which makes the host look like an exception somebody forgot to fix.

Step 0 of the substrate bootstrap is installing the container runtime (07-the-substrate.md). A host that ran inside a container could not perform it — it would need the thing it is there to install. On a machine with no runtime, nothing would ever start.

That is not the only reason, but it is the sufficient one:

  • It would break ADR 0041. Copy it onto a machine and run it stops being true when the machine must already have a container runtime.
  • The isolation would be fiction. To write /etc, install packages, manage units and run containers, it would need the host's mount, PID and network namespaces plus the runtime's own socket. A container with all of those is a process with extra steps.

So the host is a plain process on the machine, and everything above tier 0 is a container. That split is the tier boundary made concrete rather than an inconsistency.

What it needs from an init, and why that is not a dependency

The host needs four things from whatever supervises it: start at boot, restart when it exits, give up after repeated failures, and run something else when it gives up.

Every machine the mesh targets already has systemd, and the host already treats the service manager as a detected capability rather than an assumption. This is not a dependency in ADR 0041's sense — 0041 is about what must be installed before the host works, and an init is not installed, it is what the machine already is.

Abstracting over init systems is not done, because there is no second one to abstract over. The unit file is the only systemd-specific artefact, it belongs to the package rather than the binary, and a machine with a different supervisor would ship a different package — which is where that difference belongs.

It never manages its own unit

The host's own service file is not a resource the host applies. The temptation is obvious — it manages units, and its own unit is a unit — and it ends with a host stopping itself half way through an apply, leaving a machine in a state nothing is running to fix.

So the boundary is: the installation owns the host; the host owns everything else. A declaration that names the host's own unit is refused rather than obeyed.

It is installed as a package, and a tarball is the floor

Two mechanisms, and the second is not a fallback for the first failing — it is what makes the first possible.

package — pacman -S nox-mesh-host the ordinary path. Carries the binary, the unit file, the state directory, and an upgrade path
tarball — curl … | tar -xz the floor. One static binary, no repository, no distribution assumed

Why a package rather than only a binary. ADR 0041 says copying the binary onto a machine is the whole installation, and that remains true of the binary. But a unit file, a state directory and an upgrade path are real, and something has to own them. A package that installs one statically linked binary plus a unit file adds no runtime dependency — 0041 is about what must already be present for the host to work, not about how the bytes arrived.

Why the tarball must keep working. The package lives in a repository, and the mesh's own repository is hosted on the mesh. A first node cannot fetch from a mesh that does not exist yet, and neither can a node whose mesh is down — which is exactly when somebody is trying to fix it. Any path that requires the mesh to install the thing that joins the mesh is a circle, so the tarball is the path that is never allowed to acquire a dependency.

The host may replace its own binary; it may not stop its own unit

The first draft of this record said the mesh must not upgrade the host at all. That was too broad, and it conflated two different acts.

Replacing the binary is safe. Unix keeps the running executable's inode open, so a package upgrade writes a new file and the running process continues on the old one, undisturbed. Stopping the unit is what is unsafe — that is the host killing itself part-way through an apply, leaving a machine with nothing running to finish or fix it.

So the host may apply a package naming itself. What it must never do is ask the service manager to restart it.

The restart happens by exiting, not by asking. When the host notices its own executable has been replaced, it finishes the apply it is in, reports what it did, and exits cleanly. The supervisor's Restart=always starts it again, on the new binary. Nothing stops the host; the host stops, having finished.

Three conditions, and they are the whole safety argument:

  • after the apply completes and its outcomes are recorded — never mid-way;
  • only when the executable actually changed, which Linux reports plainly: a replaced /proc/self/exe reads as the old path marked deleted;
  • exit zero, so a restart is what a supervisor does next rather than a failure it backs off from.

This makes a fleet-wide host upgrade an ordinary declaration, which the first draft gave up.

Changes are pushed. The timer is for drift, and only for drift

The host does not poll for work. A new declaration arrives as a message on the link, and the host applies it then (ADR 0001). Polling for updates over a connection that already exists would be strictly worse in both directions: slower to land, and constant traffic to learn nothing.

Four triggers, and only one of them is a clock:

Trigger Kind Why
a declaration arrives pushed the ordinary path — this is how changes land
start event the machine may have changed while nothing was running
reconnect event declarations may have been missed while disconnected
a timer periodic drift, and nothing else

The timer cannot be replaced by an event, and the reason is definitional. Drift is change the mesh did not make — somebody edited a managed file, a distribution upgrade replaced a config, a container was stopped by hand. Nothing will ever send a message about it, because the thing that did it is not part of the mesh. A local periodic check is the only way to see it at all.

Without it, owned reports what the host applied rather than what is there, which is ADR 0035 violated by omission.

Ten minutes, configurable. The check is cheap: it asks the package database, the service manager and the container runtime about resources the host already knows it owns.

It reports upward on a heartbeat

Separate from reconciling, and easy to conflate with it: the node tells the mesh what it is — its running version, what it holds, what it last applied — on link, after every apply, and periodically while idle.

The heartbeat is what makes silence mean something. Without it, the mesh cannot distinguish a node that is fine and has had nothing to do from one that stopped. With it, last heard from is a fact per node, and ADR 0059 is what handles the case where the node cannot report at all.

Consequences

  • Adoption becomes two concrete steps, which is the point of writing this down: install the package, then hand it a token. Nothing else.
  • The host gains a mode it does not have, and it is the larger half of stage 3. Today every entry point runs and exits.
  • Refusing to manage its own unit needs enforcing, not just stating. A declaration naming the host's unit must be refused by name, and that refusal is a test.
  • A host that cannot reach the mesh keeps reconciling from its store, which is ADR 0036 made operational rather than aspirational: a disconnected node is not merely tolerated, it is actively holding its machine in the last state it was told to hold.
  • The timer makes drift visible and also makes it loud. A resource the host cannot apply will now fail every ten minutes rather than once. That is correct and it needs somewhere to go other than a log nobody reads — which is observability's, and it does not exist yet.
  • A host upgrade is an ordinary declaration, which is worth the care it needs: the exit path is the only place the host deliberately stops, and a bug there is a node that restarts in a loop or never comes back. It wants a test that the host does not exit when its binary is unchanged, as much as one that it does when it changed.
  • A version-skewed fleet is now normal and needs saying. Nodes restart onto the new binary at whatever moment their apply finishes, so "the fleet is upgraded" is a range rather than an instant. What a node reports as its version must be the running one, not the installed one, or the mesh will believe an upgrade landed before it took effect.

References

  • ADR 0041 — the property the package must not break.
  • ADR 0036 — what a node is, which this makes operational.
  • ADR 0039 — the link the process exists to hold.
  • ADR 0051 — the second of the two adoption steps.