The host was described as a component and never as something that runs for years on a machine somebody else also uses. 09 covers every state a machine can be in and every transition between them. Four states: unmanaged, hosted, enrolled, disconnected. Only the last two are nodes, and they are the same node in two situations. `hosted` -- the host installed but never told which mesh it belongs to -- had no name before and is where a machine sits between the two adoption commands. Things that were unclear and now are not: The first node walks the same path in an unusual order: reconcile from the bundle, the control plane it just raised issues a token, enrol against it. Its specialness lasts two commands. A side effect worth having -- enrolment is exercised on node one, rather than being written and first used on node two. Enrolment reports profile and inventory BEFORE the control plane decides anything. The profile is the input to that decision, not a diagnostic; the control plane cannot decide what a machine should run without knowing what it can run. Rebooting mid-apply is safe by construction. The store records each resource after it worked, so a host that dies half way through comes back and applies the rest. The rule that stops the host lying about what it did also makes it crash-safe. Retiring splits in two. Graceful is a final empty declaration. A node that is gone will reconcile its last declaration forever -- the honest consequence of making disconnection ordinary. The answer is not to make the host expire but that the node holds nothing that outlives revocation: every grant is a per-node credential revoked at the provider. A lost node keeps running and stops being able to reach anything. Said plainly rather than implying the mesh can switch a machine off, which it cannot and should not. Losing the store is quiet and permanent, so it gets its own section. The host re-enrols and re-applies fine; what does not come back is removal, because resources it no longer has a record of become unowned and sit there indefinitely. Also corrects 0057, which said the mesh must not upgrade the host at all. That conflated two acts. Replacing the binary is safe -- Unix keeps the running inode. Stopping the unit is not. So the host may apply a package naming itself, and restarts by finishing its apply and exiting cleanly, letting the supervisor start it on the new binary. It never asks the service manager to restart it. That makes a fleet-wide host upgrade an ordinary declaration, which the first draft gave up on. 0057 remains proposed.
152 lines
8.2 KiB
Markdown
152 lines
8.2 KiB
Markdown
---
|
|
status: proposed
|
|
date: 2026-08-27
|
|
deciders: jochen
|
|
reconstructed: false
|
|
extends: 0041-the-host-depends-on-nothing.md
|
|
---
|
|
|
|
# 57. The host is a root service, installed as a package, that never manages itself
|
|
|
|
## Context
|
|
|
|
[`05-the-node-host.md`](../03-DESIGN/01-to-be/05-the-node-host.md) describes what the host *does*
|
|
and never says what it *is* at runtime. Searched: the words *daemon*, *long-running*, *interval*,
|
|
*poll* and *heartbeat* appear in none of it, nor in
|
|
[ADR 0038](0038-a-node-joins-by-linking-first.md) or
|
|
[ADR 0039](0039-the-link-is-the-security-boundary.md).
|
|
|
|
What exists today is a command that runs and exits — `mesh-host apply FILE`. What the design
|
|
requires is a process holding an outbound link to the control plane. Nobody wrote down that
|
|
these are different things, so several questions have no answer: does it reconcile on a timer,
|
|
what happens when the link drops, and who installs the unit that starts it — given that the host
|
|
is the thing that installs units.
|
|
|
|
## Decision
|
|
|
|
### It runs on every node, and that is what a node is
|
|
|
|
[ADR 0036](0036-a-node-is-a-managed-machine.md) defines a node as a managed machine. **The host
|
|
is what makes it managed**, so a machine without one is not a node with a missing component; it
|
|
is not a node. There is no partial mode, no agentless node, and no second way in.
|
|
|
|
### It is a root service
|
|
|
|
**Root**, because there is no useful subset of its job that is not privileged: it writes under
|
|
`/etc`, installs packages, manages units, and runs containers. A host that dropped privilege
|
|
could apply almost nothing, and the almost is where the confusion would live.
|
|
|
|
**A service rather than a command**, because it holds the link, and something must survive a
|
|
reboot to hold it. The command-line entry points remain — they are how a person inspects and
|
|
rescues a machine — but the ordinary case is a unit that is always up.
|
|
|
|
### It never manages its own unit
|
|
|
|
**The host's own service file is not a resource the host applies.** The temptation is obvious —
|
|
it manages units, and its own unit is a unit — and it ends with a host stopping itself half way
|
|
through an apply, leaving a machine in a state nothing is running to fix.
|
|
|
|
So the boundary is: **the installation owns the host; the host owns everything else.** A
|
|
declaration that names the host's own unit is refused rather than obeyed.
|
|
|
|
### It is installed as a package, and a tarball is the floor
|
|
|
|
Two mechanisms, and the second is not a fallback for the first failing — it is what makes the
|
|
first possible.
|
|
|
|
| | |
|
|
|---|---|
|
|
| **package** — `pacman -S nox-mesh-host` | the ordinary path. Carries the binary, the unit file, the state directory, and an upgrade path |
|
|
| **tarball** — `curl … \| tar -xz` | the floor. One static binary, no repository, no distribution assumed |
|
|
|
|
**Why a package rather than only a binary.** [ADR 0041](0041-the-host-depends-on-nothing.md) says
|
|
copying the binary onto a machine is the whole installation, and that remains true of the
|
|
*binary*. But a unit file, a state directory and an upgrade path are real, and something has to
|
|
own them. A package that installs one statically linked binary plus a unit file adds no runtime
|
|
dependency — 0041 is about what must already be present for the host to work, not about how the
|
|
bytes arrived.
|
|
|
|
**Why the tarball must keep working.** The package lives in a repository, and the mesh's own
|
|
repository is hosted on the mesh. A first node cannot fetch from a mesh that does not exist yet,
|
|
and neither can a node whose mesh is down — which is exactly when somebody is trying to fix it.
|
|
**Any path that requires the mesh to install the thing that joins the mesh is a circle**, so the
|
|
tarball is the path that is never allowed to acquire a dependency.
|
|
|
|
### The host may replace its own binary; it may not stop its own unit
|
|
|
|
The first draft of this record said the mesh must not upgrade the host at all. That was too
|
|
broad, and it conflated two different acts.
|
|
|
|
**Replacing the binary is safe.** Unix keeps the running executable's inode open, so a package
|
|
upgrade writes a new file and the running process continues on the old one, undisturbed.
|
|
**Stopping the unit is what is unsafe** — that is the host killing itself part-way through an
|
|
apply, leaving a machine with nothing running to finish or fix it.
|
|
|
|
So the host may apply a `package` naming itself. What it must never do is ask the service
|
|
manager to restart it.
|
|
|
|
**The restart happens by exiting, not by asking.** When the host notices its own executable has
|
|
been replaced, it finishes the apply it is in, reports what it did, and **exits cleanly**. The
|
|
supervisor's `Restart=always` starts it again, on the new binary. Nothing stops the host; the
|
|
host stops, having finished.
|
|
|
|
Three conditions, and they are the whole safety argument:
|
|
|
|
- **after** the apply completes and its outcomes are recorded — never mid-way;
|
|
- **only** when the executable actually changed, which Linux reports plainly: a replaced
|
|
`/proc/self/exe` reads as the old path marked deleted;
|
|
- **exit zero**, so a restart is what a supervisor does next rather than a failure it backs off
|
|
from.
|
|
|
|
This makes a fleet-wide host upgrade an ordinary declaration, which the first draft gave up.
|
|
|
|
### It reconciles on four triggers
|
|
|
|
| Trigger | Why |
|
|
|---|---|
|
|
| **start** | the machine may have changed while nothing was running |
|
|
| **a declaration arrives** | the ordinary path |
|
|
| **a timer** | drift. Something other than the host changed the machine — a person, a package upgrade replacing a config file |
|
|
| **reconnect** | it may have missed declarations while disconnected ([ADR 0036](0036-a-node-is-a-managed-machine.md)) |
|
|
|
|
**The timer is what makes the store's claim true.** Without it a machine that drifted stays
|
|
drifted until somebody changes a declaration, and `owned` reports what the host *applied* rather
|
|
than what is *there* — which is
|
|
[ADR 0035](0035-a-picture-is-read-from-what-runs.md) violated by omission.
|
|
|
|
**Ten minutes**, configurable. Short enough that drift is bounded by something a person would
|
|
notice anyway, long enough that a fleet is not doing constant work. The reconcile is cheap: it
|
|
asks the package database, the service manager and the container runtime about resources the
|
|
host already knows it owns.
|
|
|
|
## Consequences
|
|
|
|
- **Adoption becomes two concrete steps**, which is the point of writing this down:
|
|
install the package, then hand it a token. Nothing else.
|
|
- **The host gains a mode it does not have**, and it is the larger half of stage 3. Today every
|
|
entry point runs and exits.
|
|
- **Refusing to manage its own unit needs enforcing, not just stating.** A declaration naming
|
|
the host's unit must be refused by name, and that refusal is a test.
|
|
- **A host that cannot reach the mesh keeps reconciling from its store**, which is
|
|
[ADR 0036](0036-a-node-is-a-managed-machine.md) made operational rather than aspirational: a
|
|
disconnected node is not merely tolerated, it is actively holding its machine in the last
|
|
state it was told to hold.
|
|
- **The timer makes drift visible and also makes it loud.** A resource the host cannot apply
|
|
will now fail every ten minutes rather than once. That is correct and it needs somewhere to go
|
|
other than a log nobody reads — which is `observability`'s, and it does not exist yet.
|
|
- **A host upgrade is an ordinary declaration**, which is worth the care it needs: the exit
|
|
path is the only place the host deliberately stops, and a bug there is a node that restarts in
|
|
a loop or never comes back. It wants a test that the host does **not** exit when its binary is
|
|
unchanged, as much as one that it does when it changed.
|
|
- **A version-skewed fleet is now normal and needs saying.** Nodes restart onto the new binary
|
|
at whatever moment their apply finishes, so "the fleet is upgraded" is a range rather than an
|
|
instant. What a node reports as its version must be the **running** one, not the installed
|
|
one, or the mesh will believe an upgrade landed before it took effect.
|
|
|
|
## References
|
|
|
|
- [ADR 0041](0041-the-host-depends-on-nothing.md) — the property the package must not break.
|
|
- [ADR 0036](0036-a-node-is-a-managed-machine.md) — what a node is, which this makes operational.
|
|
- [ADR 0039](0039-the-link-is-the-security-boundary.md) — the link the process exists to hold.
|
|
- [ADR 0051](0051-the-enrolment-token-carries-the-mesh.md) — the second of the two adoption steps.
|