The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18, 19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only the archaeology of what used to be there. Renumbered contiguously. Renames run in ascending order, so every target number is already free and no two files ever collide. The reference rewrite is one simultaneous pass rather than a sequence of replacements. Numbers moved into slots other numbers were vacating -- the node host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time would have cascaded and silently pointed things at the wrong record. Seven plain-text references survived the merges as prose rather than links, naming records that no longer existed: the enrolment token, the link boundary, what a declaration is, reachability, the repository structure. Each mapped to the consolidated record that now holds it. Verified rather than assumed: every [ADR NNNN](path) link now has matching text and target, checked across the whole repository, and the checker passes. Frontmatter `consolidates:` lists dropped -- they named records that are gone, and each consolidated record already says in prose what it absorbed.
8.2 KiB
status, date, deciders, reconstructed
| status | date | deciders | reconstructed |
|---|---|---|---|
| accepted | 2026-08-28 | jochen | false |
16. The node host
Consolidated 2026-08-28 from eight records. Tier 0 is one component and was decided over a week; the reasoning is kept, the fragmentation is not.
It applies; it does not decide
The host makes a machine match what it was told, and never works out what that should be.
This is the line the whole tier rests on, and it is not about privilege — it is about what a single machine can know. Deciding needs knowledge the machine does not have: which nodes should run a store, which peers belong in an overlay, whether a node has been unreachable for a week. Anything needing a second node is the control plane's.
The practical form: the host never queries the mesh's database and holds no credential to it. Two modules in the current mesh do, and they are the reason every node permanently carries a database credential.
It depends on nothing that must be installed first
A statically linked binary. Copy it onto a machine and run it — that is the whole installation. Written in Go, because the job is system-level and because a runtime that must be installed first would make the host depend on the thing it exists to install.
What it needs from the machine is not a dependency in this sense. An init is not installed; it is what the machine already is. A package manager is the distribution. Those are what a machine is, not what must be put on it before the host works.
It is built per operating system
systemd and pacman are the Arch host's implementation, not abstractions the mesh grows.
They are not independent choices: a machine has pacman because it is Arch, and the package
manager, service manager and packaging format arrive together as one decision somebody made at
install time.
mesh-host-arch pacman · systemctl · a container runtime
mesh-host-alpine apk · rc-service
mesh-host-android neither — a partial host
Abstracting them was rejected on correctness, not effort. The service applier reads systemd's
LoadState to tell not installed apart from stopped — which is what stops it reporting
absence as success — and OpenRC has no equivalent. An interface spanning both must drop it, and
the lowest common denominator is exactly where that fault lives.
Almost all of it is shared. The declaration vocabulary, the store, the apply loop, the read-back discipline, the refusal model and the link are portable. Two appliers differ.
A host that cannot implement a shape refuses it. Android has no package manager it may drive
and no init it may register with, so it implements file, directory and action and refuses
the rest — the same refusal an unknown type gets, with a different reason. Those three are the
portable floor, and they are what makes a partial host a real thing rather than a broken one.
It is a root service, and it never manages its own unit
Root, because no useful part of the job is unprivileged: it writes under /etc, installs
packages, manages units and runs containers.
It cannot run in a container, and the reason is decisive rather than stylistic: installing the container runtime is a step of the bootstrap, so a host inside a container would need the thing it exists to install. Everything above tier 0 is a container; the host is not. That split is the tier boundary made concrete.
The installation owns the host; the host owns everything else. It manages service resources
and its own unit is one — the temptation is obvious and it ends with a host stopping itself half
way through an apply, leaving a machine with nothing running to fix it.
An init is asked for one thing
Start this at boot. That is all, and every init can express it — systemd, OpenRC, runit, s6.
Everything else is a launcher the host ships, which supervises it: restart it when it exits, count consecutive failures, roll back after too many, halt after that. Policy in a unit file can only be read and hoped for; a script with a counter can be tested, and this is the one piece that must work on a machine where the host does not.
The launcher does not exec the host, it supervises it — so restarting is ours rather than the init's. The cost is signals: a supervisor that exits while its child runs leaves the host to be killed rather than to stop, and an apply interrupted that way is the half-configured machine this design is about. So it traps the shutdown signal, passes it down, and waits.
A clean exit is the upgrade path, and it is the easiest thing to get wrong — twice now. The host stands aside for a new binary by exiting zero, so anything supervising must restart on a zero exit and must not count it as a failure.
Recovery is local, and detection is the mesh's. Nothing dials a node and a host that cannot start cannot report, so the node must recover itself. But a local supervisor sees one process failing and cannot tell a broken machine from a broken release — only something watching every node can, which is why a host rollout is staged and stops when nodes go quiet.
A host may be episodic
Resident or episodic, and both are hosts. A phone has no init to register with and nothing worth supervising, because a supervisor would be killed alongside what it supervises. So it runs when the platform allows and is killed when the platform wants the memory — and that is disconnection, which is already an ordinary situation.
It needs no keep-alive and no new mechanism: the store is already authoritative while disconnected, reconcile already happens on start, and last heard from is already reported rather than alarmed on. An episodic host cannot be the first node, because every bootstrap step is a shape it refuses.
What a declaration is
An ordered list of resources the host owns. JSON, because Go's standard library carries a JSON parser and no YAML, and the one binary whose argument is that it needs nothing must not gain a parser to buy authoring comfort in a machine-written document.
Ordered, because ordering is a decision. The host does not sort and does not resolve dependencies — that would be deciding, and deciding the thing most likely to differ between what the control plane intended and what the machine does.
Every resource has a stable identity — a name the control plane keeps across declarations, not a position and not a hash of content. It is what lets the store say this is the same resource I applied last time, which is what makes removal possible at all.
Unknown is refused, whole. A field the host does not know is something the control plane believes it asked for. A declaration naming one is rejected entirely, naming every problem at once — a host that applied the parts it understood would leave a machine that looks configured and is not.
Six shapes: file, directory, service, package, container, action. Every addition
widens what a compromised control plane can express, so the list is a security artefact and grows
deliberately.
The bundle may carry actions; the link may not
An action runs a command, and the host never learns what it means. It is needed because the
bootstrap creates a database before there is any mesh to ask for one, and the host must not learn
what a database is.
Permitted from the bundle, refused from the link, and the asymmetry is the whole point: a bundle arrives with the binary, so anyone able to put a hostile action there could have put it in the host itself — refusing it buys nothing and costs the bootstrap. The link is a separate party, reachable separately, and an action there is an unbounded blast radius.
An action must carry its own verification, which is also its idempotency check. The host does not know what a database is, so is it already there is a question only the declaration can ask.
Consequences
- The migration is smaller than it looks. A joining node never needs mesh-wide state — it needs an identity, an address and one peer, and the rest arrives as declarations.
- What is applied is recorded after it works, never before. A failed apply leaves the machine in whatever state it reached, and nothing must claim otherwise.
- A second operating system is additive: two appliers and a four-line init file.