Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.
the node host 8 -> 1 applies not decides, depends on nothing,
per operating system, root service, the
launcher, episodic, what a declaration is,
actions from the bundle only
a node and how it joins 4 -> 1 what a node is, joining, the link as
security boundary, the enrolment token
modules and the graph 7 -> 1 everything is a module, no domain modules,
three edges, provisioning, the core library
substrate and control 6 -> 1 the test, seven contexts, one control plane,
plane the authority is not a database, the named
products, the pinned bundle
connectivity 3 -> 1 a route is a grant, reachability declared,
filter rules
delivery 5 -> 1 reconciliation not a pipeline, artifacts,
the three silos, a failed step, the verdict
the lab 5 -> 1 (earlier)
how this repository 10 -> 1 (earlier)
works
Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.
The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.
The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
8.3 KiB
status, date, deciders, reconstructed, consolidates
| status | date | deciders | reconstructed | consolidates | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| accepted | 2026-08-28 | jochen | false |
|
37. The node host
Consolidated 2026-08-28 from eight records. Tier 0 is one component and was decided over a week; the reasoning is kept, the fragmentation is not.
It applies; it does not decide
The host makes a machine match what it was told, and never works out what that should be.
This is the line the whole tier rests on, and it is not about privilege — it is about what a single machine can know. Deciding needs knowledge the machine does not have: which nodes should run a store, which peers belong in an overlay, whether a node has been unreachable for a week. Anything needing a second node is the control plane's.
The practical form: the host never queries the mesh's database and holds no credential to it. Two modules in the current mesh do, and they are the reason every node permanently carries a database credential.
It depends on nothing that must be installed first
A statically linked binary. Copy it onto a machine and run it — that is the whole installation. Written in Go, because the job is system-level and because a runtime that must be installed first would make the host depend on the thing it exists to install.
What it needs from the machine is not a dependency in this sense. An init is not installed; it is what the machine already is. A package manager is the distribution. Those are what a machine is, not what must be put on it before the host works.
It is built per operating system
systemd and pacman are the Arch host's implementation, not abstractions the mesh grows.
They are not independent choices: a machine has pacman because it is Arch, and the package
manager, service manager and packaging format arrive together as one decision somebody made at
install time.
mesh-host-arch pacman · systemctl · a container runtime
mesh-host-alpine apk · rc-service
mesh-host-android neither — a partial host
Abstracting them was rejected on correctness, not effort. The service applier reads systemd's
LoadState to tell not installed apart from stopped — which is what stops it reporting
absence as success — and OpenRC has no equivalent. An interface spanning both must drop it, and
the lowest common denominator is exactly where that fault lives.
Almost all of it is shared. The declaration vocabulary, the store, the apply loop, the read-back discipline, the refusal model and the link are portable. Two appliers differ.
A host that cannot implement a shape refuses it. Android has no package manager it may drive
and no init it may register with, so it implements file, directory and action and refuses
the rest — the same refusal an unknown type gets, with a different reason. Those three are the
portable floor, and they are what makes a partial host a real thing rather than a broken one.
It is a root service, and it never manages its own unit
Root, because no useful part of the job is unprivileged: it writes under /etc, installs
packages, manages units and runs containers.
It cannot run in a container, and the reason is decisive rather than stylistic: installing the container runtime is a step of the bootstrap, so a host inside a container would need the thing it exists to install. Everything above tier 0 is a container; the host is not. That split is the tier boundary made concrete.
The installation owns the host; the host owns everything else. It manages service resources
and its own unit is one — the temptation is obvious and it ends with a host stopping itself half
way through an apply, leaving a machine with nothing running to fix it.
An init is asked for one thing
Start this at boot. That is all, and every init can express it — systemd, OpenRC, runit, s6.
Everything else is a launcher the host ships, which supervises it: restart it when it exits, count consecutive failures, roll back after too many, halt after that. Policy in a unit file can only be read and hoped for; a script with a counter can be tested, and this is the one piece that must work on a machine where the host does not.
The launcher does not exec the host, it supervises it — so restarting is ours rather than the init's. The cost is signals: a supervisor that exits while its child runs leaves the host to be killed rather than to stop, and an apply interrupted that way is the half-configured machine this design is about. So it traps the shutdown signal, passes it down, and waits.
A clean exit is the upgrade path, and it is the easiest thing to get wrong — twice now. The host stands aside for a new binary by exiting zero, so anything supervising must restart on a zero exit and must not count it as a failure.
Recovery is local, and detection is the mesh's. Nothing dials a node and a host that cannot start cannot report, so the node must recover itself. But a local supervisor sees one process failing and cannot tell a broken machine from a broken release — only something watching every node can, which is why a host rollout is staged and stops when nodes go quiet.
A host may be episodic
Resident or episodic, and both are hosts. A phone has no init to register with and nothing worth supervising, because a supervisor would be killed alongside what it supervises. So it runs when the platform allows and is killed when the platform wants the memory — and that is disconnection, which is already an ordinary situation.
It needs no keep-alive and no new mechanism: the store is already authoritative while disconnected, reconcile already happens on start, and last heard from is already reported rather than alarmed on. An episodic host cannot be the first node, because every bootstrap step is a shape it refuses.
What a declaration is
An ordered list of resources the host owns. JSON, because Go's standard library carries a JSON parser and no YAML, and the one binary whose argument is that it needs nothing must not gain a parser to buy authoring comfort in a machine-written document.
Ordered, because ordering is a decision. The host does not sort and does not resolve dependencies — that would be deciding, and deciding the thing most likely to differ between what the control plane intended and what the machine does.
Every resource has a stable identity — a name the control plane keeps across declarations, not a position and not a hash of content. It is what lets the store say this is the same resource I applied last time, which is what makes removal possible at all.
Unknown is refused, whole. A field the host does not know is something the control plane believes it asked for. A declaration naming one is rejected entirely, naming every problem at once — a host that applied the parts it understood would leave a machine that looks configured and is not.
Six shapes: file, directory, service, package, container, action. Every addition
widens what a compromised control plane can express, so the list is a security artefact and grows
deliberately.
The bundle may carry actions; the link may not
An action runs a command, and the host never learns what it means. It is needed because the
bootstrap creates a database before there is any mesh to ask for one, and the host must not learn
what a database is.
Permitted from the bundle, refused from the link, and the asymmetry is the whole point: a bundle arrives with the binary, so anyone able to put a hostile action there could have put it in the host itself — refusing it buys nothing and costs the bootstrap. The link is a separate party, reachable separately, and an action there is an unbounded blast radius.
An action must carry its own verification, which is also its idempotency check. The host does not know what a database is, so is it already there is a question only the declaration can ask.
Consequences
- The migration is smaller than it looks. A joining node never needs mesh-wide state — it needs an identity, an address and one peer, and the rest arrives as declarations.
- What is applied is recorded after it works, never before. A failed apply leaves the machine in whatever state it reached, and nothing must claim otherwise.
- A second operating system is additive: two appliers and a four-line init file.