Decided after measuring what renumbering actually costs: 96 references in code comments across two repositories, none of which would have failed to compile. They would have pointed at the wrong reasoning, which is worse than a broken link because nothing reports it. So a number identifies a record and never changes. It cannot also be a position -- a position moves when the set changes, and an identity that moves is not one. The reading order moves into an index generated from each record's `topic:`. Six topics, in the order somebody learns the system. The index is WRITTEN rather than only generated on demand, which reverses what this repository previously said. The reason it said otherwise is that a hand-written index drifts -- but a reader looking at the folder on a forge sees the folder, not a command, and the drift objection is answered by checking rather than by refusing to write one. That is §5's own rule: a rule states how it is checked. Two checks, both confirmed to bite. index.py fails when the written order no longer matches the records. records.py fails when a record has no topic or one nobody defined -- the quiet failure being a record that vanishes from the order rather than appearing in the wrong place.
8.2 KiB
topic, status, date, deciders, reconstructed
| topic | status | date | deciders | reconstructed |
|---|---|---|---|---|
| the tiers | accepted | 2026-08-28 | jochen | false |
5. The node host
Consolidated 2026-08-28 from eight records. Tier 0 is one component and was decided over a week; the reasoning is kept, the fragmentation is not.
It applies; it does not decide
The host makes a machine match what it was told, and never works out what that should be.
This is the line the whole tier rests on, and it is not about privilege — it is about what a single machine can know. Deciding needs knowledge the machine does not have: which nodes should run a store, which peers belong in an overlay, whether a node has been unreachable for a week. Anything needing a second node is the control plane's.
The practical form: the host never queries the mesh's database and holds no credential to it. Two modules in the current mesh do, and they are the reason every node permanently carries a database credential.
It depends on nothing that must be installed first
A statically linked binary. Copy it onto a machine and run it — that is the whole installation. Written in Go, because the job is system-level and because a runtime that must be installed first would make the host depend on the thing it exists to install.
What it needs from the machine is not a dependency in this sense. An init is not installed; it is what the machine already is. A package manager is the distribution. Those are what a machine is, not what must be put on it before the host works.
It is built per operating system
systemd and pacman are the Arch host's implementation, not abstractions the mesh grows.
They are not independent choices: a machine has pacman because it is Arch, and the package
manager, service manager and packaging format arrive together as one decision somebody made at
install time.
mesh-host-arch pacman · systemctl · a container runtime
mesh-host-alpine apk · rc-service
mesh-host-android neither — a partial host
Abstracting them was rejected on correctness, not effort. The service applier reads systemd's
LoadState to tell not installed apart from stopped — which is what stops it reporting
absence as success — and OpenRC has no equivalent. An interface spanning both must drop it, and
the lowest common denominator is exactly where that fault lives.
Almost all of it is shared. The declaration vocabulary, the store, the apply loop, the read-back discipline, the refusal model and the link are portable. Two appliers differ.
A host that cannot implement a shape refuses it. Android has no package manager it may drive
and no init it may register with, so it implements file, directory and action and refuses
the rest — the same refusal an unknown type gets, with a different reason. Those three are the
portable floor, and they are what makes a partial host a real thing rather than a broken one.
It is a root service, and it never manages its own unit
Root, because no useful part of the job is unprivileged: it writes under /etc, installs
packages, manages units and runs containers.
It cannot run in a container, and the reason is decisive rather than stylistic: installing the container runtime is a step of the bootstrap, so a host inside a container would need the thing it exists to install. Everything above tier 0 is a container; the host is not. That split is the tier boundary made concrete.
The installation owns the host; the host owns everything else. It manages service resources
and its own unit is one — the temptation is obvious and it ends with a host stopping itself half
way through an apply, leaving a machine with nothing running to fix it.
An init is asked for one thing
Start this at boot. That is all, and every init can express it — systemd, OpenRC, runit, s6.
Everything else is a launcher the host ships, which supervises it: restart it when it exits, count consecutive failures, roll back after too many, halt after that. Policy in a unit file can only be read and hoped for; a script with a counter can be tested, and this is the one piece that must work on a machine where the host does not.
The launcher does not exec the host, it supervises it — so restarting is ours rather than the init's. The cost is signals: a supervisor that exits while its child runs leaves the host to be killed rather than to stop, and an apply interrupted that way is the half-configured machine this design is about. So it traps the shutdown signal, passes it down, and waits.
A clean exit is the upgrade path, and it is the easiest thing to get wrong — twice now. The host stands aside for a new binary by exiting zero, so anything supervising must restart on a zero exit and must not count it as a failure.
Recovery is local, and detection is the mesh's. Nothing dials a node and a host that cannot start cannot report, so the node must recover itself. But a local supervisor sees one process failing and cannot tell a broken machine from a broken release — only something watching every node can, which is why a host rollout is staged and stops when nodes go quiet.
A host may be episodic
Resident or episodic, and both are hosts. A phone has no init to register with and nothing worth supervising, because a supervisor would be killed alongside what it supervises. So it runs when the platform allows and is killed when the platform wants the memory — and that is disconnection, which is already an ordinary situation.
It needs no keep-alive and no new mechanism: the store is already authoritative while disconnected, reconcile already happens on start, and last heard from is already reported rather than alarmed on. An episodic host cannot be the first node, because every bootstrap step is a shape it refuses.
What a declaration is
An ordered list of resources the host owns. JSON, because Go's standard library carries a JSON parser and no YAML, and the one binary whose argument is that it needs nothing must not gain a parser to buy authoring comfort in a machine-written document.
Ordered, because ordering is a decision. The host does not sort and does not resolve dependencies — that would be deciding, and deciding the thing most likely to differ between what the control plane intended and what the machine does.
Every resource has a stable identity — a name the control plane keeps across declarations, not a position and not a hash of content. It is what lets the store say this is the same resource I applied last time, which is what makes removal possible at all.
Unknown is refused, whole. A field the host does not know is something the control plane believes it asked for. A declaration naming one is rejected entirely, naming every problem at once — a host that applied the parts it understood would leave a machine that looks configured and is not.
Six shapes: file, directory, service, package, container, action. Every addition
widens what a compromised control plane can express, so the list is a security artefact and grows
deliberately.
The bundle may carry actions; the link may not
An action runs a command, and the host never learns what it means. It is needed because the
bootstrap creates a database before there is any mesh to ask for one, and the host must not learn
what a database is.
Permitted from the bundle, refused from the link, and the asymmetry is the whole point: a bundle arrives with the binary, so anyone able to put a hostile action there could have put it in the host itself — refusing it buys nothing and costs the bootstrap. The link is a separate party, reachable separately, and an action there is an unbounded blast radius.
An action must carry its own verification, which is also its idempotency check. The host does not know what a database is, so is it already there is a question only the declaration can ask.
Consequences
- The migration is smaller than it looks. A joining node never needs mesh-wide state — it needs an identity, an address and one peer, and the rest arrives as declarations.
- What is applied is recorded after it works, never before. A failed apply leaves the machine in whatever state it reached, and nothing must claim otherwise.
- A second operating system is additive: two appliers and a four-line init file.