Commit Graph
11 Commits
Author SHA1 Message Date
jschoubben 10365f2eae Consolidate the design layer: one place per topic
Jochen: a jungle of specs that slightly contradict or patch each other, and
what matters is a working state rather than history. Both are fair and both are
mine.

Measured rather than assumed. 05-the-node-host and 09-the-node-lifecycle both
covered enrolment, the install commands, the unit file, the launcher and
reconcile -- I wrote 09 without taking anything out of 05, so the same things
were said twice and could drift apart.

Split by what each document IS. 05 is the component: what the host is, its
parts, the declaration vocabulary, the build order, how it is verified. 09 is
what happens to it: install, enrol, run, upgrade, retire. The whole "The
process" section left 05, and the unit file moved to 09 where installing is
described. 05 goes from 338 lines to 245 and now points at 09 rather than
restating it.

09 also carried a 105-line "Resolved" section -- six mechanisms framed as
"these were open and here is the answer". The content is needed; the framing is
history, and history is what makes a document read as a changelog rather than a
description. Renamed to what it actually is and the was-open phrasing removed.

Also added 10-delivery.md, which did not exist: four accepted decisions --
0054, 0063, 0064, 0065 -- had no design document at all, which is the specific
reason the delivery picture felt scattered. It is now one document covering
modules, the three edges, the core library, and how a change becomes a running
thing, with a table of what each property is designed against and what must
exist before it can be built.
2026-08-28 18:40:46 +02:00
jschoubben f1b1cd9aa0 Review: three ADRs no longer said what we had concluded
A sweep for claims overtaken by the last few days. Annotated rather than
rewritten, following the pattern already in 0049 -- what changed and why is the
useful part, and an accepted record should not quietly become something else.

0057's init section was wrong on all three of its claims. It said the host
needs FOUR things from an init; 0061 reduced that to one. It said every machine
the mesh targets already has systemd; Alpine does not, and it is the intended
first node. It said there is no second init to abstract over; there is now, and
the answer is still not an abstraction -- it is a four-line file per system.
What survives is the part that was always right: an init is not a dependency in
0041's sense, because it is not installed, it is what the machine already is.

0048 named Docker as the container runtime. It is now docker or podman,
detected rather than chosen -- because adoption keeps what a machine already
has, so naming one contradicted a rule already decided. That row is the only
one of the five that names two, and the record now says why.

0060 claimed the bundle is portable across operating systems. Its mechanism is;
its contents are not -- package names, unit names, service names all differ, so
an Arch host embeds an Arch bundle. That was my error, and it is the exact
confusion behind the question that found it.

The design layer had the same drift: 07 and 09 said "Docker" where they meant a
container runtime, 09 said systemd restarts the host after an upgrade when the
launcher does, and both install snippets assumed Arch. They now show Alpine and
Arch side by side, which makes the point better than prose did -- step 1
differs per system, step 2 never does.

Checked and NOT changed: 0047's "the vocabulary grows by one shape" is a claim
about the rate, not the count, and is still true. 0037 lists docker among tools
the host manages, which it does. 0041 says nothing about either.
2026-08-28 00:43:47 +02:00
jschoubben c557f99cba Record what testing podman actually showed
0060 said the container runtime was a separate decision. It is now made, and
the reasoning is worth keeping because it is the opposite answer to the same
question one paragraph earlier.

Abstracting service managers is lossy -- systemd and OpenRC are different
models and LoadState has no equivalent. Container runtimes converged on one CLI
deliberately, so almost nothing is lost: checked against podman 6.1.0, run,
rm -f and docker's own template syntax for state and labels all work unchanged.
Only the probe differs. So: a two-entry lookup, not an interface.

The difference that is NOT in the CLI is the one that would have shipped
silently. Podman accepts --restart unless-stopped, records it, and has no
daemon to act on it -- containers do not return after a reboot unless
podman-restart.service is enabled, which by default it is not. Every command
reports success and the effect does not happen.

That belongs in the declaration rather than the host: a node using podman is
told to enable the unit. Which is what made the service shape's missing 'boot'
field visible, and it is now built.
2026-08-27 23:59:05 +02:00
jschoubben e1ad39b500 Per-OS hosts, and an init asked for only start and restart
0060 -- the host is built per operating system. systemd and pacman are the Arch
host's implementation, not abstractions the mesh has to grow. They are not
independent choices: a machine has pacman because it is Arch, and the package
manager, service manager and packaging format arrive together as one decision
somebody made at install time.

Rejected abstracting them, and the reason is correctness rather than effort.
The service applier reads LoadState to tell "not installed" apart from
"stopped", which is what stops it reporting absence as success. An interface
spanning systemd and OpenRC degrades to what both express, and the lowest
common denominator is exactly where that fault lives.

Almost all of it is shared -- the vocabulary, store, apply loop, read-back
discipline, refusal model, bundle and link are portable. Two appliers differ.
And delivery was already per-OS, since a .pkg.tar.zst is an Arch artifact, so
this is the seam that already existed.

Android is the interesting case rather than Debian: no service manager, no
package installation, usually no root. Such a host implements file, directory
and action and refuses the rest -- the same refusal a host already gives an
unknown type, with a different reason. Those three are the portable floor.

The container runtime is deliberately left open: it is not an OS split, since
Arch runs docker or podman.

0061 -- the init is asked for start-at-boot and restart-on-exit, and nothing
else. Both are expressible in OpenRC, runit, s6 and an Android init.rc.
Counting failed starts and rolling back moves into a launcher, because that is
the one piece which must work when the host does not, and a script with a
counter can be tested where OnFailure= can only be hoped for. Supersedes 0059,
keeping its reasoning in full.

The checker found all six places citing 0059 and refused the commit until they
named the replacement.
2026-08-27 23:46:11 +02:00
jschoubben 19997d56c3 Approve 0057-0059, with four corrections from review
Not approved as drafted -- four things came out of checking them against each
other, and one was a bug that would have broken every upgrade.

The bug: 0059 specified Restart=on-failure while 0057 has the host restart onto
a new binary by exiting CLEANLY. on-failure does not restart a process that
exited zero, so every upgraded node would have been left stopped, having
successfully upgraded. Found by reading the two records against each other
rather than by either alone. Now Restart=always in all three places that
mention it.

The host cannot run in a container, and the reason is decisive rather than
stylistic: step 0 of the substrate bootstrap installs the container runtime, so
a host inside a container would need the thing it exists to install. It would
also break 0041 -- copy it onto a machine and run it stops being true when the
machine must already have a runtime. Everything above tier 0 is a container;
the host is not. That split is the tier boundary, not an inconsistency.

systemd is named rather than abstracted. An init is not a dependency in 0041's
sense: 0041 is about what must be installed before the host works, and an init
is not installed, it is what the machine already is. The unit file is the only
systemd-specific artefact and it belongs to the package, so a machine with a
different supervisor ships a different package.

The mesh is a watchdog, and my first draft was half an answer. Recovery must be
local -- nothing dials a node, and a host that cannot start cannot report. But
detection is the mesh's, and a local supervisor structurally cannot do it: it
sees one process failing and cannot tell a broken machine from a broken
release. Only something watching every node can, and that distinction decides
whether the response is "fix this machine" or "stop shipping this version". So
a host rollout is staged -- a few nodes, wait for heartbeats, continue or stop
on silence. Local rollback still needed, because the canary nodes break and
because a node offline during the rollout gets the declaration later with no
batch around it.

The first declaration is the overlay and nothing else. Forced, because a node's
address and peers are assigned rather than chosen. But also the way back in: a
node reachable over the overlay can be fixed by hand if a later declaration
breaks it, and a large first declaration risks a node that is broken and
unreachable at once.

Also stated plainly, because it reads as a contradiction: nodes reach each
other over the overlay and every node consumes from the broker; what 0039
forbids is an inbound CONTROL surface, not reachability.

And in 06: no node holds a credential to any control-plane store, for reads or
writes. Four ADRs already say this separately and none of them said it in one
place. Nodes state over the broker; the owning context writes. With a note that
most high-frequency writes are observability's, not the registry's -- routing
logs into the registry would be the shared-schema mistake arriving through a
door marked performance.
2026-08-27 22:12:12 +02:00
jschoubben 3ab11c96ef Say what the host process is: a root service, installed as a package
The design described what the host does and never what it is at runtime. The
words daemon, long-running, interval, poll and heartbeat appeared nowhere in it
or in the relevant decisions. What exists is a command that runs and exits;
what the design needs is a process holding a link. Nobody had written down that
those differ, so several questions had no answer.

0057 settles them. It runs on every node -- the host is what makes a machine
managed, so a machine without one is not a node. Root, because no useful subset
of the job is unprivileged. A systemd unit, because something must survive a
reboot to hold the link.

It never manages its own unit. The temptation is obvious and it ends with a
host stopping itself half way through an apply, leaving a machine with nothing
running to fix it. The installation owns the host; the host owns everything
else.

Installed as a package, with a tarball as the floor. The package carries the
unit file, the state directory and an upgrade path, which a bare binary does
not. But the mesh's package repository is hosted on the mesh, so any route that
needs the mesh to install the thing that joins the mesh is a circle -- the
tarball is the path that must never acquire a dependency.

Reconciles on start, on a declaration, on a timer and on reconnect. The timer
is the one easy to leave out, and without it `owned` reports what the host
applied rather than what is there -- ADR 0035 violated by omission.

The records checker caught this commit on its first attempt: 05 listed 0057 in
its frontmatter while 0057 is still proposed, and a to-be document may not rest
on an unaccepted record. The section now says so in the body instead.
2026-08-27 21:06:00 +02:00
jschoubben 2330d74c1b The host's vocabulary is complete; 05 and 07 said otherwise
All six shapes are built. 07 still said the last three did not exist, and 05
still described stage 2 as having built three of six.

Records what the lab still cannot do, because that is now the only thing
between here and an end-to-end substrate bootstrap: a sealed scenario cannot
fetch an image and its machines carry no container runtime, so package,
container and action were verified against a real machine instead.
2026-08-27 20:36:58 +02:00
jschoubben c631cbd07c The bootstrap starts a step earlier than recorded
Asked whether postgres has to be installed, and the answer exposed a missing
step. The store is a container, so something must run containers before anything
else happens — and a container runtime is a PACKAGE, not a container.

Step 0 is where several threads meet. It is what the host's capability detection
already reports, and the first use of that report by something other than a
person. It is adopted rather than installed when the machine already has a
runtime with configuration somebody chose. And it is a package, needing the
machine's own package manager and a network, both of which ADR 0046 permits.

So the host's bootstrap vocabulary is six shapes: package, container, file,
directory, service, action. Stage 2 built three of them.

The node host design now names which three remain and why the lab cannot yet
exercise them — a sealed scenario fetches nothing and its machines carry no
container runtime, which is lab-installation work rather than a constraint on
the design, because production machines have a network.
2026-08-26 23:58:19 +02:00
jschoubben bea052753e ADR 0043 — what a declaration is
Stage 2 could not start without it. Three constraints already bound the shape
and between them they decide most of it.

JSON, because the standard library carries it and carries no YAML, and a YAML
declaration would put a third-party parser inside the one binary whose whole
argument is that it needs nothing — to gain authoring comfort in a document
generated by a machine and read by a machine.

An ordered list, because ordering is a DECISION. A host deriving order from
declared dependencies would be deciding the thing most likely to differ between
what the control plane intended and what the machine does. The control plane
knows what depends on what; it says so by saying when.

Unknown is refused, never skipped — an unknown version, type or field refuses
the whole declaration. A host that skipped what it did not understand would
apply most of a declaration and report success, which is 04-ISSUES/003 with the
declaration on the other side of the wire.

Complete for what the host OWNS, and only that. It removes what it previously
applied and is no longer declared, which it knows from the store rather than by
inference, and never removes what it did not create — a converger that treats
'not declared' as 'must not exist' deletes what the mesh never put there.

Two consequences arriving earlier than the build order suggested: the store is
load-bearing at stage 2, because nothing can be removed without knowing what was
applied. And a closed address space bounds the first vocabulary to what needs no
network, because a scenario has no route to a package repository.
2026-08-26 02:04:41 +02:00
jschoubben 92e8c74ce4 ADR 0041 and the build handoff for the node host
Building tier 0 forced the question "the one binary installed by hand" had been
carrying unexamined. A TypeScript host needs a runtime present before it runs,
so the thing installed by hand becomes two — and the second must be installed by
the means the host exists to replace.

So the host is a statically linked binary that requires nothing present, written
in Go. Rejected: a runtime installed first, which breaks the property the tier
rests on; and bundling the runtime into the executable, which carries ninety
megabytes to preserve a language choice and puts a young feature at the bottom
of the stack.

The argument that decided it is architectural rather than about taste. 0037
means the host never queries the mesh database and 0039 means it only receives
declarations, so the host shares NO code with any other tier — not a client, not
a schema, not the SDK. The language boundary falls exactly on a boundary that
already exists, and a second language usually costs duplicated logic where here
there is none to duplicate.

§8 gains a scope: it said "TypeScript throughout" when everything was a service
or a surface, and is now scoped to those with tier 0 named. Another sync owed.

Playbook 04 steps 2 and 4: repos.md records mesh-host as existing, the design
takes code: [mesh-host] and status: in-progress.
2026-08-26 00:16:07 +02:00
jschoubben b9facf9375 Design the node host
Playbook 02 step 3, on four recorded decisions. Tier 0 has one job — apply
declared state on this machine — and the six absorbed concerns are instances of
it, not additions to it.

Specifies the six parts and what each owns, and the two properties that make
apply trustworthy rather than merely present: every applier reads back, because
setting a value is not evidence the value took; and what was applied is recorded
after it works, never before, because a failed apply leaves the machine wherever
it reached and nothing must claim otherwise.

Build order is staged so each stage is verifiable in the lab before the next
exists. Stage 1 is profile and inventory — no control plane, no declarations, no
network — and it is deliberately the smallest useful thing, because `place:` has
nothing to place and the lab therefore raises empty machines. Stage 1 ends that,
and every later stage is tested by a lab that already works.

Stage 2 is the one that could invalidate the tier boundary: whether one host can
raise the substrate alone is Move 1's assumption and has never been proved.

Every decision the design rests on is given the test that asserts it, per 0034 —
including the dependency-direction lint, which is what makes "the host never
queries the mesh database" a rule rather than an intention.

Six things left open and named, including the one that host-size.md could not
measure: zero dependencies, but still six vocabularies.
2026-08-26 00:10:24 +02:00