Commit Graph
52 Commits
Author SHA1 Message Date
jschoubben 5218b06c02 Fold the control plane's build decisions into 0006 and 0008
Back to 23 records. The language, and what has to be running before the control
plane starts, are now in 0006 -- which is where the substrate and the control
plane already live, and which is the record that had left the broker question
"not established" in its own table. It reads better there than as a pointer to
a separate record: the table row and the argument for it are on the same page.

The store mechanics went into 0008. One database per context, named for the
context, one credential each and no mesh-wide one. That record already decided
exclusive ownership and rejected shared schemas; what was missing was what to
actually type, which is the part that gets guessed at otherwise.

Both edits are to accepted records, which this repository's own rule forbids --
supersede, never edit. Recorded here so it is visible rather than silent. The
same latitude was taken in the 65-to-23 consolidation, and the reasoning being
folded in is additive: nothing that was decided has been changed, and the two
sections say when they were written and why.
2026-08-29 03:08:50 +02:00
jschoubben 82a3065f82 Tier 2 exists, and the token was missing a quarter of itself
mesh-control is built as far as it can honestly go: one context of seven,
inventory, with its schema and the command that applies it. The repos map and
the control plane design say so, and point at ADR 0024 for what it took.

Separately, and more importantly: this repository described the enrolment token
as carrying three things when ADR 0004 says four. The missing one is the
control plane's signing identity -- the reason a node does not have to trust
the broker it dials.

Without it the control plane's authority is transitive through the broker, and
0004 spells out what that costs: a compromised broker could forge declarations,
and since the host applies whatever the link delivers, that is the whole
machine. The record has the argument in full; the design doc had dropped the
conclusion.

Found by reading the two together while deciding what the control plane must
store, which is roughly the only way it would have been found -- both documents
are internally consistent and only disagree with each other.
2026-08-29 02:49:58 +02:00
jschoubben 84f4425fd6 The broker precedes the control plane, and it is written in Go
Two things found by trying to build tier 2.

The substrate design asked whether the message broker has to be running before
the control plane, and framed it as depending on whether the control plane's
own parts talk to each other over it. They do not -- it is one process -- so
under that framing the broker stays out of the bundle.

The framing cannot answer the question. What decides it is how the control
plane reaches a node, and the answer was already decided: only ever over the
link, and the link is the broker. So provisioning the broker would require the
broker. The first node does not escape this by being local, because it enrols
the ordinary way, by dialling the broker at the address in its token -- which
was deliberate, and worth keeping.

The bundle is two images now. The record says what that costs, including a
certificate the broker needs at a moment when there is no mesh to issue one.

The language had never been decided for tier 2. Go, for the same reason the
host is: the bundle pins this image by digest and runs it where nothing can
check it, so the image should hold the program and nothing else.

Also corrects something already built: the bootstrap created one database and
called it 'mesh'. ADR 0008 grants a context only what it exclusively owns and
ADR 0006 says the mesh database names a thing that will not exist. One database
per context, so one today, called inventory.
2026-08-29 02:32:46 +02:00
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00
jschoubben e1febe8e0f Renumber the records 1 to 23
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18,
19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only
the archaeology of what used to be there.

Renumbered contiguously. Renames run in ascending order, so every target number
is already free and no two files ever collide.

The reference rewrite is one simultaneous pass rather than a sequence of
replacements. Numbers moved into slots other numbers were vacating -- the node
host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time
would have cascaded and silently pointed things at the wrong record.

Seven plain-text references survived the merges as prose rather than links,
naming records that no longer existed: the enrolment token, the link boundary,
what a declaration is, reachability, the repository structure. Each mapped to
the consolidated record that now holds it.

Verified rather than assumed: every [ADR NNNN](path) link now has matching text
and target, checked across the whole repository, and the checker passes.

Frontmatter `consolidates:` lists dropped -- they named records that are gone,
and each consolidated record already says in prose what it absorbed.
2026-08-28 23:28:34 +02:00
jschoubben 77f3a4cea7 Consolidate: 65 decision records to 23
Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.

  the node host          8 -> 1    applies not decides, depends on nothing,
                                   per operating system, root service, the
                                   launcher, episodic, what a declaration is,
                                   actions from the bundle only
  a node and how it joins 4 -> 1   what a node is, joining, the link as
                                   security boundary, the enrolment token
  modules and the graph   7 -> 1   everything is a module, no domain modules,
                                   three edges, provisioning, the core library
  substrate and control   6 -> 1   the test, seven contexts, one control plane,
    plane                          the authority is not a database, the named
                                   products, the pinned bundle
  connectivity            3 -> 1   a route is a grant, reachability declared,
                                   filter rules
  delivery                5 -> 1   reconciliation not a pipeline, artifacts,
                                   the three silos, a failed step, the verdict
  the lab                 5 -> 1   (earlier)
  how this repository     10 -> 1  (earlier)
    works

Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.

The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.

The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
2026-08-28 20:03:24 +02:00
jschoubben 5e83ac2c22 Consolidate: 65 decision records to 52
Jochen: a normal application has 3-5 ADRs, maybe 10 for a large one, and we are
at 65. Fair, and the cause is mine -- I recorded every FINDING as a decision
rather than every fork in the road.

Two merges, both cases where one decision had been split across many records
because it was taken over several days rather than at once.

0019 absorbs ten records about how this repository works: what it is and that
it is public, the folder flow, the two design layers, the issue front door,
status in frontmatter, playbooks, the naming rule, the product name. Those were
never ten decisions -- they were one, seen from ten angles as the repository
took shape.

0016 absorbs the five about the lab: a node is a virtual machine, a router is
scenery, a scenario declares the underlay, a scenario is a closed address
space, and the two scenario classes. Same pattern -- one design, split by the
order it was worked out in.

The consolidated 0019 also raises the bar for what earns a record, since that
is what produced 65: a record is warranted when there is a genuine fork -- a
direction reversed, an alternative that will be proposed again, something
contested. A finding is not a decision, and a bug is certainly not. Everything
else belongs in the design document where the reasoning is actually read.

The checker earned its place here. Deleting nine records left 13 dangling links
across the repository and it named every one, including in AGENTS.md. Nothing
was found by reading.

Remaining clusters worth the same treatment: the host (8 records), delivery
(5), modules (6), connectivity (4), substrate and control plane (4). That would
be 52 down to roughly 30.
2026-08-28 18:53:19 +02:00
jschoubben 10365f2eae Consolidate the design layer: one place per topic
Jochen: a jungle of specs that slightly contradict or patch each other, and
what matters is a working state rather than history. Both are fair and both are
mine.

Measured rather than assumed. 05-the-node-host and 09-the-node-lifecycle both
covered enrolment, the install commands, the unit file, the launcher and
reconcile -- I wrote 09 without taking anything out of 05, so the same things
were said twice and could drift apart.

Split by what each document IS. 05 is the component: what the host is, its
parts, the declaration vocabulary, the build order, how it is verified. 09 is
what happens to it: install, enrol, run, upgrade, retire. The whole "The
process" section left 05, and the unit file moved to 09 where installing is
described. 05 goes from 338 lines to 245 and now points at 09 rather than
restating it.

09 also carried a 105-line "Resolved" section -- six mechanisms framed as
"these were open and here is the answer". The content is needed; the framing is
history, and history is what makes a document read as a changelog rather than a
description. Renamed to what it actually is and the was-open phrasing removed.

Also added 10-delivery.md, which did not exist: four accepted decisions --
0054, 0063, 0064, 0065 -- had no design document at all, which is the specific
reason the delivery picture felt scattered. It is now one document covering
modules, the three edges, the core library, and how a change becomes a running
thing, with a table of what each property is designed against and what must
exist before it can be built.
2026-08-28 18:40:46 +02:00
jschoubben ba0d01788e 0062: a host may be episodic; 0060's Android gap closed
0060 named the gap and did not close it: everywhere else an init runs the
launcher at boot, and Android grants neither an init to register with nor
anything worth supervising, because a supervisor would be killed alongside what
it supervises.

Closed by narrowing what is required rather than building something. A host is
resident or episodic, and both are hosts. Being killed by the platform is
disconnection, which 0036 already made ordinary -- and every mechanism an
episodic host needs already exists because it was built for laptops that close.

A partial host can join a mesh and cannot be the first node, since every
bootstrap step is a shape it refuses. Its bundle says so.

Two consequences that are easy to miss: last-heard-from means much less on an
episodic host, so a healthy phone reads as a dead server unless the reader
knows which kind it is; and a declaration may take a long time to land, which
makes 0058's outstanding-versus-failed distinction load-bearing.

Still open, and in that order: what an Android node is FOR, and only then how
it is started.
2026-08-28 01:24:07 +02:00
jschoubben f1b1cd9aa0 Review: three ADRs no longer said what we had concluded
A sweep for claims overtaken by the last few days. Annotated rather than
rewritten, following the pattern already in 0049 -- what changed and why is the
useful part, and an accepted record should not quietly become something else.

0057's init section was wrong on all three of its claims. It said the host
needs FOUR things from an init; 0061 reduced that to one. It said every machine
the mesh targets already has systemd; Alpine does not, and it is the intended
first node. It said there is no second init to abstract over; there is now, and
the answer is still not an abstraction -- it is a four-line file per system.
What survives is the part that was always right: an init is not a dependency in
0041's sense, because it is not installed, it is what the machine already is.

0048 named Docker as the container runtime. It is now docker or podman,
detected rather than chosen -- because adoption keeps what a machine already
has, so naming one contradicted a rule already decided. That row is the only
one of the five that names two, and the record now says why.

0060 claimed the bundle is portable across operating systems. Its mechanism is;
its contents are not -- package names, unit names, service names all differ, so
an Arch host embeds an Arch bundle. That was my error, and it is the exact
confusion behind the question that found it.

The design layer had the same drift: 07 and 09 said "Docker" where they meant a
container runtime, 09 said systemd restarts the host after an upgrade when the
launcher does, and both install snippets assumed Arch. They now show Alpine and
Arch side by side, which makes the point better than prose did -- step 1
differs per system, step 2 never does.

Checked and NOT changed: 0047's "the vocabulary grows by one shape" is a claim
about the rate, not the count, and is still true. 0037 lists docker among tools
the host manages, which it does. 0041 says nothing about either.
2026-08-28 00:43:47 +02:00
jschoubben c557f99cba Record what testing podman actually showed
0060 said the container runtime was a separate decision. It is now made, and
the reasoning is worth keeping because it is the opposite answer to the same
question one paragraph earlier.

Abstracting service managers is lossy -- systemd and OpenRC are different
models and LoadState has no equivalent. Container runtimes converged on one CLI
deliberately, so almost nothing is lost: checked against podman 6.1.0, run,
rm -f and docker's own template syntax for state and labels all work unchanged.
Only the probe differs. So: a two-entry lookup, not an interface.

The difference that is NOT in the CLI is the one that would have shipped
silently. Podman accepts --restart unless-stopped, records it, and has no
daemon to act on it -- containers do not return after a reboot unless
podman-restart.service is enabled, which by default it is not. Every command
reports success and the effect does not happen.

That belongs in the declaration rather than the host: a node using podman is
told to enable the unit. Which is what made the service shape's missing 'boot'
field visible, and it is now built.
2026-08-27 23:59:05 +02:00
jschoubben e1ad39b500 Per-OS hosts, and an init asked for only start and restart
0060 -- the host is built per operating system. systemd and pacman are the Arch
host's implementation, not abstractions the mesh has to grow. They are not
independent choices: a machine has pacman because it is Arch, and the package
manager, service manager and packaging format arrive together as one decision
somebody made at install time.

Rejected abstracting them, and the reason is correctness rather than effort.
The service applier reads LoadState to tell "not installed" apart from
"stopped", which is what stops it reporting absence as success. An interface
spanning systemd and OpenRC degrades to what both express, and the lowest
common denominator is exactly where that fault lives.

Almost all of it is shared -- the vocabulary, store, apply loop, read-back
discipline, refusal model, bundle and link are portable. Two appliers differ.
And delivery was already per-OS, since a .pkg.tar.zst is an Arch artifact, so
this is the seam that already existed.

Android is the interesting case rather than Debian: no service manager, no
package installation, usually no root. Such a host implements file, directory
and action and refuses the rest -- the same refusal a host already gives an
unknown type, with a different reason. Those three are the portable floor.

The container runtime is deliberately left open: it is not an OS split, since
Arch runs docker or podman.

0061 -- the init is asked for start-at-boot and restart-on-exit, and nothing
else. Both are expressible in OpenRC, runit, s6 and an Android init.rc.
Counting failed starts and rolling back moves into a launcher, because that is
the one piece which must work when the host does not, and a script with a
counter can be tested where OnFailure= can only be hoped for. Supersedes 0059,
keeping its reasoning in full.

The checker found all six places citing 0059 and refused the commit until they
named the replacement.
2026-08-27 23:46:11 +02:00
jschoubben dcc4b8339c Say who consumes the broker and who writes the registry
Left implicit by the previous commit, which said the owning context writes
without saying what does the consuming.

The control plane is the consumer, and there is one of it. Seven contexts but
one deployable, so it is one process dispatching internally rather than seven
consumers racing -- which matters because the as-is records two consumers
accidentally sharing a queue and silently splitting the traffic, each getting
half of what it expected. With one consumer that cannot arise.

The broker is also the buffer while the control plane is down: nodes keep
publishing, messages queue, the control plane drains them on return. That is
what makes a single control plane tolerable -- an outage delays the mesh's
knowledge rather than losing it.

One consequence named because it will otherwise be discovered: an unbounded
queue grows until the broker's disk is full, and the broker is the component
every node depends on. The bound is per queue and undecided -- dropping the
oldest health report is obviously right, dropping the oldest declaration
acknowledgement is not.
2026-08-27 22:20:03 +02:00
jschoubben 19997d56c3 Approve 0057-0059, with four corrections from review
Not approved as drafted -- four things came out of checking them against each
other, and one was a bug that would have broken every upgrade.

The bug: 0059 specified Restart=on-failure while 0057 has the host restart onto
a new binary by exiting CLEANLY. on-failure does not restart a process that
exited zero, so every upgraded node would have been left stopped, having
successfully upgraded. Found by reading the two records against each other
rather than by either alone. Now Restart=always in all three places that
mention it.

The host cannot run in a container, and the reason is decisive rather than
stylistic: step 0 of the substrate bootstrap installs the container runtime, so
a host inside a container would need the thing it exists to install. It would
also break 0041 -- copy it onto a machine and run it stops being true when the
machine must already have a runtime. Everything above tier 0 is a container;
the host is not. That split is the tier boundary, not an inconsistency.

systemd is named rather than abstracted. An init is not a dependency in 0041's
sense: 0041 is about what must be installed before the host works, and an init
is not installed, it is what the machine already is. The unit file is the only
systemd-specific artefact and it belongs to the package, so a machine with a
different supervisor ships a different package.

The mesh is a watchdog, and my first draft was half an answer. Recovery must be
local -- nothing dials a node, and a host that cannot start cannot report. But
detection is the mesh's, and a local supervisor structurally cannot do it: it
sees one process failing and cannot tell a broken machine from a broken
release. Only something watching every node can, and that distinction decides
whether the response is "fix this machine" or "stop shipping this version". So
a host rollout is staged -- a few nodes, wait for heartbeats, continue or stop
on silence. Local rollback still needed, because the canary nodes break and
because a node offline during the rollout gets the declaration later with no
batch around it.

The first declaration is the overlay and nothing else. Forced, because a node's
address and peers are assigned rather than chosen. But also the way back in: a
node reachable over the overlay can be fixed by hand if a later declaration
breaks it, and a large first declaration risks a node that is broken and
unreachable at once.

Also stated plainly, because it reads as a contradiction: nodes reach each
other over the overlay and every node consumes from the broker; what 0039
forbids is an inbound CONTROL surface, not reachability.

And in 06: no node holds a credential to any control-plane store, for reads or
writes. Four ADRs already say this separately and none of them said it in one
place. Nodes state over the broker; the owning context writes. With a note that
most high-frequency writes are observability's, not the registry's -- routing
logs into the registry would be the shared-schema mistake arriving through a
door marked performance.
2026-08-27 22:12:12 +02:00
jschoubben 605c9fd441 Changes are pushed, not polled; and a stuck host rolls itself back
Two corrections and one new decision, all from Jochen catching things.

Pushed, not polled. I described updates as landing "on the next reconcile",
which reads as polling and is not the design. A declaration arrives as a
message on a link that is already open; the host applies it then. Polling over
an existing connection would be slower to land AND constant traffic to learn
nothing.

The timer is for drift and nothing else, and it cannot be replaced by an event
for a definitional reason: drift is change the mesh did not make -- somebody
edited a managed file, a distribution upgrade replaced a config -- so nothing
will ever publish a message about it. Only looking finds it.

Separated the heartbeat from the reconcile timer, which I had been conflating.
They point in opposite directions and answer different questions: the timer
looks at the machine and asks whether it still matches; the heartbeat reports
upward and is what makes silence mean something. A node with nothing to do
sends nothing, and without a heartbeat that is indistinguishable from a node
that stopped.

0059 -- a host that cannot start is rolled back by the service manager. I had
left this open on the grounds that recovery meant the host judging its own
health. That objection does not survive being asked properly: a keepalive is
something else judging the host. The watchdog must be local, because nothing
dials a node and a host that cannot start cannot report -- so it is the service
manager, which is already there.

The failure it prevents is sharper than "the node is down": a host that will
not start looks exactly like a machine somebody switched off, which is the one
condition this design has deliberately decided not to alarm on. So a bad
release reaches every node, each goes quiet, and the mesh reports a fleet of
sleeping laptops.

Confirmed means started and completed one reconcile -- deliberately not "the
link is up", or a laptop on a train would roll itself back. The rollback is a
script shipped by the package, not a host subcommand, because a binary that
will not start cannot be its own recovery. It rolls back once: a second failure
means the machine is the problem, not the binary.

Also refined the records checker, which produced a false positive: a proposed
record may extend another proposed one, because decisions are drafted in chains
and the alternative is marking things accepted to satisfy a check. An accepted
document resting on a proposed record still fails, and that was verified.

0057, 0058 and 0059 are all proposed.
2026-08-27 22:04:26 +02:00
jschoubben aeea2a9f9a Resolve the host lifecycle's open items, and say how the host is delivered
The upgrade question turned out to be a delivery question, so 0058 answers
both.

Today's third silo runs once per node and sends each one a command to install
and start. That is where the as-is records a package install that 404ed from
every mirror while the job went green, an image pull failure that did not fail
the deploy, and a verify stage that was built and never scheduled because it
was missing from a list.

The shape underneath all of those is that the thing reporting success was not
the thing doing the work. Meanwhile ADR 0037 has given every node a component
that applies state, reads back and reports -- so two mechanisms now change a
node and only one checks its work.

0058: a pipeline ends when the declaration is updated. Deploy stops sending
commands to nodes and becomes one write. The host applies it on its next
reconcile, and the host cannot report success it did not verify. The verify
stage disappears as a stage, which is the point -- verification stops being a
step that can be left off a list.

A pipeline result now means "the declaration is updated, and here is which
nodes have applied it". It does not wait for every node, because a node may be
legitimately switched off for a week. Outstanding is reported separately from
failed, since conflating them is how the old system produced a stall with no
error anywhere.

The host is delivered by exactly this path and needs no new resource type: a
`file` writes the package manager's config pointing at the mesh's repository, a
`package` names the version. Added a step I had missed -- before exiting for a
restart, the host runs the new binary once. A package can install something
that does not execute here, and that turns "the node never came back" into "the
apply failed and said why".

Six open items resolved: re-enrolment is decided when the token is issued and
revokes the previous identity; the mesh keeps a recovery copy of what each node
reports it owns, which un-strands the orphans; last-contact is reported with no
threshold, because a laptop off for three weeks is doing nothing wrong;
adoption always completes but a failed line makes a node ineligible for
assignment; a briefing is a structured document whose outcome is computed from
its lines; and the token is printed once and carried by hand, which is the
property that makes it worth anything.

Still open and named: automatic rollback of a host version that will not start.

0057 and 0058 are both proposed.
2026-08-27 21:53:38 +02:00
jschoubben 2204b01909 Design the node lifecycle end to end
The host was described as a component and never as something that runs for
years on a machine somebody else also uses. 09 covers every state a machine can
be in and every transition between them.

Four states: unmanaged, hosted, enrolled, disconnected. Only the last two are
nodes, and they are the same node in two situations. `hosted` -- the host
installed but never told which mesh it belongs to -- had no name before and is
where a machine sits between the two adoption commands.

Things that were unclear and now are not:

The first node walks the same path in an unusual order: reconcile from the
bundle, the control plane it just raised issues a token, enrol against it. Its
specialness lasts two commands. A side effect worth having -- enrolment is
exercised on node one, rather than being written and first used on node two.

Enrolment reports profile and inventory BEFORE the control plane decides
anything. The profile is the input to that decision, not a diagnostic; the
control plane cannot decide what a machine should run without knowing what it
can run.

Rebooting mid-apply is safe by construction. The store records each resource
after it worked, so a host that dies half way through comes back and applies
the rest. The rule that stops the host lying about what it did also makes it
crash-safe.

Retiring splits in two. Graceful is a final empty declaration. A node that is
gone will reconcile its last declaration forever -- the honest consequence of
making disconnection ordinary. The answer is not to make the host expire but
that the node holds nothing that outlives revocation: every grant is a per-node
credential revoked at the provider. A lost node keeps running and stops being
able to reach anything. Said plainly rather than implying the mesh can switch a
machine off, which it cannot and should not.

Losing the store is quiet and permanent, so it gets its own section. The host
re-enrols and re-applies fine; what does not come back is removal, because
resources it no longer has a record of become unowned and sit there
indefinitely.

Also corrects 0057, which said the mesh must not upgrade the host at all. That
conflated two acts. Replacing the binary is safe -- Unix keeps the running
inode. Stopping the unit is not. So the host may apply a package naming itself,
and restarts by finishing its apply and exiting cleanly, letting the supervisor
start it on the new binary. It never asks the service manager to restart it.
That makes a fleet-wide host upgrade an ordinary declaration, which the first
draft gave up on.

0057 remains proposed.
2026-08-27 21:16:44 +02:00
jschoubben 3ab11c96ef Say what the host process is: a root service, installed as a package
The design described what the host does and never what it is at runtime. The
words daemon, long-running, interval, poll and heartbeat appeared nowhere in it
or in the relevant decisions. What exists is a command that runs and exits;
what the design needs is a process holding a link. Nobody had written down that
those differ, so several questions had no answer.

0057 settles them. It runs on every node -- the host is what makes a machine
managed, so a machine without one is not a node. Root, because no useful subset
of the job is unprivileged. A systemd unit, because something must survive a
reboot to hold the link.

It never manages its own unit. The temptation is obvious and it ends with a
host stopping itself half way through an apply, leaving a machine with nothing
running to fix it. The installation owns the host; the host owns everything
else.

Installed as a package, with a tarball as the floor. The package carries the
unit file, the state directory and an upgrade path, which a bare binary does
not. But the mesh's package repository is hosted on the mesh, so any route that
needs the mesh to install the thing that joins the mesh is a circle -- the
tarball is the path that must never acquire a dependency.

Reconciles on start, on a declaration, on a timer and on reconnect. The timer
is the one easy to leave out, and without it `owned` reports what the host
applied rather than what is there -- ADR 0035 violated by omission.

The records checker caught this commit on its first attempt: 05 listed 0057 in
its frontmatter while 0057 is still proposed, and a to-be document may not rest
on an unaccepted record. The section now says so in the body instead.
2026-08-27 21:06:00 +02:00
jschoubben 2330d74c1b The host's vocabulary is complete; 05 and 07 said otherwise
All six shapes are built. 07 still said the last three did not exist, and 05
still described stage 2 as having built three of six.

Records what the lab still cannot do, because that is now the only thing
between here and an end-to-end substrate bootstrap: a sealed scenario cannot
fetch an image and its machines carry no container runtime, so package,
container and action were verified against a real machine instead.
2026-08-27 20:36:58 +02:00
jschoubben e1f4c7d9e0 Approve 0054-0056, apply them, and fix the two smaller findings
0003 is now superseded by 0056. Nothing is left proposed.

Applied:
- 06 corrected from ten contexts to seven plus the api, each row now stating
  why it passes the more-than-one-node test. work, knowledge and stream are
  named as mesh-hosted rather than dropped; `ai` folds into config; `record`
  is deferred explicitly rather than listed. Its frontmatter now cites 0055.
- how-we-build §4 amended per 0054, and the derived page republished by
  playbook 05.

The sync found the drift the playbook exists to catch: the published §4 and
the source did not say the same thing. The source said "four accidents, not
four boundaries"; the published page said "one intent expressed four times",
and only the published page carried the scope caveat. Same rule, two texts,
already diverging. Verified the republish by reading back -- the new rule is
present and the old section's body returns nothing -- rather than trusting the
success message.

The two smaller findings:
- 0051 separated the transport identity from the declaring authority. It said
  the token carries "an address" and "the identity to expect" without saying
  what the node dials. It dials the broker, so pinning only that would make the
  control plane's authority transitive and let a compromised broker forge
  declarations -- which, since the host applies whatever the link delivers, is
  the whole machine. The token now carries four things, and declarations are
  signed and verified per declaration. Cost recorded: rotating the signing
  identity is fleet-wide.
- 0026 no longer restates 0022's rule about generated views. 0022's own words
  are "prose does not restate status; one place, and two is one too many",
  which is what 0026 was doing to it.
2026-08-27 02:21:34 +02:00
jschoubben ef5dd0751b Approve 0049-0053; drop a to-be item superseded by ADR 0044
The 'domain grouping' item cited ADR 0017 as live guidance. 0044 superseded
it -- there is no domain module to group into, so there is no domain list to
settle.
2026-08-27 01:00:29 +02:00
jschoubben ccbbfa9c8a One node runs the control plane, and nothing takes over
Closes the two open questions in 06 and 08, which turned out to be one
question: how many control planes run, and what happens when the hub is down.
Both were drifting toward redundancy by default -- a standby plane, a second
hub, an election to pick between them. That is not one feature but a property
every layer must then honour, and each layer gets it wrong independently.

Not wanted, and not needed. A handful of machines with one node hosting the
registry is not a distributed system.

The argument for why this is sound rather than merely cheap is that the design
already tolerates it by construction. ADR 0036 makes reachability state rather
than class; the host reconciles from its own store (0043) and never needed to
ask anybody to hold the state it was last given. So the control plane being
down is not a new failure mode -- it is every node in the ordinary disconnected
situation at once. What is lost is change, not operation.

No node holds a contended role: the control plane is assigned like any other
module, and the overlay hub is declared (0050). No promotion, no quorum, no
fencing, no split brain, no replicated store, and no "which node is
authoritative" recurring at every layer.

Two consequences stated plainly rather than buried. The control-plane node is a
single point of failure -- deliberate, and said out loud so it stays
deliberate. And recovery is restore rather than failover, which makes backup
the availability story rather than hygiene.

The sharpest one is the clock: the control plane owns certificate issuance
(0049), so an outage outlasting a renewal window expires every public name.
That bounds how long recovery may take, and nothing measures it today.
2026-08-27 00:55:10 +02:00
jschoubben 4e80820e2f Design connectivity in full: overlay, resolution, exposure, filtering, certificates
Written as one document because the five are one design. They share inputs,
they must agree, and every one of them today is computed in a different place
by a different module from a different copy of the same facts.

The through-line is that none of the five can be answered by a machine alone,
so all five are decided centrally and delivered as `file` resources. That costs
no new host vocabulary and removes both remaining direct database connections
from nodes -- wireguard and traefik are the only two, and both are connectivity.

Three decisions fall out, all proposed:

0050 -- reachability is declared, not inferred from an address. The RFC1918
regex is wrong for carrier-grade NAT (100.64/10 tests as public, so an endpoint
is written to an address nothing can reach), wrong for IPv6, and wrong for a
routable address behind a closed firewall. The lab needing TEST-NET-3 to
satisfy the regex is the same bug from the other side. Also kills hub election
by address prefix, which fails silently and makes renumbering an outage.

0051 -- the enrolment token carries where the mesh is and how to recognise it.
Closes two circles with one mechanism: verifying the mesh needed the CA, and
obtaining the CA meant trusting whoever handed it over; and a node had to reach
the mesh before it could resolve any mesh name. An address plus a fingerprint,
carried out of band, resolves both -- and closes the CA question 0049 deferred.

0052 -- a filter rule names its source. `scope:` is declared in five manifests,
is part of no rule type, and is referenced by no code, so those manifests
appear to restrict ports and restrict nothing. Removed rather than implemented;
the general fix is refusing unknown keys, which the host already does and
manifests do not.

Also corrects two claims in 0049 asserting wireguard was already handled.
Research 006 says both modules still reach upward; neither is.
2026-08-27 00:36:49 +02:00
jschoubben 8d9282d86b Resolve the ingress gap: a route is a grant
ADR 0048 named ingress as an unclosed hole -- nothing said what terminates
TLS, how a public name reaches a container, or which tier owned it. Resolving
it needed no new concepts, which is why it survived: nobody had applied the
rules already written to it.

Ingress is not substrate. The control plane does not need a route to start,
and no node needs one to reach it -- the node dials out and has no listening
control surface. It grants itself a route afterwards, like a bucket.

A route is an instantiation edge under ADR 0044. The direction mirrors a
database -- the consumer supplies a target and receives a name rather than
credentials -- but it is the same edge.

The substantive finding is that exposure is three facts at two scopes: name
resolution and certificate issuance need to know which node is publicly
reachable, and only the proxy mapping is a single machine's business. That is
why it belongs to the connectivity context, and why Traefik doing all three on
the node is wrong.

Which matters beyond tidiness: research 006 counted traefik as one of two
modules opening a direct Postgres connection, reading nodes and mesh_ca. That
violates 0037, 0045 and 0039 at once, and is why every node permanently holds
a credential to the control plane's database. Deriving the config centrally and
delivering it as `file` resources removes it, costs zero new host vocabulary,
and closes the set 0039 identified -- wireguard was the other.

Left open deliberately: the mesh's internal CA is the other thing traefik
reads, and it belongs to the link's mutual authority, not to exposure.
Conflating the two is what made the gap hard to see.

Also fixes an inconsistency from the previous commit: 06 still claimed the
virtual host was raised from the bundle.

Proposed, not accepted -- for review.
2026-08-27 00:22:52 +02:00
jschoubben 4d19e93900 Name the substrate's actual products
The design layer described every service by role and never once by name:
Postgres appeared in zero design documents. That was over-application of the
research rule "never identify the mesh it observed", which is about node names
and domains, not software.

Two things were actually broken by it. substrate.lock pins images by digest and
a digest belongs to a named image, so the bundle could not be written from the
design. And a reader could not tell a settled choice from an unexamined one --
"a relational store" reads identically either way.

ADR 0048 names them: PostgreSQL, LavinMQ, MinIO, an OCI registry, Docker. The
argument for each is continuity, which is a real argument -- replacing a
substrate service migrates the mesh's own state. Role and product are now both
written, because the design depends on the protocol while the installer needs
the product.

Also separates two questions the substrate doc had merged: being substrate and
being in the bundle. Only Postgres must precede the control plane; the rest are
substrate by role and ordinary by delivery. Whether the bus joins it is left
open, because it turns on the control plane's internal shape.

Names the forge as Gitea, and records ingress/Traefik as an unclosed gap rather
than a naming one -- nothing says what terminates TLS or which tier owns it.

Fixes a miscount: the host's bootstrap vocabulary is six shapes, not five.
2026-08-27 00:11:38 +02:00
jschoubben c631cbd07c The bootstrap starts a step earlier than recorded
Asked whether postgres has to be installed, and the answer exposed a missing
step. The store is a container, so something must run containers before anything
else happens — and a container runtime is a PACKAGE, not a container.

Step 0 is where several threads meet. It is what the host's capability detection
already reports, and the first use of that report by something other than a
person. It is adopted rather than installed when the machine already has a
runtime with configuration somebody chose. And it is a package, needing the
machine's own package manager and a network, both of which ADR 0046 permits.

So the host's bootstrap vocabulary is six shapes: package, container, file,
directory, service, action. Stage 2 built three of them.

The node host design now names which three remain and why the lab cannot yet
exercise them — a sealed scenario fetches nothing and its machines carry no
container runtime, which is lab-installation work rather than a constraint on
the design, because production machines have a network.
2026-08-26 23:58:19 +02:00
jschoubben 93470f6162 ADR 0047 — the bundle may carry actions the link may not
The bootstrap's sharpest open question, and the framing was wrong. "State on
this machine" was being read as the filesystem and the service manager. A
service running on this machine IS part of this machine — writing a file and
creating a database in a local store differ in mechanism, not in scope.

The real question was underneath: must the host learn what a database is? It
must not. Giving it a `database` resource type means tier 0 knows Postgres, then
a bucket, then a virtual host — the host acquiring the substrate's vocabulary
one service at a time, which is what ADR 0037 exists to stop.

So the bundle declares an ACTION and the host runs it and verifies it. What a
database means stays with the module that provides one; the host knows only how
to run a declared action against something local and check the result. Its
vocabulary grows by one shape rather than by one resource type per service.

Actions are permitted in the bundle and forbidden over the link, and the
asymmetry is deliberate. A bundle arrives WITH the binary: anyone able to put a
hostile action in it could equally have put it in the host itself, so refusing
actions there buys nothing and costs the bootstrap. The link is a separate
party, reachable separately, and an action there is the unbounded blast radius
ADR 0039 refuses. That decision stands unchanged.

And ongoing provisioning is not the host's at all — the control plane does it
once a mesh exists — so the asymmetry costs nothing.

Which dissolves the earlier worry about one mechanism with a tier boundary
inside it: there are two mechanisms, with different actors, scopes and trust
models, and that is the answer rather than a compromise.

Named rather than hidden: this is the escape hatch research 011 warned about,
arbitrary code in the place hardest to remove later. It is bounded by being
bundle-only and by every action having to declare how it verifies itself, and
that boundary is the whole defence.
2026-08-26 23:56:39 +02:00
jschoubben 5b3d0ebd4f ADR 0046 — the installer fetches what it pins
The blocking question was where a container image comes from, and the version
that blocked assumed the machine might have no network. That assumption came
from the LAB: a scenario is a closed address space by design, which is what lets
two scenarios hold the same addresses without meeting. Production is not sealed
— a machine being adopted has a network, and one that does not is a machine
where very little works anyway.

So substrate.lock carries references, not payload: an image name and a digest,
fetched at apply time. A first node pulls from upstream because no mesh registry
exists yet; every node after that pulls from the mesh's own. The lab is the
exception and places images itself, the way it already places the host binary —
a property of a test environment, and letting it dictate the production design
would be the tail wagging the dog.

Pinned by DIGEST rather than tag. Reproducibility comes from pinning the
identity of a thing, not from carrying its bytes, which is what makes fetching
acceptable rather than a compromise.

ADR 0041 survives untouched, which was the point. "Copy it onto a machine and
run it" stays literally true — one binary, a few megabytes, which then fetches
what it was told to. Carrying images would have quietly redefined the property
that decision rests on.

Costs accepted and named: an apply can now fail because something is
unreachable, which a self-contained artifact could not, so it must fail legibly
— naming what it could not fetch and from where. And the lab needs a way to
place images into a machine that also has no container runtime, both of which
are lab-installation concerns and neither solved here.

Research 012's build-time-versus-apply-time reframing narrows accordingly: it
still holds for what a tailored installer contains, and no longer has to hold
for images.
2026-08-26 23:52:46 +02:00
jschoubben 60aea14935 Define the substrate, and answer 006's four-or-five conditionally
Same gap as the control plane: load-bearing and unpinned.

The substrate is what the control plane CONSUMES AND CANNOT GRANT ITSELF. Every
module needing a database asks provisioning for one; the control plane needs one
too and cannot ask itself, because it is not running yet. That circularity is
not an awkwardness to work around — it is the definition, and anything on the
wrong side of it must be raised by the bundle the host carries.

Which answers 006's open question in the honest form rather than with a number.
The identity provider is substrate only if the control plane DELEGATES
authentication — then it cannot serve anybody before the provider exists and
cannot grant itself a client. If it authenticates natively, the provider is an
ordinary hosted service. So the count follows from a decision not yet taken, and
asserting four was asserting that decision.

The test also rules out the tempting wrong answer: an identity provider, a mail
server and an analytics service are all infrastructure by any ordinary reading,
and none are substrate, because the control plane starts and runs without them.
Important is not the test.

Records why the bundle is pinned by hand — it is applied when no mesh exists, so
nothing can resolve a version or ask a registry — and why it must be
self-contained, which makes it an artifact built on a machine with a network for
a machine that may have none.
2026-08-26 23:39:11 +02:00
jschoubben 148395ca54 Define the control plane, which was used 79 times and defined nowhere
Nineteen files, seventy-nine mentions, no definition. That is how-we-build §5
failing on this repository's own vocabulary — ubiquitous language is checked,
not assumed.

The definition, and it is not arbitrary: the control plane is everything that
needs to know about MORE THAN ONE NODE. It follows from ADR 0037, which has the
host applying rather than deciding precisely because deciding needs knowledge
the machine does not have. So the line falls exactly there — writing a file is
the host's, choosing which nodes run the store is the control plane's, and
anything a single machine could answer alone does not belong here at all.

That last consequence is worth having: putting a single-machine concern in tier
2 is a mistake the tier rule will NOT catch, because the dependency direction
stays correct.

Also states what it is not — not the thing that changes machines, not a surface,
not the substrate, and not privileged on a node beyond what the declaration
vocabulary allows. And the property that makes tier 2 unlike the others: it is
itself a consumer, with the same requirements as any module, which is the
circularity the bundle exists to resolve rather than hide.

Scoped deliberately: this defines the term and does not design the contexts
inside it. Ten is the skeleton's claim rather than a settled list, and research
006 still asks whether the record belongs here or in the substrate.
2026-08-26 23:33:20 +02:00
jschoubben bea052753e ADR 0043 — what a declaration is
Stage 2 could not start without it. Three constraints already bound the shape
and between them they decide most of it.

JSON, because the standard library carries it and carries no YAML, and a YAML
declaration would put a third-party parser inside the one binary whose whole
argument is that it needs nothing — to gain authoring comfort in a document
generated by a machine and read by a machine.

An ordered list, because ordering is a DECISION. A host deriving order from
declared dependencies would be deciding the thing most likely to differ between
what the control plane intended and what the machine does. The control plane
knows what depends on what; it says so by saying when.

Unknown is refused, never skipped — an unknown version, type or field refuses
the whole declaration. A host that skipped what it did not understand would
apply most of a declaration and report success, which is 04-ISSUES/003 with the
declaration on the other side of the wire.

Complete for what the host OWNS, and only that. It removes what it previously
applied and is no longer declared, which it knows from the store rather than by
inference, and never removes what it did not create — a converger that treats
'not declared' as 'must not exist' deletes what the mesh never put there.

Two consequences arriving earlier than the build order suggested: the store is
load-bearing at stage 2, because nothing can be removed without knowing what was
applied. And a closed address space bounds the first vocabulary to what needs no
network, because a scenario has no route to a package repository.
2026-08-26 02:04:41 +02:00
jschoubben 92e8c74ce4 ADR 0041 and the build handoff for the node host
Building tier 0 forced the question "the one binary installed by hand" had been
carrying unexamined. A TypeScript host needs a runtime present before it runs,
so the thing installed by hand becomes two — and the second must be installed by
the means the host exists to replace.

So the host is a statically linked binary that requires nothing present, written
in Go. Rejected: a runtime installed first, which breaks the property the tier
rests on; and bundling the runtime into the executable, which carries ninety
megabytes to preserve a language choice and puts a young feature at the bottom
of the stack.

The argument that decided it is architectural rather than about taste. 0037
means the host never queries the mesh database and 0039 means it only receives
declarations, so the host shares NO code with any other tier — not a client, not
a schema, not the SDK. The language boundary falls exactly on a boundary that
already exists, and a second language usually costs duplicated logic where here
there is none to duplicate.

§8 gains a scope: it said "TypeScript throughout" when everything was a service
or a surface, and is now scoped to those with tier 0 named. Another sync owed.

Playbook 04 steps 2 and 4: repos.md records mesh-host as existing, the design
takes code: [mesh-host] and status: in-progress.
2026-08-26 00:16:07 +02:00
jschoubben b9facf9375 Design the node host
Playbook 02 step 3, on four recorded decisions. Tier 0 has one job — apply
declared state on this machine — and the six absorbed concerns are instances of
it, not additions to it.

Specifies the six parts and what each owns, and the two properties that make
apply trustworthy rather than merely present: every applier reads back, because
setting a value is not evidence the value took; and what was applied is recorded
after it works, never before, because a failed apply leaves the machine wherever
it reached and nothing must claim otherwise.

Build order is staged so each stage is verifiable in the lab before the next
exists. Stage 1 is profile and inventory — no control plane, no declarations, no
network — and it is deliberately the smallest useful thing, because `place:` has
nothing to place and the lab therefore raises empty machines. Stage 1 ends that,
and every later stage is tested by a lab that already works.

Stage 2 is the one that could invalidate the tier boundary: whether one host can
raise the substrate alone is Move 1's assumption and has never been proved.

Every decision the design rests on is given the test that asserts it, per 0034 —
including the dependency-direction lint, which is what makes "the host never
queries the mesh database" a rule rather than an intention.

Six things left open and named, including the one that host-size.md could not
measure: zero dependencies, but still six vocabularies.
2026-08-26 00:10:24 +02:00
jschoubben 4bf7a35568 Close the record on the lab
Playbook 02 and 04 were followed for the substance — decisions before design,
design before build — and skipped for the bookkeeping. This closes that.

004 graduates. Its one open item was "not yet stood up"; the lab is stood up,
and the substitution the effort turned on is now enforced by the validator
before anything is raised rather than left as a thing to remember. Its
certificate conclusion has a home in 01-end-to-end-testing and is designed but
not built — implementation is a third axis, and an effort graduates on its
conclusions.

One item leaves 004 without a home and is recorded rather than lost: the reverse
proxy does not set caServer, so it defaults to the production endpoint.

The two lab designs read `designed` while running in production of a sort, so
they become `in-progress`.

And the lab gets an as-is document, which it did not have. It records what runs
including the parts nobody would choose again: that `place:` is refused and the
lab therefore raises EMPTY MACHINES, that the drawing shipped with no design
document behind it, that a router is tagged as a machine for a reason found by a
bug, and that the integration suite raises two of five scenarios while both
faults found so far lived in the three it does not.

006 stays active, deliberately. Two of its open questions ARE the tier 0 design
— whether absorbing six concerns makes the host too large, and whether an
unprivileged node earns a place in the inventory. Playbook 04 is explicit that
an open question is a reason to research, not to build around.
2026-08-25 01:55:52 +02:00
jschoubben 8efa063f21 The snapshot question is answered by a test
The lifecycle design asked whether a scenario snapshot needs the machines
stopped. The integration test answered it on its first run: no, but they
must be flushed.

A snapshot captures disk and not memory, so a write still in the guest's
page cache is absent from it — not stale, absent. A file written seconds
before a snapshot did not survive the restore.

Flushing first buys write-durability. It does not buy
application-consistency: anything mid-transaction is still captured
mid-transaction, and that limit is now stated rather than left implied.
2026-08-24 22:26:47 +02:00
jschoubben eab4598494 ADR 0033: a router is scenery, not a node
ADR 0016 makes a lab node a virtual machine, and its reasoning is
fidelity: a node boots a stock image and runs the real install, so it has
to be a real machine or the thing under test is not the thing that ships.

That reasoning does not reach a router. Nothing under test runs on one, it
holds no identity, the mesh never installs anything on it, and no assertion
is ever made about its internals. It exists so packets behave the way they
behave in the world, which is the definition of scenery.

So a router is a system container. What it must reproduce is kernel
behaviour — translation, connection tracking, filtering, forwarding — and a
container has the same kernel.

Verified before deciding rather than assumed. In a plain unprivileged
container: ip_forward and ipv6 forwarding both settable, nftables
masquerade accepted and listed back, and the conntrack timeouts that
mapping_ttl depends on both writable. No privileged mode, no nesting, no
capability grants.

Rejected letting the hypervisor provide NAT, on a stronger ground than
speed: it makes the lab provide what the declaration is supposed to own,
and it cannot express a mapping that expires, a gateway that refuses to
forward, or policy between siblings. The model would shrink to fit the
tool.

The distinction is now load-bearing and has to stay legible: node means
something under test, scenery means something that makes the test real. If
the mesh ever installs anything on a router, it has become a node and this
record no longer covers it.
2026-08-24 01:28:08 +02:00
jschoubben e88b448145 The fix is real: 76x, verified. And how the lab installs on a clean machine
Snapshot 9.9s -> 0.13s. Restore 10.4s -> 0.80s. Three snapshots sharing
1.36 GB instead of costing 4.8 GB. The projected four-machine reset cycle
falls from ~90s, unbounded at worst, to ~15s dominated by a boot that
cannot be avoided. ADR 0029's inner-loop argument holds with copy-on-write
and did not without it.

The consistency matters as much as the speed: three consecutive snapshots
took 0.13, 0.12 and 0.13 seconds, against a dir second snapshot that never
finished.

One honest counter-observation recorded: launching onto the fresh
copy-on-write pool was slower, 20.2s against 14.3s, because the image had
to be unpacked into a pool that had never seen it. Paid once per pool, and
dwarfed by what snapshotting saves, but it went the other way.

Doing the measurement produced the answer to how the lab installs on a
clean machine, because both failure modes appeared while doing it.

Installed is not available: the daemon was present with units disabled and
no group. Issue 007.

Available is not adequate, and this is worse: with the storage tooling
absent everything worked and snapshots were seventy-six times slower.
Nothing failed, nothing warned. That is a variant the mesh has not
catalogued — its usual failure is reported success and did nothing; this is
reported success and did it seventy-six times slower, which no error
surface catches because nothing is wrong.

So the lab verifies CAPABILITY, never installation, and refuses to run
degraded rather than warning — a warning about a slow inner loop is read
once and ignored forever. Prerequisites may arrive from a mesh module or
from the lab's own bootstrap, and the second path is required rather than
convenient: a lab installable only by a mesh cannot host the development
of the mesh that installs it.

The lab is the second thing installed by hand, after the node host, and for
the same reason: something has to be first, and pretending otherwise
produces a circularity papered over by a script nobody exercises.
2026-08-24 00:14:57 +02:00
jschoubben 98bcd5cc49 Measure the lab's inner loop — it is too slow, for a fixable reason
The lifecycle design closed on a question that was measurable rather than
arguable, so it was measured. One virtual machine on a workstation with
hardware virtualisation and NVMe.

Raising: the launch call returns in 3.4s, the machine is actually usable
after 14.3s. The gap is a design constraint — raise must wait for the
second number, because reporting the first would be transport reported as
effect, which is the mesh's own recurring failure.

Snapshot: 9.9s and 1.6 GB for a 1.5 GB instance. A dir snapshot is a full
copy; nothing is shared. Restore: 10.4s, usable again after 20.1s.

The second snapshot exceeded two minutes and never completed. That is the
more troubling number: snapshot cost here is not merely high, it is
unpredictable, and a loop with a variable multi-minute step is one nobody
trusts.

Projected to a four-machine scenario, a reset-and-rerun cycle is about a
minute and a half at best and unbounded at worst, before any of the mesh's
own work begins. That is too slow for an inner loop, and ADR 0029's whole
argument — that making the bootstrap path the inner loop turns the
least-exercised code into the most-exercised — holds only while resetting
is cheap.

The cause is not virtual machines. Hardware virtualisation is present and
machines boot in fourteen seconds. It is that the daemon offers exactly one
storage driver, dir, which has no copy-on-write and therefore no cheap
snapshot. The btrfs kernel module is available; btrfs-progs is simply not
installed, which is the entire reason the driver is absent.

The copy-on-write comparison was deliberately NOT run, because running it
would mean installing a package by hand — which the rules forbid and which
would have made the measurement unreproducible. So the honest statement is
that the current configuration is too slow and the likely fix is known but
unverified, rather than that btrfs fixes it.
2026-08-24 00:07:22 +02:00
jschoubben a253afe020 Scenario lifecycle, and how two scenarios coexist
ADR 0032: a scenario is a closed address space. Every segment materialises
as its own isolated link belonging to one instance, so two scenarios raised
from the same declaration hold the same addresses and never meet. The
declaration keeps its literal addresses and they mean what they say —
allocating from a pool would have made them a fiction, so a scenario
reproducing a specific topology would stop reproducing it.

The constraint that follows shapes everything: the lab never reaches into a
scenario over IP. It talks to machines through the virtualisation layer's
own channel. If it reached them by address, the workstation would need a
route into each scenario, and two carrying the same prefix would give it
two routes to one destination — failing not with an error but by one
scenario's traffic arriving in another.

That also makes reachability an honest question. Can this machine reach
that one is asked from INSIDE, by executing on the first, rather than
probed from a workstation that is not on the network and whose opinion
would be a different question with a misleadingly similar answer.

The lifecycle itself: six verbs, of which raise and destroy are enough to
be useful and the rest are what make repetition cheap. Raising is
convergent rather than incremental, because a lab behaving differently
from the thing it tests teaches the wrong habit.

A failed raise leaves the wreckage standing. Tearing down on failure
destroys the only evidence, which is backwards — a scenario that failed to
raise is more interesting than one that succeeded.

Snapshots are whole-scenario. Per-machine would be cheaper and wrong: the
mesh keeps state spanning nodes, so restoring one machine while its peers
move on produces a mesh that has never existed, and faults found there
would be artefacts of the lab.

Closes the declaration's open question about running several scenarios at
once.
2026-08-23 23:53:00 +02:00
jschoubben e88a6df924 Public networks are unrelated, and routed rather than bridged
Caught in review: every public address sat in one /24, which made the
three of them look like one network. They are not. The internet is a very
large number of unrelated networks routing to each other, and a machine in
one is many hops from a machine in another with no shared broadcast domain
between them.

Putting them in one prefix would have quietly made four false things true
in the lab: machines resolving each other by ARP and talking directly, TTL
never decrementing, broadcast and multicast crossing between them, and any
two being adjacent.

The third is not hypothetical. The as-is layer records that mesh names are
deliberately not multicast names, after a delay and a one-node-only failure
mode. A lab where the internet is one broadcast domain would let a node
discover a peer by multicast that it could never discover in production,
and report success — the exact false green this effort exists to prevent.

So a scenario has one public segment per public NETWORK, each with its own
unrelated prefix, wired together through a router and never onto a shared
bridge. That is a property of how the lab wires them rather than a field
anyone sets, because no correct scenario has two public networks adjacent.

Addresses now spread across all three RFC 5737 ranges plus RFC 3849 /48s,
chosen to look nothing like each other, and a foreign private network uses
someone else's RFC 1918 range rather than a documentation one.

All three examples in the document rewritten, since two of them still
showed a single flat internet segment and contradicted the new rule.
2026-08-23 23:49:38 +02:00
jschoubben a873088140 A worked example: the whole model applied to an ordinary mesh
The shape research 004 identified — one machine with a routable address,
one publicly named but behind a household connection, one stationary on
that network, one that roams — written out with every field the model has,
in role names and documentation addresses.

It shows the ISP modem doing nothing, because in bridge mode it is a media
converter: it changes the physical medium and leaves the packets alone, so
it creates no IP-level fact and appears nowhere. In router mode it would be
a second gateway and publishing would need a rule on both, which is the one
case the model still cannot express.

It shows two segments sharing one gateway declaration, which means one
gateway machine, and a policy rule between them that is asymmetric because
useful ones almost always are.

And it shows what is deliberately absent. Research 004 recorded overlay
addresses, hub election and names for exactly this topology, and none of
them appear: a scenario must not state what the mesh is responsible for.
Given the declaration, whether a hub is elected, whether the NATed
machine's endpoint is learned, and whether the roaming machine re-forms
after moving are all observed rather than arranged. The absence is the
point.

The run at the end moves one identity through four positions — home,
foreign network, asleep, home again — against a foreign gateway whose
mapping expires in 30 seconds, which is why a number is there rather than
a boolean.
2026-08-23 23:43:09 +02:00
jschoubben b1874f1d0f Segment policy, shared gateways, and what the model leaves out
Asked whether a real setup is coverable — router, modem, access points —
the answer splits, and one part was a genuine gap.

Most equipment is invisible and the omission is deliberate. The test: does
the device change what an IP packet can do? A switch moves frames within a
segment. An access point bridges wireless clients onto one — a machine on
wifi and a machine on cable are the same machine to IP. A controller
configures equipment and has no packets of its own. Modelling any of them
adds a fixture with no fault to catch.

Two entries in that list do matter. A modem in bridge mode is a media
converter and invisible; in router mode it is a second gateway, which is
double NAT — expressible as nested segments, but publishing through two
gateways still is not, and that is now named as the one real absence.

And VLANs are segments, which exposed the gap: inter-segment policy was
inexpressible. inbound: is a HOST firewall, per machine. A segmented
router enforcing rules between networks is a different thing and blocks
traffic regardless of what the destination thinks — a node behind such a
rule cannot be reached even by a peer that knows exactly where it is.

policy: states it as a fact about a pair rather than a property of either,
defaulting to allowed and asymmetric by design, because the useful
configuration is almost always one-directional.

Segments may also share a gateway: identical gateway declarations mean one
gateway machine, not two, because that is what a VLAN-capable router is —
and two routers sharing an address would not work anyway.
2026-08-23 23:32:30 +02:00
jschoubben 274bd3b304 Close the missing axes — and address family changes the model
Address family was not a field. IPv6 usually has no NAT, so a machine
behind a household gateway is typically unforwardable on v4 and DIRECTLY
ATTACHED on v6, at the same moment. The three positions therefore apply
per family, and reachability is a property of (machine, family) rather
than of a machine.

The consequence is bigger than the syntax: 'can these two nodes reach each
other' stops being a yes/no question. It is asked once per family, and the
asymmetric answers are the interesting ones. A mesh treating reachability
as one fact per node reaches a peer over one family, fails over the other,
and reports whichever it tried. That distinction did not exist in the model
and would have been found by a failure rather than by reading.

Two fields follow from it. inbound: allow|deny became necessary because
with NAT unreachability was implied by topology, while a globally routable
v6 address is reachable unless something refuses — so refusing has to be
sayable or v6 addressing silently implies reachability. And nat: became a
list of families rather than a boolean, because a real gateway translates
v4 and routes v6 and a boolean cannot say that.

mapping_ttl closes the keepalive gap: a mesh holding a connection through
NAT without refreshing it works perfectly until the far side goes quiet
for longer than the mapping lives.

segments[].mtu closes the fragmentation gap: an overlay adds a header, so
a tunnel over a reduced-MTU path establishes a connection and then
silently drops large packets.

at: takes a list, so a multi-homed machine is expressible — which the
model already implicitly required, since a border machine sits on two
segments.

v6 uses RFC 3849 documentation space, the exact counterpart of the RFC
5737 rule and load-bearing for the same reason.

Remaining: nested forwarding and an address changing in place, both
extensible when needed. Path quality stays deliberately out — it changes
performance, not correctness, and modelling it makes a network simulator
rather than a fixture.
2026-08-23 22:59:31 +02:00
jschoubben b944904f1a Audit the scenario model for generality, and fix what it found
The question is not whether the model covers our mesh but whether it can
express any mesh. Audited against the axes a deployment varies along, with
the standard being every property that changes how the mesh BEHAVES rather
than every property a network has — bandwidth does not change correctness,
MTU does.

One real bug, now fixed. A segment with no gateway was read as the
internet, which made an isolated network inexpressible: a LAN with no route
out would have been treated as public and forced onto documentation
addresses. Segments now state kind: public or private, and a private
segment with no gateway is an island. A mesh spanning a site with no
internet is a real topology.

One modelling error, now corrected. The three positions were framed by
ownership — a gateway you control versus one you do not. The axis is
forwardability. Carrier-grade NAT is your own connection and is still
unforwardable, so it belongs with the café network. Gateways gain
forwardable:, independent of nat:, and publishing through an unforwardable
one is a declaration error because that is the constraint being reproduced.

Three genuine gaps recorded in priority order. Address family: cidr is
implicitly v4, and a v6-only node is not exotic — a mesh that assumes v4
fails there completely rather than partially, which makes this a second
world rather than a refinement. Expiring NAT mappings: without them
keepalive behaviour is hoped for rather than tested, and for a mesh mostly
behind NAT that is the fault that shows up after an idle night. MTU:
tunnels fragment, and a smaller-MTU path establishes a connection that then
silently drops large packets — the exact shape this effort exists to stop
shipping.

Latency and loss are deliberately out: they change performance, not
correctness, and modelling them makes a network simulator rather than a
fixture.

Also adds a NAT primer, because the three positions are consequences of it
and the document should not assume the reader already knows why a mesh
dials outward and never inward.
2026-08-23 22:48:46 +02:00
jschoubben e65e5809dc The scenario declaration gets a real network model
forwarded: [443] was the tell. It implied a destination-NAT rule while
never saying from which address, and the address is the whole point: a
household's public address is what a peer records as the endpoint when a
machine there dials out, and what a public name for a published machine
there resolves to. It was decoration in the old shape and is load-bearing
in this one.

The model now names three positions a machine can be in, because they are
genuinely different and the mesh has to cope with all three. Directly
attached, with its own routable address. Behind a gateway you control,
reachable only through a forwarded port at the gateway's address. Behind a
gateway you do not control, reachable not at all, with an apparent address
belonging to someone else's router that changes when the machine moves.
The third is the hard one and the one that breaks reachability
assumptions first.

A gateway now carries three facts instead of a boolean: the parent
segment, the address the world sees the network as, and whether addresses
are translated — so a routed range is expressible as well as ordinary
household NAT. published names the gateway it forwards through, which is
how a machine on a LAN that itself has a public address is stated, and
publishing on a foreign gateway is a declaration error because that is
exactly the constraint being reproduced.

Moving a machine between positions becomes a lifecycle operation rather
than a declaration: the same identity at home, then on a foreign network,
then asleep, in one run. Whether the overlay survives that and notices the
endpoint changed is observed, never arranged.

All three RFC 5737 ranges are now allocated a job — the internet segment,
a foreign network, and a spare — with private segments kept byte-identical
to production because those addresses mean the same everywhere.

New open question worth having: a real gateway forgets NAT mappings after
a timeout, and whether a scenario can say so decides whether keepalive
behaviour is testable or merely hoped for.
2026-08-23 22:42:27 +02:00
jschoubben a72fea5342 ADR 0031 and the scenario declaration
The lab provides the underlay; the mesh builds the overlay. This is the
boundary that decides whether the lab is worth having: a scenario that
assigns overlay addresses, elects the hub and writes peer configuration
certifies its own work — if the mesh's peering is broken, that scenario
still comes up green. The most valuable thing the lab can test is exactly
the part pre-building would replace.

So a scenario declares what a hosting provider and a home router would
provide: segments, which machine sits where at which address, what NAT is
between them, which ports are forwarded, which machines are detached. It
declares nothing about overlay addresses, hubs, peering, names or
certificates, all of which become outcomes to observe.

The declaration has four parts — segments, machines, place, snapshot — and
the two scenario classes differ only in place. That is what makes one a
strict subset of the other rather than a fork.

Research 004's most important finding becomes a format constraint rather
than a footnote: the routable segment must use RFC 5737 documentation
space, because the mesh decides public versus private by matching the
address, and a private range there makes the hub test as unreachable while
the mesh silently never forms. A segment without behind: is routable, and a
non-documentation address in it should be refused before anything is
raised — ADR 0008 applied to a configuration file, since the failure it
prevents has no error at all.

Four things left open, including the one that matters most: a lab machine
is always privileged, so the user and edge profiles have no scenario that
exercises them.
2026-08-23 22:33:52 +02:00
jschoubben 4d387d2998 Hand the lab design off to mesh-lab
Playbook 04: the design names its owner and flips to in-progress. code
moves from hal to mesh-lab, and decisions gains 0029 and 0030 — the design
now rests on three records rather than one.

repos.md marks mesh-lab as the one target repository that exists. The
other six remain the target, not the present, and saying so is the point:
a map that lists repositories which do not exist is a map that will be
believed.
2026-08-23 22:26:30 +02:00
jschoubben b4904fec7e The lab comes first, and its first scenario has no pipeline
The lab was designed around a module under test, with a scenario being a
complete mesh — forge, coordinator, cascade, verify. That is unusable for
building the new mesh, because all four are tier 2 and do not exist yet.

And research 009 had the sequence backwards. It placed the lab at phase B
as verification of tiers already built, but tier 0 is the component that
takes over a machine's packages, services and network. It cannot be
developed against a machine anyone needs. The lab has to exist before the
thing it will test.

ADR 0029 splits scenarios into two classes. The bootstrap scenario is
virtual machines, the host binary and a pinned bundle, with the verdict
coming from what the host reports about the state it reconciled. The full
scenario is the designed one. The first is a strict subset of the second —
same virtualisation, same networking, same lifecycle, stopping before a
control plane exists — so the second is reached by addition rather than
rework.

The consequence worth having: raising a node from nothing stops being the
least-exercised path in the system and becomes the inner development loop.

It also settles the runner's two jobs. Scenario lifecycle is needed
immediately, because something must materialise and reset a mesh before
anything can be written against it. Assertion execution waits for the full
scenario.

Corrects a stale claim in the design while amending it: it argued
scenarios were affordable with system containers and would not be with
virtual machines. ADR 0016 superseded that reasoning and the text had not
followed.

Issue 007: the lab's first requirement is installed and unusable. The
virtualisation package is present and explicitly installed; both units are
disabled, the operator is in no group, and the client reports the server
unreachable. Not issue 001 again — that is an install failing while
reporting success. This is an install succeeding when success was not the
point. A package is files; a capability is a running service and an
identity permitted to reach it, and the module model has no vocabulary for
the second.
2026-08-23 21:57:14 +02:00
jschoubben 93a1231e00 Retire the HAL name where it points forward
Skills take the hq- prefix: they are HQ process workflows, not mesh
workflows, and HQ is company-scoped now. hq-new-research, hq-graduate,
hq-new-issue, hq-diagnose, hq-amend-design, hq-handoff,
hq-sync-constitution, hq-status.

Forward-looking prose becomes Novox Mesh or simply the mesh — the root
README, AGENTS.md, the 00-META README, the mission's module example, and
one to-be document that addressed 'someone working on HAL'.

Three categories deliberately keep HAL, per ADR 0027:

The monorepo is still called hal on the forge. repos.md, every code: field
and every located-in: field name a repository that exists under that name,
and renaming them in prose would make them false.

The as-is layer and the research that measured it describe the system that
runs, and that system is called HAL. 124 modules, 9 daemons, a dead
containerised node — those are observations, not intentions.

Records 0001-0026 are immutable. A record says what was decided when it
was decided, and no record is edited for a name.

Also repoints ADR 0022's link at the renamed skill — a path fix, which the
immutability rule permits, not a change of meaning.
2026-08-23 21:26:09 +02:00
jschoubben daf3e17c32 self-hosting, provisioning and delivery efforts, and the dotfiles origin
The identity provider is settled as not-substrate: the mesh does not
require one, tier 2 authenticates natively, and it is a hosted service
like any other. Four substrate services, not five. The tier test's second
step gains the verb that matters — can the control plane START without it,
not function fully without it.

That verb answers the forge and the registries. They are not substrate and
they are not duplicated: the control plane starts and manages nodes
without a forge, it just cannot change itself. One gitea module, tier 4,
and the mesh's own instance is distinguished by what it is bound to rather
than by being a different module — the same answer as postgres, from the
same test. It also buys a property worth having: if the forge dies the
mesh keeps running.

Delivery needing them is not an upward dependency, resolved the way the
constitution already says to: tier 2 declares requirements, tier 4
provides implementations, the binding is data. The mechanism is
provisioning, and the new idea is that the control plane is itself a
consumer.

Self-hosting therefore becomes a state the mesh REACHES, not a
precondition. A first node comes up from pinned external artifacts and
re-binds to internal providers once they exist. Today's mesh assumes the
second state from the first moment, which is why the first-node path needs
a script that papers over an impossibility and is the least-exercised code
in the system. Made explicit, the transition is also reversible.

Research 007 and 008 opened for the two areas flagged as important and
complex, scoped from the weaknesses the as-is layer already documents
rather than started blank.

And the origin: this began as a dotfiles repository. The first two days
adopt dotfiles, add per-node overrides, and introduce service symlinking
with an ignore file. The flat one-directory-per-tool catalogue, linking
over copying, adoption of already-configured machines, per-node overrides
and the desktop modules are all inherited rather than chosen for a mesh.
That is the single most useful fact for anyone changing the catalogue, it
strengthens ADR 0018 — the case for links was never made for a mesh — and
it explains research 005's silent fifty: dotfiles-era entries for one tool
never shared a domain because they never had one.
2026-08-23 20:55:14 +02:00