Commit Graph
182 Commits
Author SHA1 Message Date
jschoubben 554f6bd7a4 A capability may carry a value, and adding one is not free
Recorded while building the seat detector. A capability is a named fact about a
machine: its presence gates an assignment and its detail can carry a value, so
"can this run here" and "what should it be configured as" are the same fact
read two ways. A verdict has always had a detail beside its yes or no, so
panel: oled needs no new concept.

Two things that keep the set honest, both worth writing down before anyone adds
the fiftieth capability. It must be detected and the detector must say how it
knows -- so nobody can add one they cannot check, which is the whole of issue
007. And detectors ship inside the host, which is one static binary, so adding
a capability means shipping a new host everywhere. That argues for a small
general vocabulary rather than a specific one.
2026-08-29 21:12:21 +02:00
jschoubben f140303257 A module claims; it does not list its rivals. And flavor is retired.
Three decisions, all Jochen's, and the first is the one that unlocked it.

Exclusivity is not a property of a module. It is a property of a singular
resource the module takes over. Two shells compete for nothing and any number
may be installed; two display servers both want the seat. So a module declares
what it CLAIMS, and two modules claiming the same thing cannot both be assigned
within that claim's scope.

Not "xorg conflicts with wayland". Pairwise exclusion has a property that only
shows up later: adding a third display server means editing xorg and wayland to
know about it. Every new module requires changing modules nobody who wrote it
owns, and the edits grow as the square of the count. With a claim the third one
says what it claims and nothing else changes anywhere.

Claims have a scope -- node, site, mesh -- which is not new. The mesh already
enforces exactly one hub with a unique index. Scope is that idea said once
rather than hard-coded per case.

And some conflicts need no claim at all: two modules declaring the same file or
binding the same port are visible from what they declare. A claim is only
written for the abstract ones.

A requirement with several answers is refused, never guessed. One candidate is
assigned silently because there was no choice to make; none is refused naming
what is missing; several is refused naming them. That is what makes a solver
unnecessary -- counting candidates has no surprising behaviour, and a solver
can be added later without changing a single manifest.

Flavor is retired. It was carrying three unrelated meanings: variants of a
thing, a subset of a module a node installs, and whatever the current system
does, which earned two knowledge-base entries about going wrong. A word with
three meanings cannot be reasoned about. What it reached for is two ordinary
things -- different modules providing the same thing, and one module with a
setting.
2026-08-29 21:00:13 +02:00
jschoubben 974985b3d1 Four things the lab found about the private network
All on the first three machines to actually run it, and all invisible from the
mesh's own state: the graph was right, the files were right, the services were
up, every node reported success, and the network did not work.

A running interface does not re-read its configuration, so a node joining left
every existing node carrying a network that no longer existed. A hub sharing a
site with a spoke was emitted twice, which WireGuard refuses. Two nodes at one
site that neither can be dialled were peered directly, so nobody opened the
path and the more specific route blackholed -- this document's own warning
arriving in its implementation. And Docker sets the FORWARD policy to DROP, so
a hub with forwarding enabled still carried nothing between its spokes.

The last one is the sharpest: the substrate at tier 1 silently breaks the
network at tier 2, and nothing in either tier's state says so.

None of these is reachable by reasoning, and each was found within minutes of a
real machine trying it. That is the argument for the lab in one line.
2026-08-29 18:03:39 +02:00
jschoubben 6bcf0e4f9f Issue 010 fixed: origins keep the bundle and the mesh apart
The store records where each resource came from and each origin removes only
its own. Verified on the scenario that caused it -- eleven resources raised,
enrolled, sent the same two-resource declaration, and the store, broker and
control plane were all still running. A later declaration dropping a resource
still removed it, so removal by omission survived the fix.

Two more faults found while fixing it, both the same shape. A report published
to a routing key nobody bound vanishes: the broker accepts it, finds no queue,
drops it, and tells the publisher nothing -- so nodes announced what they had
applied into a void. And publishReport was discarding its error, so a node that
could not tell the mesh looked exactly like one that had.

Reports are mandatory now, so an unroutable one comes back and is said out
loud, and the binding covers every key a node may publish.
2026-08-29 16:43:28 +02:00
jschoubben 594ea10b07 Issue 010: the first declaration destroys the substrate
Found in the lab, doing the ordinary thing: raise a first node, enrol it, send
it a declaration. Both declared resources applied correctly and every container
on the machine was removed -- the store, the broker, and the control plane that
had sent the message. The link died mid-sentence because the broker carrying it
had just been torn down by what it carried.

Nothing is behaving incorrectly. Apply removes what the store holds and the
declaration does not name, which is what reconciliation means. The fault is
that the carried bundle and mesh declarations share one store, so the host
cannot tell what this machine raised for itself before there was a mesh from
what the mesh told it to have.

It is invisible until those two meet, which happens exactly once per mesh: on
the first node, after enrolment, the moment the control plane first speaks.

The report says what is not the answer, including the tempting one -- having
the control plane send the substrate back. It cannot: it was never told what
the bundle contained, and the bundle exists precisely because there was no
control plane to ask.
2026-08-29 16:23:01 +02:00
jschoubben 02afb7516b What connecting to the mesh is, and what a node presents
Two things this record never said, both asked directly.

Connecting to the mesh is one outbound AMQP connection from the node to the
broker, held open. There is no second connection and nothing is ever dialled at
a node. Being in the mesh means that connection is up.

Two different things ride on it and conflating them is what made this murky. An
AMQP account, which the mesh issues per node at enrolment, answers whether the
connection is accepted at all -- per node rather than shared, because a shared
one lets any node consume another's queue, which is the shared-credential fault
this record exists to remove reappearing at the transport.

The node's own keypair answers which node is speaking, on every message. It is
not made redundant by the account: with only an account the control plane knows
who is speaking because the broker says so, and that is the same transitive
authority this record already refuses in the other direction. A compromised
broker could attribute reports to whichever node it liked.

So a node holds two things after enrolment -- a credential the mesh issued for
reaching the broker, and a key it generated that the mesh only sees the public
half of. Both are its own, neither reaches anything else.
2026-08-29 15:25:18 +02:00
jschoubben 004057d85c A node's identity is a keypair it generates. This was never open.
I have been treating "what a node presents to prove it is that node" as an
undecided design question for weeks, and blocking on it. It was decided.
08-connectivity says of the overlay keys: each node generates its own keypair,
the private key never leaves the machine, the public key is published to the
mesh -- and says explicitly that this IS ADR 0004's "a node holds its own
identity", applied. Nobody had applied it to the thing 0004 is actually about.

What caused it was a word. The lifecycle said a joining node receives its own
durable identity, which reads as the mesh issuing something, and then the
question is what. The mesh issues nothing. A node arrives holding its identity;
what it receives is being known. That line now says what happens: it presents
the one-time secret and its own public key, which the mesh records.

The rule above it then holds literally rather than aspirationally. The mesh
stores a public key, so a copy of the mesh's database grants nothing, and
compromise of a node really is compromise of only that node.

Also recorded, since it was asked directly: same principle as SSH, own key, not
the machine's SSH host key. Host keys are regenerated by reinstalls and image
clones, which would silently un-enrol a node; their lifecycle belongs to sshd
rather than the mesh; and a partial host has no SSH daemon at all, so an
identity scheme resting on one excludes a supported kind of node.

The good half of that idea is kept: the mesh knows every node, so it can
distribute host keys the way it distributes authorised keys, and node-to-node
SSH stops depending on trust-on-first-use.
2026-08-29 15:21:36 +02:00
jschoubben 5fd522b8da A node is a machine; the session is a feature of it
Correcting an overstatement from the previous commit, where I had written that
a node IS a conversation. It is not. A node is a machine inside the mesh, and
the session is one of the things running on it -- like the host, like any
workload.

That also dissolves the conflict I flagged as unresolved rather than needing
anyone to decide it. 0001 says a node does not authenticate to a model
provider, agents do. Still true: the session authenticates, and the session is
not the machine. The node does not think, something on the node does. I had
manufactured the contradiction by promoting a feature into an identity.

0001's summary row is corrected the same way, and says explicitly that neither
the node's session nor a hired worker makes the node itself a thinking thing --
both run on a machine, which is what leaves that line untouched.
2026-08-29 14:22:29 +02:00
jschoubben 066f14b5f8 A node is a conversation, and that is not the employee model
Moving this out of 0003 and out of its vocabulary. I had spent three attempts
fitting the node's own session into the agent-as-employee record, each time
bending hired, draining, reassigned and retired to cover something none of them
describe. 0003 is back to its original text.

It belongs in 0004, under what a node is, because that is what it is -- not a
program installed on a node but part of the node. It holds one session
permanently, anything in the mesh can message it, and it remembers across
callers and across weeks. Its system prompt is the engram, which is recorded
here for the first time despite running on every node.

Also recorded: it has its own narrower tool list, so it can go and look rather
than only report about itself; there is no authorisation between nodes, because
every node is the operator's own; and how a node passes a question on is its
own business rather than a protocol field.

Switched off it still answers, and that is the point of having an off state
rather than an absent one. A node with nothing there is a silence somebody has
to diagnose. A node that says it is switched off is not. Same rule the host
follows about a service that does not exist.

0001's summary is corrected too: it had one row for "agents", which is the
conflation being complained about. Two rows now. A node's own session and a
hired worker are built from the same parts and run on entirely different terms.

Left standing and NOT resolved here: 0001 says a node does not authenticate to
a model provider, agents do. A node that holds a session does. That is a real
conflict between what is recorded and what runs, and it needs deciding rather
than a fourth reconciliation from me.
2026-08-29 14:17:25 +02:00
jschoubben 079c488d5e Provisioned and immutable beats exempt
Replacing the framing I wrote an hour ago. I had the node's own agent sitting
outside the lifecycle as an exemption, which is a rule somebody has to
remember. Provisioned the ordinary way and constrained is a rule the system
enforces, and it is one row like any other rather than a category every query
listing agents has to special-case.

It also reads the original sentence more carefully. "Exempt from the hiring
lifecycle" is exempt from hiring, not from having a lifecycle. Its lifecycle is
the node's -- provisioned at enrolment, retired when the node is retired. Same
states, a different thing driving them, and no exemption needed.

The constraints are now the four nonsense states written as things that cannot
happen rather than as an argument: not retirable, reassignable or deletable
while its node exists; exactly one per node. And a distinction that was missing
-- its existence is immutable, its engram is not. Freezing the personality
would remove the way a node is configured.

Disabling is the better half of this. A node with no agent is a silence
somebody has to diagnose; a node whose agent is disabled answers saying so,
immediately, with no model invoked -- the queue is still consumed and the state
is the reply. That is the host's own rule about a service that does not exist,
applied one tier up: absence must never be indistinguishable from a failure to
answer.
2026-08-29 14:06:33 +02:00
jschoubben fd7f7557bd The node's own session, and why it is not hired
Answering a question that was asked three times and that I kept not answering:
should the node's session just be an agent per node, since otherwise the
functionality exists at two levels?

Same mechanism, different lifecycle. A persistent session, accumulating memory,
a system prompt, a scoped tool list, addressable by message -- identical, and
building that twice is the duplication the question was worried about. What
must not be shared is the lifecycle, because if a node's own voice were an
ordinary hired agent it could be retired, leaving a node nothing can talk to;
reassigned, moving one machine's mind onto another; hired twice, with no answer
to which one replies; or never hired, leaving a node mute. The exemption in
this record exists to make those four unreachable.

I had this backwards earlier today and said so out loud: I called "a node
itself is an agent of a kind exempt from the hiring lifecycle" a fossil of the
old model and recommended striking it. It is the design. And it does not
conflict with 0001 -- "the two agent rows per node merge" means one per node,
not zero. I read merge as delete and invented a contradiction between two
records that agree.

Engrams are recorded for the first time. They are in use on every node and
appear in no record, which is how a decided thing comes to look accidental.
The engram is the node's system prompt, and it is what makes one node's answers
recognisably its own rather than generic.

Also recorded: there is no authorisation between nodes, because every node is
the operator's own and a prompt from one is a prompt from them. The consequence
is stated once rather than left to be discovered -- the mesh boundary is the
security boundary, which is what puts the whole perimeter on the token and the
overlay.

And how a node passes a question on is the node's choice, not a protocol field.
A node may say who is asking or may simply ask, the way a person relaying a
question decides how to phrase it. That follows from the engram. The cost is
that there is no machine-readable chain of who ultimately asked; each node
still holds what it was asked and by whom.
2026-08-29 14:01:00 +02:00
jschoubben 88ba81e9c1 Agents reaching nodes is the capability, not a hole in it
Correcting what I wrote an hour ago. I had recorded node-to-node SSH as "not a
mesh function" and "a second control path through the back door", reasoning
from ADR 0004's rule that the host has no inbound control surface. That
conflated two different things and got the product backwards.

There is no node-to-node SSH to forbid. The actor is always an agent; a node is
only where it happens to be running -- ADR 0001 already says a node is a place
where an agent can run and that is the entire relationship. An agent hired onto
one node reaching another to do work is the capability the whole arrangement
exists to provide.

The credential is the agent's, in its own credential directory, which ADR 0001
already established. So a node's authorized_keys lists agents and never nodes,
and three things follow: no node holds a key reaching another node, so 0004's
"a node holds its own identity and nothing else" stays literally true; a
compromised node costs the credentials of the agents that were on it rather
than a way into everything; and who may reach what stays a mesh-wide fact,
which is why it is identity's.

The rule I misapplied is about how a node's declared state changes -- over the
broker, never by being dialled. An agent with a shell is not the mesh
reconfiguring a machine, it is what a person with a terminal has always been,
and this design already depends on that working: the overlay is the way back in
when a declaration breaks something. What such a session leaves behind is
drift, and drift is what reconciliation is for.

0001 also stops underselling the fourth layer. It read as "the layer the other
three exist to carry", which is true and flat. The value is that an agent can
work across a set of machines as though they were one -- centrally configurable
machines are ordinary; that is not.
2026-08-29 13:06:40 +02:00
jschoubben 918dc04916 What this actually is, and three things that were assumed
Four things settled by talking them through, all of which had been true in
somebody's head and written nowhere.

It is not a mesh in the peer-to-peer sense and will not become one. 0001 now
says what it is instead: machines linked by a private network, one node holding
knowledge of all of them, modules as the way anything is built and delivered,
and agents hired onto nodes to do the work. The word describes what machines
can reach, not how they are governed. "Master" overstates it the other way --
nothing needs that node to keep running, only to change.

0006 gains the option that would make it a real mesh, recorded as considered
rather than rejected by silence: every node holding the whole inventory, a
replication process, an elected master with promotion on failure. What settles
it is not the complexity but that it still would not deliver the name, because
application databases are not replicated -- so a genuine peer-to-peer mesh
means becoming a replicated database system for every consumer's data too. That
is a larger product than the thing it would support.

Also in 0006: three central roles, not one. Losing the control plane costs
change, losing the broker costs being told anything, and losing the hub costs
nodes in different places reaching each other at all -- which is operation, not
administration. Whether they are one node is not decided.

And SSH access is identity's. It appeared three times as something that uses
the overlay and never as something the mesh provides, which reads as settled
when nothing decided it. Nobody else could: the mesh is the only thing that
knows which humans and agents exist and which nodes they may reach. Node to
node SSH stays out -- the host has no inbound control surface by decision, and
nodes reaching each other that way is a second control path through the back
door.

0007 gains the requirement underneath all of it. Reachability was recorded as a
fact to track and never as a thing some node must have. The broker's node and
the hub must be dialable by every node at a stable address, or nothing can join
and a disconnected node cannot return. A mesh entirely behind NAT cannot be
raised. That is a precondition and it belongs with the others.

The link staying on the underlay is also argued now rather than asserted. At
join time it is forced; afterwards it is a choice, and the reason is that a
repair channel carried over the thing being repaired is not one. Moving it onto
the overlay, with fallback, is recorded as open with what it would have to get
right -- a WireGuard interface has no link state to test, and a silent fallback
is this repository's recurring fault in a new place.

0010 says in one line what was the intention throughout: the module system is
the CI/CD. Not a pipeline beside the mesh. Build, test, publish and deploy are
one reconciliation seen at four points, which is why a thing that cannot be a
module cannot be delivered.
2026-08-29 13:01:36 +02:00
jschoubben 5218b06c02 Fold the control plane's build decisions into 0006 and 0008
Back to 23 records. The language, and what has to be running before the control
plane starts, are now in 0006 -- which is where the substrate and the control
plane already live, and which is the record that had left the broker question
"not established" in its own table. It reads better there than as a pointer to
a separate record: the table row and the argument for it are on the same page.

The store mechanics went into 0008. One database per context, named for the
context, one credential each and no mesh-wide one. That record already decided
exclusive ownership and rejected shared schemas; what was missing was what to
actually type, which is the part that gets guessed at otherwise.

Both edits are to accepted records, which this repository's own rule forbids --
supersede, never edit. Recorded here so it is visible rather than silent. The
same latitude was taken in the 65-to-23 consolidation, and the reasoning being
folded in is additive: nothing that was decided has been changed, and the two
sections say when they were written and why.
2026-08-29 03:08:50 +02:00
jschoubben 82a3065f82 Tier 2 exists, and the token was missing a quarter of itself
mesh-control is built as far as it can honestly go: one context of seven,
inventory, with its schema and the command that applies it. The repos map and
the control plane design say so, and point at ADR 0024 for what it took.

Separately, and more importantly: this repository described the enrolment token
as carrying three things when ADR 0004 says four. The missing one is the
control plane's signing identity -- the reason a node does not have to trust
the broker it dials.

Without it the control plane's authority is transitive through the broker, and
0004 spells out what that costs: a compromised broker could forge declarations,
and since the host applies whatever the link delivers, that is the whole
machine. The record has the argument in full; the design doc had dropped the
conclusion.

Found by reading the two together while deciding what the control plane must
store, which is roughly the only way it would have been found -- both documents
are internally consistent and only disagree with each other.
2026-08-29 02:49:58 +02:00
jschoubben 84f4425fd6 The broker precedes the control plane, and it is written in Go
Two things found by trying to build tier 2.

The substrate design asked whether the message broker has to be running before
the control plane, and framed it as depending on whether the control plane's
own parts talk to each other over it. They do not -- it is one process -- so
under that framing the broker stays out of the bundle.

The framing cannot answer the question. What decides it is how the control
plane reaches a node, and the answer was already decided: only ever over the
link, and the link is the broker. So provisioning the broker would require the
broker. The first node does not escape this by being local, because it enrols
the ordinary way, by dialling the broker at the address in its token -- which
was deliberate, and worth keeping.

The bundle is two images now. The record says what that costs, including a
certificate the broker needs at a moment when there is no mesh to issue one.

The language had never been decided for tier 2. Go, for the same reason the
host is: the bundle pins this image by digest and runs it where nothing can
check it, so the image should hold the program and nothing else.

Also corrects something already built: the bootstrap created one database and
called it 'mesh'. ADR 0008 grants a context only what it exclusively owns and
ADR 0006 says the mesh database names a thing that will not exist. One database
per context, so one today, called inventory.
2026-08-29 02:32:46 +02:00
jschoubben 6a2b107fb8 Restore a consequence the consolidation dropped
I said nothing was lost when 65 records became 23. That was too strong, and
here is a counterexample: ADR 0046's consequence that the lab needs a way to
place images did not survive into the merged substrate record. The compression
kept the decision and dropped one of the things it implied.

It was not lost from the repository -- 04-ISSUES/009 had already picked it up,
which is why it was found at all. But the record no longer carried it, and the
record is where somebody would look.

Restored, now as a resolved fact rather than an open consequence: the lab
raises a registry inside the scenario, which is the real path since that is
what every node after the first pulls from. The digests it serves are its own,
and that satisfies the pinning rule -- what is required is a reference that is
exact and cannot move.

Worth recording the wrong assumption too, because it is what made this look
impossible for two days: I took "pinned by digest" to mean the UPSTREAM digest
had to be preserved. It does not. Any digest that is exact and immutable
satisfies the rule, and a registry assigns one.
2026-08-29 00:06:05 +02:00
jschoubben 087a8f4144 Close 009: a sealed machine now pulls by digest
The resolution was the one the issue predicted -- a registry inside the
scenario -- and it is the real path rather than a stand-in, since that is what
every node after the first pulls from.

The digests are the lab registry's own, which satisfies the pinning rule: what
is required is a reference that is exact and cannot move, and one this registry
assigned is both. That was the insight that unblocked it; I had assumed the
upstream digest had to be preserved, which is what made it look impossible.

The fault worth keeping is recorded in the issue: the read-back checked that
the catalog endpoint answered by matching the substring 'repositories', which
an empty catalog also contains. It passed on a registry holding nothing. This
repository's own subject, arriving in the tooling built to catch it.
2026-08-29 00:05:20 +02:00
jschoubben b4607dfc03 Numbers are identity; the reading order is a generated, checked index
Decided after measuring what renumbering actually costs: 96 references in code
comments across two repositories, none of which would have failed to compile.
They would have pointed at the wrong reasoning, which is worse than a broken
link because nothing reports it.

So a number identifies a record and never changes. It cannot also be a
position -- a position moves when the set changes, and an identity that moves
is not one.

The reading order moves into an index generated from each record's `topic:`.
Six topics, in the order somebody learns the system.

The index is WRITTEN rather than only generated on demand, which reverses what
this repository previously said. The reason it said otherwise is that a
hand-written index drifts -- but a reader looking at the folder on a forge sees
the folder, not a command, and the drift objection is answered by checking
rather than by refusing to write one. That is §5's own rule: a rule states how
it is checked.

Two checks, both confirmed to bite. index.py fails when the written order no
longer matches the records. records.py fails when a record has no topic or one
nobody defined -- the quiet failure being a record that vanishes from the order
rather than appearing in the wrong place.
2026-08-28 23:39:18 +02:00
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00
jschoubben e1febe8e0f Renumber the records 1 to 23
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18,
19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only
the archaeology of what used to be there.

Renumbered contiguously. Renames run in ascending order, so every target number
is already free and no two files ever collide.

The reference rewrite is one simultaneous pass rather than a sequence of
replacements. Numbers moved into slots other numbers were vacating -- the node
host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time
would have cascaded and silently pointed things at the wrong record.

Seven plain-text references survived the merges as prose rather than links,
naming records that no longer existed: the enrolment token, the link boundary,
what a declaration is, reachability, the repository structure. Each mapped to
the consolidated record that now holds it.

Verified rather than assumed: every [ADR NNNN](path) link now has matching text
and target, checked across the whole repository, and the checker passes.

Frontmatter `consolidates:` lists dropped -- they named records that are gone,
and each consolidated record already says in prose what it absorbed.
2026-08-28 23:28:34 +02:00
jschoubben 77f3a4cea7 Consolidate: 65 decision records to 23
Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.

  the node host          8 -> 1    applies not decides, depends on nothing,
                                   per operating system, root service, the
                                   launcher, episodic, what a declaration is,
                                   actions from the bundle only
  a node and how it joins 4 -> 1   what a node is, joining, the link as
                                   security boundary, the enrolment token
  modules and the graph   7 -> 1   everything is a module, no domain modules,
                                   three edges, provisioning, the core library
  substrate and control   6 -> 1   the test, seven contexts, one control plane,
    plane                          the authority is not a database, the named
                                   products, the pinned bundle
  connectivity            3 -> 1   a route is a grant, reachability declared,
                                   filter rules
  delivery                5 -> 1   reconciliation not a pipeline, artifacts,
                                   the three silos, a failed step, the verdict
  the lab                 5 -> 1   (earlier)
  how this repository     10 -> 1  (earlier)
    works

Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.

The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.

The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
2026-08-28 20:03:24 +02:00
jschoubben 5e83ac2c22 Consolidate: 65 decision records to 52
Jochen: a normal application has 3-5 ADRs, maybe 10 for a large one, and we are
at 65. Fair, and the cause is mine -- I recorded every FINDING as a decision
rather than every fork in the road.

Two merges, both cases where one decision had been split across many records
because it was taken over several days rather than at once.

0019 absorbs ten records about how this repository works: what it is and that
it is public, the folder flow, the two design layers, the issue front door,
status in frontmatter, playbooks, the naming rule, the product name. Those were
never ten decisions -- they were one, seen from ten angles as the repository
took shape.

0016 absorbs the five about the lab: a node is a virtual machine, a router is
scenery, a scenario declares the underlay, a scenario is a closed address
space, and the two scenario classes. Same pattern -- one design, split by the
order it was worked out in.

The consolidated 0019 also raises the bar for what earns a record, since that
is what produced 65: a record is warranted when there is a genuine fork -- a
direction reversed, an alternative that will be proposed again, something
contested. A finding is not a decision, and a bug is certainly not. Everything
else belongs in the design document where the reasoning is actually read.

The checker earned its place here. Deleting nine records left 13 dangling links
across the repository and it named every one, including in AGENTS.md. Nothing
was found by reading.

Remaining clusters worth the same treatment: the host (8 records), delivery
(5), modules (6), connectivity (4), substrate and control plane (4). That would
be 52 down to roughly 30.
2026-08-28 18:53:19 +02:00
jschoubben 10365f2eae Consolidate the design layer: one place per topic
Jochen: a jungle of specs that slightly contradict or patch each other, and
what matters is a working state rather than history. Both are fair and both are
mine.

Measured rather than assumed. 05-the-node-host and 09-the-node-lifecycle both
covered enrolment, the install commands, the unit file, the launcher and
reconcile -- I wrote 09 without taking anything out of 05, so the same things
were said twice and could drift apart.

Split by what each document IS. 05 is the component: what the host is, its
parts, the declaration vocabulary, the build order, how it is verified. 09 is
what happens to it: install, enrol, run, upgrade, retire. The whole "The
process" section left 05, and the unit file moved to 09 where installing is
described. 05 goes from 338 lines to 245 and now points at 09 rather than
restating it.

09 also carried a 105-line "Resolved" section -- six mechanisms framed as
"these were open and here is the answer". The content is needed; the framing is
history, and history is what makes a document read as a changelog rather than a
description. Renamed to what it actually is and the was-open phrasing removed.

Also added 10-delivery.md, which did not exist: four accepted decisions --
0054, 0063, 0064, 0065 -- had no design document at all, which is the specific
reason the delivery picture felt scattered. It is now one document covering
modules, the three edges, the core library, and how a change becomes a running
thing, with a table of what each property is designed against and what must
exist before it can be built.
2026-08-28 18:40:46 +02:00
jschoubben 9d091c81e0 A build edge, a core library that is a domain, and 0063 corrected
Three things from walking a real dev cycle through 0063, all of which Jochen
caught by pushing on where I had glossed.

0064 -- a build edge is a third kind. Research 011 established presence and
instantiation, and both are RUNTIME edges: they answer what a module needs in
order to run. Delivery needs a different question -- what has to be rebuilt when
this changes -- and that relationship is fixed inside an artifact rather than
negotiated when it runs. So the graph as designed could not drive delivery,
which is the real reason 0063 was not approvable.

It is derived rather than declared, read from what a module actually imports,
because a declared list and the imports it describes drift and the imports are
the true ones. The runtime edges stay declared, and that asymmetry is not an
inconsistency: a runtime edge is an intention somebody has, a build edge is a
fact about code that exists.

It also makes design quality measurable. A module with many inbound build edges
is one whose every change is expensive, and the current shared library is
exactly that -- nobody could see it because nothing drew the edges.

0065 -- the core library is the mesh's domain. Jochen disagreed with 0030's
"types, not behaviour" and was right: that guard is aimed at the wrong thing. A
library everything depends on is a hub whether it holds types or code, and the
fan-in is what makes a change expensive. So types ship with the module that
owns them -- trading one wide edge for several narrow ones -- and the core
library holds what is true of the mesh regardless of context, which research
011 already found: a module, a node, an assignment.

The test is "would this still mean the same thing in a context that had never
heard of the one it came from". A node does; a pipeline stage does not.
Domain-driven is the point rather than the label: "who else might want this"
always answers yes, which is how the current one grew.

And it changes the check for the better. "The build output contains no runtime
code" would have enforced a rule now withdrawn. Inbound build edges is a
measurement rather than a prohibition, and it is visible while a hub is forming
rather than after.

0063 revised on both counts, plus a third: I had written "the lab judges it" as
though that were a step. A lab run takes tens of seconds, occupies a VM, and
fails for environmental reasons -- and a shared-library change produces dozens.
One expensive non-deterministic gate fails both ways, and neither failure looks
like itself. Verdicts are now tiered, and a run that failed environmentally is
explicitly not a verdict.

0063 also now carries what must exist before it can be implemented, rather than
leaving that to be discovered.
2026-08-28 18:25:28 +02:00
jschoubben 4ab8a0507f Delivery is reconciliation, not a pipeline; research 008 closes
Jochen: don't rebuild the current coordinator, use it as a pitfall list. That
reframed the last open question rather than answering it.

0058 stopped deploy being a stage that pushes to nodes, and said plainly what
it did not fix: detection. A merge that created no pipeline, and nothing said
so. That is not a defect in the detector -- it is what happens when correctness
depends on an event ARRIVING.

0063 applies 0058's move one level up. The control plane holds what source
exists and what has been built from it, and builds the difference. A change
becomes a build because source is ahead of artifacts, which is a comparison
answerable at any moment. An event makes it fast; nothing makes it necessary,
so a missed webhook costs latency and cannot cost correctness.

The mesh becomes one idea at two layers: the control plane reconciles artifacts
against source, the host reconciles machine state against declarations. The
pipeline as a state machine disappears, and with it the stage list that a
verify step was once omitted from.

That reframing answered the three questions still open in 008, so it graduates
with all six closed. A deployed state is two comparisons rather than an event.
A verdict is about an ARTIFACT and gates whether it may be declared -- sharper
than the question expected. And "before self-hosting" mostly dissolves, because
a reconciler needs source and artifacts as bindings where a pipeline's stages
name their targets.

Four costs recorded, and one is a real risk rather than a trade: a reconciler
that cannot reach its target retries forever, and without something noticing,
the failure is silence -- the exact fault this removes, reintroduced elsewhere.
Also named: the run identity people actually use is lost, and "did my change go
out?" needs a replacement or this will be worse to live with than what it
replaces, whatever its properties.
2026-08-28 02:59:42 +02:00
jschoubben f728c3fd98 File 009: a digest-pinned image cannot be placed in the lab
Two accepted decisions collide, and testing found it rather than review.

0046 pins images by digest and has the host refuse anything unpinned. The lab
places images by exporting them from the workstation, because a sealed scenario
cannot reach a registry -- and that loses the digest, since a repo digest only
exists for an image a registry served. Measured: the load says 'Loaded image
ID:' rather than 'Loaded image:', and the image lands dangling.

So a tag is refused by the host and a digest is unusable in the lab. There is
currently no declaration the lab can raise that exercises the container shape,
which matters because the container shape IS the substrate -- every bootstrap
step past the runtime is one.

The resolution is a registry inside the scenario, and that is not a workaround:
0048 already names an OCI registry as substrate and every node after the first
pulls from the mesh's own. It also removes the lab's export-and-push mechanism
rather than repairing it.

0046 now carries a pointer, since its own consequence is where the collision
was predicted -- half of it is closed and the other half turned out to be
harder than 'not solved here' suggested.
2026-08-28 01:56:49 +02:00
jschoubben 9dc57b4712 Graduate 005; record what 0058 answered in 008
Continuing the sweep. Both were answered by records that did not cite them,
which is the same pattern 003 showed -- an effort stays active because the
decision that resolved it was reached from another direction.

005 graduates. Three of its four questions are answered: provider modules do
not group (0044), the ~50 modules that co-change with nothing stay as they are,
and 'group or leave' was never the right pair -- 0054 reframes it as authority
versus package. Worth noting the debt runs the other way too: this effort's
measurement, that reachability is the ONLY place modules genuinely co-change,
is what 0054 rests on and why connectivity is a context while nothing else
needed one.

Its fourth question moves rather than closes. Whether applications leave the
monorepo before or after they group is a sequencing question, so it belongs to
009-migration.

008 stays active, with its central question marked answered: the coordinator
converges nodes on a declaration rather than dispatching stages (0058). The
three-silo split survives with the third redefined. What 0058 explicitly does
NOT answer is how a change becomes a pipeline reliably -- detection is upstream
of everything it changed and remains the fragile input.
2026-08-28 01:40:02 +02:00
jschoubben 0a37d751e2 Graduate research 003; file the rescue that does not exist
A sweep of the nine active research efforts. 003 was answered five days ago and
nobody closed it -- the decision it asked for was taken without citing it, which
is how an effort stays `active` after being resolved.

Its recommendation is what the mesh adopted, and the match is exact rather than
approximate. "Run the daemons as containers, making Docker the supervisor for
everything" is ADR 0057. Its warning that a mesh-native supervisor inherits
fate-sharing "unless it sits outside the mesh's own process tree" is where ADR
0061 put the launcher. And its insistence that it cannot be all-or-nothing is
why the host itself is the one thing an init starts.

Its incidental finding does not graduate with it, so it is now issue 008: the
automatic node rescue the documentation describes does not exist. No unit
declares OnFailure=, nothing calls the rescue script on a timer.

That is worse than having no rescue. A rescue nobody wrote is a gap somebody
can see; a documented one that is absent is a gap nobody looks for, and the
documentation is read exactly when a node has failed and somebody is deciding
whether to intervene.

The issue names two honest resolutions -- implement it, or delete the
documentation and say a failed node needs a person -- and says the choice is
scheduling rather than technical, since the new host's recovery is built and
tested. It also says what would make the finding certain: it came from reading
the repository, and confirming it on a running node is the difference between
"no unit declares this" and "no unit in the source declares this".
2026-08-28 01:39:10 +02:00
jschoubben ba0d01788e 0062: a host may be episodic; 0060's Android gap closed
0060 named the gap and did not close it: everywhere else an init runs the
launcher at boot, and Android grants neither an init to register with nor
anything worth supervising, because a supervisor would be killed alongside what
it supervises.

Closed by narrowing what is required rather than building something. A host is
resident or episodic, and both are hosts. Being killed by the platform is
disconnection, which 0036 already made ordinary -- and every mechanism an
episodic host needs already exists because it was built for laptops that close.

A partial host can join a mesh and cannot be the first node, since every
bootstrap step is a shape it refuses. Its bundle says so.

Two consequences that are easy to miss: last-heard-from means much less on an
episodic host, so a healthy phone reads as a dead server unless the reader
knows which kind it is; and a declaration may take a long time to land, which
makes 0058's outstanding-versus-failed distinction load-bearing.

Still open, and in that order: what an Android node is FOR, and only then how
it is started.
2026-08-28 01:24:07 +02:00
jschoubben f1b1cd9aa0 Review: three ADRs no longer said what we had concluded
A sweep for claims overtaken by the last few days. Annotated rather than
rewritten, following the pattern already in 0049 -- what changed and why is the
useful part, and an accepted record should not quietly become something else.

0057's init section was wrong on all three of its claims. It said the host
needs FOUR things from an init; 0061 reduced that to one. It said every machine
the mesh targets already has systemd; Alpine does not, and it is the intended
first node. It said there is no second init to abstract over; there is now, and
the answer is still not an abstraction -- it is a four-line file per system.
What survives is the part that was always right: an init is not a dependency in
0041's sense, because it is not installed, it is what the machine already is.

0048 named Docker as the container runtime. It is now docker or podman,
detected rather than chosen -- because adoption keeps what a machine already
has, so naming one contradicted a rule already decided. That row is the only
one of the five that names two, and the record now says why.

0060 claimed the bundle is portable across operating systems. Its mechanism is;
its contents are not -- package names, unit names, service names all differ, so
an Arch host embeds an Arch bundle. That was my error, and it is the exact
confusion behind the question that found it.

The design layer had the same drift: 07 and 09 said "Docker" where they meant a
container runtime, 09 said systemd restarts the host after an upgrade when the
launcher does, and both install snippets assumed Arch. They now show Alpine and
Arch side by side, which makes the point better than prose did -- step 1
differs per system, step 2 never does.

Checked and NOT changed: 0047's "the vocabulary grows by one shape" is a claim
about the rate, not the count, and is still true. 0037 lists docker among tools
the host manages, which it does. 0041 says nothing about either.
2026-08-28 00:43:47 +02:00
jschoubben 66df0eb53e 0061: the launcher supervises; init is asked for one thing
The record said an init is asked for two things -- start at boot and restart on
exit -- which was half a change. It moved the give-up logic out of unit files
and left the restart in one, so init still decided when the host came back.

The launcher no longer execs the host. It supervises it, so restarting is ours
too, and init is asked only to run it at boot. There is an OpenRC script beside
the systemd unit now.

Records the cost honestly: not exec'ing means the launcher must trap the
shutdown signal and pass it down, because a supervisor that exits while its
child runs leaves the host to be killed rather than to stop.

And records what the implementation found: the counter counts consecutive
FAILURES, not starts. Counting starts meant a host that upgraded itself three
times rolled itself back, having worked perfectly every time -- because a clean
exit IS the upgrade path. That is now the second time a clean exit has been
mishandled, so it is called out as the thing to check.
2026-08-28 00:38:18 +02:00
jschoubben c557f99cba Record what testing podman actually showed
0060 said the container runtime was a separate decision. It is now made, and
the reasoning is worth keeping because it is the opposite answer to the same
question one paragraph earlier.

Abstracting service managers is lossy -- systemd and OpenRC are different
models and LoadState has no equivalent. Container runtimes converged on one CLI
deliberately, so almost nothing is lost: checked against podman 6.1.0, run,
rm -f and docker's own template syntax for state and labels all work unchanged.
Only the probe differs. So: a two-entry lookup, not an interface.

The difference that is NOT in the CLI is the one that would have shipped
silently. Podman accepts --restart unless-stopped, records it, and has no
daemon to act on it -- containers do not return after a reboot unless
podman-restart.service is enabled, which by default it is not. Every command
reports success and the effect does not happen.

That belongs in the declaration rather than the host: a node using podman is
told to enable the unit. Which is what made the service shape's missing 'boot'
field visible, and it is now built.
2026-08-27 23:59:05 +02:00
jschoubben e1ad39b500 Per-OS hosts, and an init asked for only start and restart
0060 -- the host is built per operating system. systemd and pacman are the Arch
host's implementation, not abstractions the mesh has to grow. They are not
independent choices: a machine has pacman because it is Arch, and the package
manager, service manager and packaging format arrive together as one decision
somebody made at install time.

Rejected abstracting them, and the reason is correctness rather than effort.
The service applier reads LoadState to tell "not installed" apart from
"stopped", which is what stops it reporting absence as success. An interface
spanning systemd and OpenRC degrades to what both express, and the lowest
common denominator is exactly where that fault lives.

Almost all of it is shared -- the vocabulary, store, apply loop, read-back
discipline, refusal model, bundle and link are portable. Two appliers differ.
And delivery was already per-OS, since a .pkg.tar.zst is an Arch artifact, so
this is the seam that already existed.

Android is the interesting case rather than Debian: no service manager, no
package installation, usually no root. Such a host implements file, directory
and action and refuses the rest -- the same refusal a host already gives an
unknown type, with a different reason. Those three are the portable floor.

The container runtime is deliberately left open: it is not an OS split, since
Arch runs docker or podman.

0061 -- the init is asked for start-at-boot and restart-on-exit, and nothing
else. Both are expressible in OpenRC, runit, s6 and an Android init.rc.
Counting failed starts and rolling back moves into a launcher, because that is
the one piece which must work when the host does not, and a script with a
counter can be tested where OnFailure= can only be hoped for. Supersedes 0059,
keeping its reasoning in full.

The checker found all six places citing 0059 and refused the commit until they
named the replacement.
2026-08-27 23:46:11 +02:00
jschoubben dcc4b8339c Say who consumes the broker and who writes the registry
Left implicit by the previous commit, which said the owning context writes
without saying what does the consuming.

The control plane is the consumer, and there is one of it. Seven contexts but
one deployable, so it is one process dispatching internally rather than seven
consumers racing -- which matters because the as-is records two consumers
accidentally sharing a queue and silently splitting the traffic, each getting
half of what it expected. With one consumer that cannot arise.

The broker is also the buffer while the control plane is down: nodes keep
publishing, messages queue, the control plane drains them on return. That is
what makes a single control plane tolerable -- an outage delays the mesh's
knowledge rather than losing it.

One consequence named because it will otherwise be discovered: an unbounded
queue grows until the broker's disk is full, and the broker is the component
every node depends on. The bound is per queue and undecided -- dropping the
oldest health report is obviously right, dropping the oldest declaration
acknowledgement is not.
2026-08-27 22:20:03 +02:00
jschoubben 19997d56c3 Approve 0057-0059, with four corrections from review
Not approved as drafted -- four things came out of checking them against each
other, and one was a bug that would have broken every upgrade.

The bug: 0059 specified Restart=on-failure while 0057 has the host restart onto
a new binary by exiting CLEANLY. on-failure does not restart a process that
exited zero, so every upgraded node would have been left stopped, having
successfully upgraded. Found by reading the two records against each other
rather than by either alone. Now Restart=always in all three places that
mention it.

The host cannot run in a container, and the reason is decisive rather than
stylistic: step 0 of the substrate bootstrap installs the container runtime, so
a host inside a container would need the thing it exists to install. It would
also break 0041 -- copy it onto a machine and run it stops being true when the
machine must already have a runtime. Everything above tier 0 is a container;
the host is not. That split is the tier boundary, not an inconsistency.

systemd is named rather than abstracted. An init is not a dependency in 0041's
sense: 0041 is about what must be installed before the host works, and an init
is not installed, it is what the machine already is. The unit file is the only
systemd-specific artefact and it belongs to the package, so a machine with a
different supervisor ships a different package.

The mesh is a watchdog, and my first draft was half an answer. Recovery must be
local -- nothing dials a node, and a host that cannot start cannot report. But
detection is the mesh's, and a local supervisor structurally cannot do it: it
sees one process failing and cannot tell a broken machine from a broken
release. Only something watching every node can, and that distinction decides
whether the response is "fix this machine" or "stop shipping this version". So
a host rollout is staged -- a few nodes, wait for heartbeats, continue or stop
on silence. Local rollback still needed, because the canary nodes break and
because a node offline during the rollout gets the declaration later with no
batch around it.

The first declaration is the overlay and nothing else. Forced, because a node's
address and peers are assigned rather than chosen. But also the way back in: a
node reachable over the overlay can be fixed by hand if a later declaration
breaks it, and a large first declaration risks a node that is broken and
unreachable at once.

Also stated plainly, because it reads as a contradiction: nodes reach each
other over the overlay and every node consumes from the broker; what 0039
forbids is an inbound CONTROL surface, not reachability.

And in 06: no node holds a credential to any control-plane store, for reads or
writes. Four ADRs already say this separately and none of them said it in one
place. Nodes state over the broker; the owning context writes. With a note that
most high-frequency writes are observability's, not the registry's -- routing
logs into the registry would be the shared-schema mistake arriving through a
door marked performance.
2026-08-27 22:12:12 +02:00
jschoubben 605c9fd441 Changes are pushed, not polled; and a stuck host rolls itself back
Two corrections and one new decision, all from Jochen catching things.

Pushed, not polled. I described updates as landing "on the next reconcile",
which reads as polling and is not the design. A declaration arrives as a
message on a link that is already open; the host applies it then. Polling over
an existing connection would be slower to land AND constant traffic to learn
nothing.

The timer is for drift and nothing else, and it cannot be replaced by an event
for a definitional reason: drift is change the mesh did not make -- somebody
edited a managed file, a distribution upgrade replaced a config -- so nothing
will ever publish a message about it. Only looking finds it.

Separated the heartbeat from the reconcile timer, which I had been conflating.
They point in opposite directions and answer different questions: the timer
looks at the machine and asks whether it still matches; the heartbeat reports
upward and is what makes silence mean something. A node with nothing to do
sends nothing, and without a heartbeat that is indistinguishable from a node
that stopped.

0059 -- a host that cannot start is rolled back by the service manager. I had
left this open on the grounds that recovery meant the host judging its own
health. That objection does not survive being asked properly: a keepalive is
something else judging the host. The watchdog must be local, because nothing
dials a node and a host that cannot start cannot report -- so it is the service
manager, which is already there.

The failure it prevents is sharper than "the node is down": a host that will
not start looks exactly like a machine somebody switched off, which is the one
condition this design has deliberately decided not to alarm on. So a bad
release reaches every node, each goes quiet, and the mesh reports a fleet of
sleeping laptops.

Confirmed means started and completed one reconcile -- deliberately not "the
link is up", or a laptop on a train would roll itself back. The rollback is a
script shipped by the package, not a host subcommand, because a binary that
will not start cannot be its own recovery. It rolls back once: a second failure
means the machine is the problem, not the binary.

Also refined the records checker, which produced a false positive: a proposed
record may extend another proposed one, because decisions are drafted in chains
and the alternative is marking things accepted to satisfy a check. An accepted
document resting on a proposed record still fails, and that was verified.

0057, 0058 and 0059 are all proposed.
2026-08-27 22:04:26 +02:00
jschoubben aeea2a9f9a Resolve the host lifecycle's open items, and say how the host is delivered
The upgrade question turned out to be a delivery question, so 0058 answers
both.

Today's third silo runs once per node and sends each one a command to install
and start. That is where the as-is records a package install that 404ed from
every mirror while the job went green, an image pull failure that did not fail
the deploy, and a verify stage that was built and never scheduled because it
was missing from a list.

The shape underneath all of those is that the thing reporting success was not
the thing doing the work. Meanwhile ADR 0037 has given every node a component
that applies state, reads back and reports -- so two mechanisms now change a
node and only one checks its work.

0058: a pipeline ends when the declaration is updated. Deploy stops sending
commands to nodes and becomes one write. The host applies it on its next
reconcile, and the host cannot report success it did not verify. The verify
stage disappears as a stage, which is the point -- verification stops being a
step that can be left off a list.

A pipeline result now means "the declaration is updated, and here is which
nodes have applied it". It does not wait for every node, because a node may be
legitimately switched off for a week. Outstanding is reported separately from
failed, since conflating them is how the old system produced a stall with no
error anywhere.

The host is delivered by exactly this path and needs no new resource type: a
`file` writes the package manager's config pointing at the mesh's repository, a
`package` names the version. Added a step I had missed -- before exiting for a
restart, the host runs the new binary once. A package can install something
that does not execute here, and that turns "the node never came back" into "the
apply failed and said why".

Six open items resolved: re-enrolment is decided when the token is issued and
revokes the previous identity; the mesh keeps a recovery copy of what each node
reports it owns, which un-strands the orphans; last-contact is reported with no
threshold, because a laptop off for three weeks is doing nothing wrong;
adoption always completes but a failed line makes a node ineligible for
assignment; a briefing is a structured document whose outcome is computed from
its lines; and the token is printed once and carried by hand, which is the
property that makes it worth anything.

Still open and named: automatic rollback of a host version that will not start.

0057 and 0058 are both proposed.
2026-08-27 21:53:38 +02:00
jschoubben 2204b01909 Design the node lifecycle end to end
The host was described as a component and never as something that runs for
years on a machine somebody else also uses. 09 covers every state a machine can
be in and every transition between them.

Four states: unmanaged, hosted, enrolled, disconnected. Only the last two are
nodes, and they are the same node in two situations. `hosted` -- the host
installed but never told which mesh it belongs to -- had no name before and is
where a machine sits between the two adoption commands.

Things that were unclear and now are not:

The first node walks the same path in an unusual order: reconcile from the
bundle, the control plane it just raised issues a token, enrol against it. Its
specialness lasts two commands. A side effect worth having -- enrolment is
exercised on node one, rather than being written and first used on node two.

Enrolment reports profile and inventory BEFORE the control plane decides
anything. The profile is the input to that decision, not a diagnostic; the
control plane cannot decide what a machine should run without knowing what it
can run.

Rebooting mid-apply is safe by construction. The store records each resource
after it worked, so a host that dies half way through comes back and applies
the rest. The rule that stops the host lying about what it did also makes it
crash-safe.

Retiring splits in two. Graceful is a final empty declaration. A node that is
gone will reconcile its last declaration forever -- the honest consequence of
making disconnection ordinary. The answer is not to make the host expire but
that the node holds nothing that outlives revocation: every grant is a per-node
credential revoked at the provider. A lost node keeps running and stops being
able to reach anything. Said plainly rather than implying the mesh can switch a
machine off, which it cannot and should not.

Losing the store is quiet and permanent, so it gets its own section. The host
re-enrols and re-applies fine; what does not come back is removal, because
resources it no longer has a record of become unowned and sit there
indefinitely.

Also corrects 0057, which said the mesh must not upgrade the host at all. That
conflated two acts. Replacing the binary is safe -- Unix keeps the running
inode. Stopping the unit is not. So the host may apply a package naming itself,
and restarts by finishing its apply and exiting cleanly, letting the supervisor
start it on the new binary. It never asks the service manager to restart it.
That makes a fleet-wide host upgrade an ordinary declaration, which the first
draft gave up on.

0057 remains proposed.
2026-08-27 21:16:44 +02:00
jschoubben 3ab11c96ef Say what the host process is: a root service, installed as a package
The design described what the host does and never what it is at runtime. The
words daemon, long-running, interval, poll and heartbeat appeared nowhere in it
or in the relevant decisions. What exists is a command that runs and exits;
what the design needs is a process holding a link. Nobody had written down that
those differ, so several questions had no answer.

0057 settles them. It runs on every node -- the host is what makes a machine
managed, so a machine without one is not a node. Root, because no useful subset
of the job is unprivileged. A systemd unit, because something must survive a
reboot to hold the link.

It never manages its own unit. The temptation is obvious and it ends with a
host stopping itself half way through an apply, leaving a machine with nothing
running to fix it. The installation owns the host; the host owns everything
else.

Installed as a package, with a tarball as the floor. The package carries the
unit file, the state directory and an upgrade path, which a bare binary does
not. But the mesh's package repository is hosted on the mesh, so any route that
needs the mesh to install the thing that joins the mesh is a circle -- the
tarball is the path that must never acquire a dependency.

Reconciles on start, on a declaration, on a timer and on reconnect. The timer
is the one easy to leave out, and without it `owned` reports what the host
applied rather than what is there -- ADR 0035 violated by omission.

The records checker caught this commit on its first attempt: 05 listed 0057 in
its frontmatter while 0057 is still proposed, and a to-be document may not rest
on an unaccepted record. The section now says so in the body instead.
2026-08-27 21:06:00 +02:00
jschoubben 2330d74c1b The host's vocabulary is complete; 05 and 07 said otherwise
All six shapes are built. 07 still said the last three did not exist, and 05
still described stage 2 as having built three of six.

Records what the lab still cannot do, because that is now the only thing
between here and an end-to-end substrate bootstrap: a sealed scenario cannot
fetch an image and its machines carry no container runtime, so package,
container and action were verified against a real machine instead.
2026-08-27 20:36:58 +02:00
jschoubben 03874f3fe2 Add a structural check over HQ's own records
Nothing in this repository was verified by anything but reading, which is how
a superseded decision stayed live in the constitution and in the to-be README
at the same time. Both were found by a person looking, and nothing stopped a
third.

Five checks: links resolve; `decisions:`/`extends:` name records that exist and
are accepted; a governing document citing a superseded record must name its
replacement in the same paragraph; supersession is symmetric; filename number
matches heading number.

Each was made to fail before it was made to pass. The live-citation check was
verified against a reconstruction of the actual incident -- the to-be README
citing ADR 0017 as live guidance -- and reports it with file and line.

It found one thing nobody had noticed: ADR 0018 never declared that it
superseded 0011, though 0011 has named 0018 as its superseder since August.
Fixed.

Deliberately not checked, and said so in the README: 02-DECISIONS and
01-RESEARCH may cite superseded records freely, because a decision record
discusses history and research records what was observed. 00-as-is may rest on
one, per 0056. Flagging those would put noise on correct documents, and a check
that cries wolf gets suppressed -- which costs more than not having it.

Two bugs found by running it: the frontmatter reader iterated an inline list as
characters, and the as-is exemption was missing entirely.
2026-08-27 20:18:35 +02:00
jschoubben e1f4c7d9e0 Approve 0054-0056, apply them, and fix the two smaller findings
0003 is now superseded by 0056. Nothing is left proposed.

Applied:
- 06 corrected from ten contexts to seven plus the api, each row now stating
  why it passes the more-than-one-node test. work, knowledge and stream are
  named as mesh-hosted rather than dropped; `ai` folds into config; `record`
  is deferred explicitly rather than listed. Its frontmatter now cites 0055.
- how-we-build §4 amended per 0054, and the derived page republished by
  playbook 05.

The sync found the drift the playbook exists to catch: the published §4 and
the source did not say the same thing. The source said "four accidents, not
four boundaries"; the published page said "one intent expressed four times",
and only the published page carried the scope caveat. Same rule, two texts,
already diverging. Verified the republish by reading back -- the new rule is
present and the old section's body returns nothing -- rather than trusting the
success message.

The two smaller findings:
- 0051 separated the transport identity from the declaring authority. It said
  the token carries "an address" and "the identity to expect" without saying
  what the node dials. It dials the broker, so pinning only that would make the
  control plane's authority transitive and let a compromised broker forge
  declarations -- which, since the host applies whatever the link delivers, is
  the whole machine. The token now carries four things, and declarations are
  signed and verified per declaration. Cost recorded: rotating the signing
  identity is fleet-wide.
- 0026 no longer restates 0022's rule about generated views. 0022's own words
  are "prose does not restate status; one place, and two is one too many",
  which is what 0026 was doing to it.
2026-08-27 02:21:34 +02:00
jschoubben f49d177a31 Draft three records for the contradictions the review found
0054 -- things that change together share an authority, not a package.
The constitution instructs agents to group "how a node is reachable" into one
module, citing superseded ADR 0017; ADR 0044 says there is no networking thing
to install. Since the constitution is injected where work is decided, the
superseded rule is the one actually steering work. The observation behind it
was right -- research 005 measured that reachability is the only place modules
genuinely change together -- but the conclusion was wrong: tight coupling means
a shared authority, not one artifact. wireguard and traefik deploy to different
node sets, so the merged module would be assigned where half is unwanted.
Requires amending how-we-build and republishing the derived page.

0055 -- the control plane is the node-coordinating contexts.
Three context lists were in circulation (0015 says nine, 06 says ten, the
README said eight) and none was decided. Research 006 said explicitly that the
change "belongs in a new record -- not written here", and the design used the
list anyway. Reconciling them shows `stream` and `ai` were dropped with no
reasoning at all. Applying 06's own test -- needs to know about more than one
node -- gives seven contexts plus the api, with work, knowledge and stream as
hosted applications and `ai` folded into config as an ordinary grant. The
record defers rather than lists.

The cost is stated rather than reassured away: a board composing across the
boundary reads more than one interface. That was raised before as "only moves
the problem up a layer", and the answer is that 0045 already requires surfaces
to read interfaces rather than stores -- what changes is the count, not the
kind of work.

0056 -- the authority is the control plane, not a database.
Every clause of 0003 has been decided against in four separate records and it
is still accepted and cited as live. The error underneath is the same category
error 0054 corrects: "source of truth" named a storage location when it meant
an authority, and once the store is the answer, shared schemas follow. The
half that was right -- the repository defines what exists, the mesh defines
what runs where -- survives untouched. Best consequence: the cache mode
disappears, so a node that has not heard from the mesh is no longer
indistinguishable from one that has.

0003 is left accepted until 0056 is.
2026-08-27 01:35:37 +02:00
jschoubben ef5dd0751b Approve 0049-0053; drop a to-be item superseded by ADR 0044
The 'domain grouping' item cited ADR 0017 as live guidance. 0044 superseded
it -- there is no domain module to group into, so there is no domain list to
settle.
2026-08-27 01:00:29 +02:00
jschoubben ccbbfa9c8a One node runs the control plane, and nothing takes over
Closes the two open questions in 06 and 08, which turned out to be one
question: how many control planes run, and what happens when the hub is down.
Both were drifting toward redundancy by default -- a standby plane, a second
hub, an election to pick between them. That is not one feature but a property
every layer must then honour, and each layer gets it wrong independently.

Not wanted, and not needed. A handful of machines with one node hosting the
registry is not a distributed system.

The argument for why this is sound rather than merely cheap is that the design
already tolerates it by construction. ADR 0036 makes reachability state rather
than class; the host reconciles from its own store (0043) and never needed to
ask anybody to hold the state it was last given. So the control plane being
down is not a new failure mode -- it is every node in the ordinary disconnected
situation at once. What is lost is change, not operation.

No node holds a contended role: the control plane is assigned like any other
module, and the overlay hub is declared (0050). No promotion, no quorum, no
fencing, no split brain, no replicated store, and no "which node is
authoritative" recurring at every layer.

Two consequences stated plainly rather than buried. The control-plane node is a
single point of failure -- deliberate, and said out loud so it stays
deliberate. And recovery is restore rather than failover, which makes backup
the availability story rather than hygiene.

The sharpest one is the clock: the control plane owns certificate issuance
(0049), so an outage outlasting a renewal window expires every public name.
That bounds how long recovery may take, and nothing measures it today.
2026-08-27 00:55:10 +02:00
jschoubben 4e80820e2f Design connectivity in full: overlay, resolution, exposure, filtering, certificates
Written as one document because the five are one design. They share inputs,
they must agree, and every one of them today is computed in a different place
by a different module from a different copy of the same facts.

The through-line is that none of the five can be answered by a machine alone,
so all five are decided centrally and delivered as `file` resources. That costs
no new host vocabulary and removes both remaining direct database connections
from nodes -- wireguard and traefik are the only two, and both are connectivity.

Three decisions fall out, all proposed:

0050 -- reachability is declared, not inferred from an address. The RFC1918
regex is wrong for carrier-grade NAT (100.64/10 tests as public, so an endpoint
is written to an address nothing can reach), wrong for IPv6, and wrong for a
routable address behind a closed firewall. The lab needing TEST-NET-3 to
satisfy the regex is the same bug from the other side. Also kills hub election
by address prefix, which fails silently and makes renumbering an outage.

0051 -- the enrolment token carries where the mesh is and how to recognise it.
Closes two circles with one mechanism: verifying the mesh needed the CA, and
obtaining the CA meant trusting whoever handed it over; and a node had to reach
the mesh before it could resolve any mesh name. An address plus a fingerprint,
carried out of band, resolves both -- and closes the CA question 0049 deferred.

0052 -- a filter rule names its source. `scope:` is declared in five manifests,
is part of no rule type, and is referenced by no code, so those manifests
appear to restrict ports and restrict nothing. Removed rather than implemented;
the general fix is refusing unknown keys, which the host already does and
manifests do not.

Also corrects two claims in 0049 asserting wireguard was already handled.
Research 006 says both modules still reach upward; neither is.
2026-08-27 00:36:49 +02:00
jschoubben 8d9282d86b Resolve the ingress gap: a route is a grant
ADR 0048 named ingress as an unclosed hole -- nothing said what terminates
TLS, how a public name reaches a container, or which tier owned it. Resolving
it needed no new concepts, which is why it survived: nobody had applied the
rules already written to it.

Ingress is not substrate. The control plane does not need a route to start,
and no node needs one to reach it -- the node dials out and has no listening
control surface. It grants itself a route afterwards, like a bucket.

A route is an instantiation edge under ADR 0044. The direction mirrors a
database -- the consumer supplies a target and receives a name rather than
credentials -- but it is the same edge.

The substantive finding is that exposure is three facts at two scopes: name
resolution and certificate issuance need to know which node is publicly
reachable, and only the proxy mapping is a single machine's business. That is
why it belongs to the connectivity context, and why Traefik doing all three on
the node is wrong.

Which matters beyond tidiness: research 006 counted traefik as one of two
modules opening a direct Postgres connection, reading nodes and mesh_ca. That
violates 0037, 0045 and 0039 at once, and is why every node permanently holds
a credential to the control plane's database. Deriving the config centrally and
delivering it as `file` resources removes it, costs zero new host vocabulary,
and closes the set 0039 identified -- wireguard was the other.

Left open deliberately: the mesh's internal CA is the other thing traefik
reads, and it belongs to the link's mutual authority, not to exposure.
Conflating the two is what made the gap hard to see.

Also fixes an inconsistency from the previous commit: 06 still claimed the
virtual host was raised from the bundle.

Proposed, not accepted -- for review.
2026-08-27 00:22:52 +02:00
jschoubben 4d19e93900 Name the substrate's actual products
The design layer described every service by role and never once by name:
Postgres appeared in zero design documents. That was over-application of the
research rule "never identify the mesh it observed", which is about node names
and domains, not software.

Two things were actually broken by it. substrate.lock pins images by digest and
a digest belongs to a named image, so the bundle could not be written from the
design. And a reader could not tell a settled choice from an unexamined one --
"a relational store" reads identically either way.

ADR 0048 names them: PostgreSQL, LavinMQ, MinIO, an OCI registry, Docker. The
argument for each is continuity, which is a real argument -- replacing a
substrate service migrates the mesh's own state. Role and product are now both
written, because the design depends on the protocol while the installer needs
the product.

Also separates two questions the substrate doc had merged: being substrate and
being in the bundle. Only Postgres must precede the control plane; the rest are
substrate by role and ordinary by delivery. Whether the bus joins it is left
open, because it turns on the control plane's internal shape.

Names the forge as Gitea, and records ingress/Traefik as an unclosed gap rather
than a naming one -- nothing says what terminates TLS or which tier owns it.

Fixes a miscount: the host's bootstrap vocabulary is six shapes, not five.
2026-08-27 00:11:38 +02:00
jschoubben c631cbd07c The bootstrap starts a step earlier than recorded
Asked whether postgres has to be installed, and the answer exposed a missing
step. The store is a container, so something must run containers before anything
else happens — and a container runtime is a PACKAGE, not a container.

Step 0 is where several threads meet. It is what the host's capability detection
already reports, and the first use of that report by something other than a
person. It is adopted rather than installed when the machine already has a
runtime with configuration somebody chose. And it is a package, needing the
machine's own package manager and a network, both of which ADR 0046 permits.

So the host's bootstrap vocabulary is six shapes: package, container, file,
directory, service, action. Stage 2 built three of them.

The node host design now names which three remain and why the lab cannot yet
exercise them — a sealed scenario fetches nothing and its machines carry no
container runtime, which is lab-installation work rather than a constraint on
the design, because production machines have a network.
2026-08-26 23:58:19 +02:00