A session for the mesh itself, addressed as the mesh, differing from a
node's in exactly three things: the context it starts in, its engram,
and its licence binding. Not a new kind of agent — the same mechanism
pointed at a different root. Two implementations of one mechanism drift,
and the vocabulary collision 0001 exists to undo began exactly that way.
It runs on the control-plane node, and the reasoning is easy to get
backwards: not "the important agent on the important machine", but that
this node is already the one place excepted from "compromise of a node
is compromise of that node". Placed anywhere else it would create a
second such place.
It is an addition to per-node messaging and never a replacement. 0001
holds that losing the control plane costs change, not operation — and a
mesh whose only conversational surface lived there would lose the
ability to ask anything while every machine kept running perfectly.
Writing it up exposed that the node session's setup was never designed
at all. 0004 gives behaviour and stops: nothing said how a session
starts, where its context lives, or how a broker message becomes a
prompt. That gap was invisible until something had to be built *like* a
node session. 15-the-agent-session.md covers both as one mechanism.
It also makes "a consumer that is not a machine" undeferrable. The
control-plane node now hosts two sessions that must hold different
licences, and a per-machine binding cannot express that at all. Noted in
14-model-access.md against the gap it was already recorded as.
Also completes the to-be index, which stopped at 10 and omitted four
documents. Pre-existing broken ADR references in the older rows are left
alone rather than guessed at.
Retired in favour of the lab rather than repaired — that answers the
first open question. The second finding is the one that generalises:
"nothing runs it, and nothing reports that nothing runs it" is not a
fact about that harness, it is a fact about any suite too expensive to
run on every push. The replacement inherited the fault it was replacing.
Records the three rules that now hold, and what the fix taught twice:
the remedy rebuilt the symptom inside itself, and the code that counts
results passed every test while reading nothing.
A service is reached at <service>.<node>.internal, so what resolves is anything
under a node's name. The mesh writes the data and runs no daemon; two roles,
two claims, because systemd-resolved cannot serve a wildcard at all.
Both prohibitions were found by a machine rather than by reasoning: an address
systemd already held, and reading resolv.conf for upstreams that now point at
itself.
The field was called needs, beside secrets, and both were name-to-path holding
something secret. What separates them is whose, not how secret — so that is
what the name says now.
The connectivity design still said a hub cannot be filtered — a gap recorded in
the morning and closed in the afternoon, left standing as though it were
current. Worse than a stale date: it would send somebody away from something
that works.
`restart-on` was described nowhere, including the part added today that lets a
service reflect a file another module put on the machine. A rule the host
enforces and no document mentions is a rule nobody can rely on.
And nine of fifteen design documents claimed an `updated:` older than their last
change, some by a week. That field is what cross-cutting views are generated
from, so it is not decoration.
A resolver takes over /etc/resolv.conf, which is a singular resource — ADR 0009
lists it in the table beside the seat and pid 1. So choosing between resolved,
dnsmasq and unbound is assigning a module, per machine, and the mesh refuses
two rather than letting them fight over the file.
Recorded because it was treated as an open question two days after being
decided, which is the argument for that table being a table.
Found by a container failing to resolve a name every machine could: a container
gets its own hosts file holding only its own hostname, and on the machine it
always worked, which is what made it easy to miss.
Declared containers are given the names. A container somebody starts by hand is
not the mesh's to configure — which is a second, different reason to want a
resolver, recorded beside the first rather than folded into it.
Asked whether a machine that drops off needs re-adopting: it does not, nothing
expires, and the only thing that forces re-enrolment is losing its own key.
The gap was the twenty or thirty seconds after a resume in which a node
believes it is in a mesh it has left — recovering on its own, which made it a
quality gap rather than a fault, and still a machine waiting to be told
something it already knew.
It meant failed-or-refused, so the question this record says must not be lost
was answerable only for the machines that broke. Out of date, never told, and
not worked out are kept apart: the remedy is the same push and they read
differently to whoever is looking.
Their subject matter has been built and proven for days and their frontmatter
still said code: [] — which is what the cross-cutting view is generated from,
so it was claiming nothing existed for the substrate, the node lifecycle and
delivery.
One reading answered three ways, holding nothing and touching no context's
store — which is the constraint the whole document is about, and the thing the
board being replaced gets wrong.
Rotation and the provisioner contract; model access as a provision answered by
a record, with ADR 0024's other two gaps left as gaps; exposure, which closes
the open question about revoking a route; and the delivery loop, which closes
the gap ADR 0010 left when it replaced a pipeline with a comparison.
The broker's fingerprint travels with its credential, and the machine's
filesystem does not travel at all — it runs in a container, which is the
arrangement working rather than a limitation to route around.
Found by the firewall: every packet filtered as declared, and the machine
reported as not doing what it was told, because the unit that loaded the rules
had finished. Stated as a gap rather than worked around silently.
Otherwise it succeeds into a state its verify rejects, and the host's report is
accurate and names nothing. Recorded where the vocabulary is described, because
it is a rule about writing an action rather than about one action.
A hub needs its overlay port open and a node that is not a hub does not, and
they are the same module — so listens, a static manifest field, cannot express
it while the overlay module's resources are computed per node. Written down
rather than left as an oversight for whoever first puts a firewall on a hub.
Issue 014: the node's serving key was stored in the host's own encoding, so
every check that reads the file passed and no server could start. Same shape as
013 — two halves of one mechanism designed separately, each correct about its
own half. Where a file exists so a third party can read it, the format is the
interface.
Issue 003 is answered in both halves: manifests are parsed strictly, and a
module says what it listens on and from where rather than carrying a key
nothing reads. The design records what was built and how each part is checked.
Issue 013 is new, found by reading while writing the first module that has
both a computed file and a service that needs it. The file arrived second.
It failed, then the next reconcile fixed it, which is why nothing caught it.
Three additions, all written by trying to write a real database module
and finding out what could not be said.
A module may mirror an image it did not write. Naming an upstream
reference directly needs every machine to reach a public registry and
pins to a tag somebody else can move.
A module may need a secret of its own — a superuser password is not FOR
anybody, so the mechanism that hands credentials to consumers cannot
express it. Per node, so three machines have three passwords.
And the provisioner watches, which is what lets it be a module rather
than a binary somebody places. It polls rather than watching the
filesystem, because the host writes atomically and a watch on a replaced
path silently stops working.
One assignment now gets a working database provider: two directories, two
pinned containers, a sealed password and the grants manifest.
A page nobody had thought to ask for turns out to be the one a person
opens first: what is not doing what it was told. Recorded with the order
that matters — broken, then quiet, then out of date — because a page
leading with the last would bury the first.
And refused stays distinct from failed all the way to the page. They are
fixed in different places, so one word for both sends half its readers to
the wrong one.
Two additions to the module-repository design, both from building it.
A build machine has its own credential and it is not a node's: read the
build queue, write the mesh exchange, nothing else. A node's queue
carries that node's declarations.
The answer goes through the exchange and never the default one, because
permission there is per exchange rather than per queue — anything allowed
to use it can publish into any node's queue. The price is that every
asker sees every result and filters by correlation, which is cheap
against a builder never needing that permission.
And every result is kept, failures included, because one that leaves no
trace is indistinguishable from a build nobody asked for. That is what a
builds view reads; the board page is corrected to say so.
A build is work, not state, and that is why it does not travel as a
declaration: as one it would either rebuild on every reconcile or carry
"and I already did this", which is state about an event rather than about
a machine. So it has its own queue and the answer comes back correlated.
Three properties recorded because they are decisions: acknowledge only
once the answer is away, one build at a time, and a failure is a result
rather than silence.
And the board page is corrected. A build result today is answered to
whoever asked and kept nowhere, so a builds view has nothing to read. A
record of past builds is the missing piece, not the builder.
Designed with no reference to what came before, which was asked for. The
system this replaces has features — several deployable units inside one
module — and they are deliberately absent.
That closes something ADR 0001 has been carrying as an open prerequisite.
It lists "named features with per-node opt-in" as required, or "every
independently deployable unit becomes a module again and the count
returns". The premise was right and the remedy already exists in another
form: several modules, assignment per node, and a module with
requirements and no files of its own. `networking` is exactly that. The
count does not return because what made it return — a module is
expensive, so put several things in one — is gone. A module here is a
manifest and usually nothing else.
The manifest in a repository names artifacts; the manifest the mesh holds
names digests. Two documents, because a digest is not knowable until
something is built and a repository carrying one is wrong the moment
anybody edits anything.
The builder runs on a node. Building needs a container runtime and a
working tree, and what the control plane may send a machine is bounded by
the declaration language. A control plane holding a container socket
would be the one component that can do anything anywhere.
And the host's vocabulary grew from six shapes to eight — user and
archive — with the reasoning for each and for the refusals that came with
them. The count is asserted by a test precisely because every addition
widens what a compromised control plane can express.
Read from the board that exists. Eight sections; four are about work and
workers and are held back with that domain. The other four are the mesh
itself, and everything behind the main one already exists here — it is a
reader, not a second source of truth.
The constraint is the point of writing this down now. The existing board
is one service that reads every context's database, because that is the
shortest path to a page showing all of them at once. That is ADR 0008
violated by the one component with a reason to violate it, and the cost
is the same one the shared library has: a boundary nothing may cross is a
boundary that can move, and one thing crossing it is enough to freeze it.
So a board reads through interfaces and stores nothing. If a question is
slow, the answer belongs in the context that owns it, where everything
else asking gets it too.
First end-to-end raise. A machine with a container runtime applied the
bundle its host carries and ended with a store, databases, schemas, a
broker holding a certificate it generated itself, and the control plane
serving. Then it took a token, checked the broker against the pinned
fingerprint, generated three keypairs and enrolled — the first node being
a node whose mesh is not up yet, observed rather than argued.
And a credential crossed. Declared the provider of a database for a
second node and pushed to over the broker, the machine ended with the
password in one file at mode 0600, and that password appears nowhere in
the declaration that crossed the broker, nowhere in the control plane's
database, and nowhere in what the node reported back. That is the whole
secrets argument, measured.
One fault, in the joining: the token did not say what the mesh calls the
machine, so enrolment needed a flag its own help said it did not, and
failed at the broker with an empty username. It is the fifth thing a
token carries now — the node cannot work its own name out, because the
broker account it authenticates as is named after it and exists before
the mesh has told it anything.
0009 already said a consumer supplies a target and receives a name. What
it did not say is that those are two separate mechanisms.
Contribution — publish me at this name, on this port — now exists.
Binding — and hand me back a credential — does not, and is the larger
half: a secret has to exist, be stored, reach one node and not the
others, and rotate with every holder informed. That is the invariant set
found violated three ways at once, so it is not something to add in
passing.
The absence had a measured cost. Exactly two modules opened a direct
connection to the control plane's database, and they are the reason every
node permanently holds a credential to it. Both were doing by hand what
this edge is for. Neither needed a new kind of thing.
Two records, from building it.
0009 has a section titled "there are no domain modules", and `networking`
now exists. It is not a contradiction and it reads as one, so the
difference is written down: what was refused contains WireGuard and a
proxy and is assigned where half of it is unwanted. What exists contains
nothing — requirements and a name — so there is no half. Every artifact
it leads to is still an ordinary module assigned on its own terms.
With the cost stated, because it is real: adding a second implementation
turns a settled question into an open one for everyone using the bundle,
not only for whoever wanted the alternative. That is the refusing rule
applied consistently, and the alternative is a default, which is the
flavor field returning under a better name.
08-connectivity gains why the network stopped being code beside the
module system: a machine was on the private network because it had an
address, and there was no way to keep one off. A manifest can now say its
resources are computed, which is what a peer list needs.
And three modules rather than one, because WireGuard is one VPN of
several. Naming a module after the job and putting one implementation
inside it is flavor wearing a generic name — the second VPN has nowhere
to go.
Three decisions, all Jochen's, and the first is the one that unlocked it.
Exclusivity is not a property of a module. It is a property of a singular
resource the module takes over. Two shells compete for nothing and any number
may be installed; two display servers both want the seat. So a module declares
what it CLAIMS, and two modules claiming the same thing cannot both be assigned
within that claim's scope.
Not "xorg conflicts with wayland". Pairwise exclusion has a property that only
shows up later: adding a third display server means editing xorg and wayland to
know about it. Every new module requires changing modules nobody who wrote it
owns, and the edits grow as the square of the count. With a claim the third one
says what it claims and nothing else changes anywhere.
Claims have a scope -- node, site, mesh -- which is not new. The mesh already
enforces exactly one hub with a unique index. Scope is that idea said once
rather than hard-coded per case.
And some conflicts need no claim at all: two modules declaring the same file or
binding the same port are visible from what they declare. A claim is only
written for the abstract ones.
A requirement with several answers is refused, never guessed. One candidate is
assigned silently because there was no choice to make; none is refused naming
what is missing; several is refused naming them. That is what makes a solver
unnecessary -- counting candidates has no surprising behaviour, and a solver
can be added later without changing a single manifest.
Flavor is retired. It was carrying three unrelated meanings: variants of a
thing, a subset of a module a node installs, and whatever the current system
does, which earned two knowledge-base entries about going wrong. A word with
three meanings cannot be reasoned about. What it reached for is two ordinary
things -- different modules providing the same thing, and one module with a
setting.
All on the first three machines to actually run it, and all invisible from the
mesh's own state: the graph was right, the files were right, the services were
up, every node reported success, and the network did not work.
A running interface does not re-read its configuration, so a node joining left
every existing node carrying a network that no longer existed. A hub sharing a
site with a spoke was emitted twice, which WireGuard refuses. Two nodes at one
site that neither can be dialled were peered directly, so nobody opened the
path and the more specific route blackholed -- this document's own warning
arriving in its implementation. And Docker sets the FORWARD policy to DROP, so
a hub with forwarding enabled still carried nothing between its spokes.
The last one is the sharpest: the substrate at tier 1 silently breaks the
network at tier 2, and nothing in either tier's state says so.
None of these is reachable by reasoning, and each was found within minutes of a
real machine trying it. That is the argument for the lab in one line.
I have been treating "what a node presents to prove it is that node" as an
undecided design question for weeks, and blocking on it. It was decided.
08-connectivity says of the overlay keys: each node generates its own keypair,
the private key never leaves the machine, the public key is published to the
mesh -- and says explicitly that this IS ADR 0004's "a node holds its own
identity", applied. Nobody had applied it to the thing 0004 is actually about.
What caused it was a word. The lifecycle said a joining node receives its own
durable identity, which reads as the mesh issuing something, and then the
question is what. The mesh issues nothing. A node arrives holding its identity;
what it receives is being known. That line now says what happens: it presents
the one-time secret and its own public key, which the mesh records.
The rule above it then holds literally rather than aspirationally. The mesh
stores a public key, so a copy of the mesh's database grants nothing, and
compromise of a node really is compromise of only that node.
Also recorded, since it was asked directly: same principle as SSH, own key, not
the machine's SSH host key. Host keys are regenerated by reinstalls and image
clones, which would silently un-enrol a node; their lifecycle belongs to sshd
rather than the mesh; and a partial host has no SSH daemon at all, so an
identity scheme resting on one excludes a supported kind of node.
The good half of that idea is kept: the mesh knows every node, so it can
distribute host keys the way it distributes authorised keys, and node-to-node
SSH stops depending on trust-on-first-use.
Four things settled by talking them through, all of which had been true in
somebody's head and written nowhere.
It is not a mesh in the peer-to-peer sense and will not become one. 0001 now
says what it is instead: machines linked by a private network, one node holding
knowledge of all of them, modules as the way anything is built and delivered,
and agents hired onto nodes to do the work. The word describes what machines
can reach, not how they are governed. "Master" overstates it the other way --
nothing needs that node to keep running, only to change.
0006 gains the option that would make it a real mesh, recorded as considered
rather than rejected by silence: every node holding the whole inventory, a
replication process, an elected master with promotion on failure. What settles
it is not the complexity but that it still would not deliver the name, because
application databases are not replicated -- so a genuine peer-to-peer mesh
means becoming a replicated database system for every consumer's data too. That
is a larger product than the thing it would support.
Also in 0006: three central roles, not one. Losing the control plane costs
change, losing the broker costs being told anything, and losing the hub costs
nodes in different places reaching each other at all -- which is operation, not
administration. Whether they are one node is not decided.
And SSH access is identity's. It appeared three times as something that uses
the overlay and never as something the mesh provides, which reads as settled
when nothing decided it. Nobody else could: the mesh is the only thing that
knows which humans and agents exist and which nodes they may reach. Node to
node SSH stays out -- the host has no inbound control surface by decision, and
nodes reaching each other that way is a second control path through the back
door.
0007 gains the requirement underneath all of it. Reachability was recorded as a
fact to track and never as a thing some node must have. The broker's node and
the hub must be dialable by every node at a stable address, or nothing can join
and a disconnected node cannot return. A mesh entirely behind NAT cannot be
raised. That is a precondition and it belongs with the others.
The link staying on the underlay is also argued now rather than asserted. At
join time it is forced; afterwards it is a choice, and the reason is that a
repair channel carried over the thing being repaired is not one. Moving it onto
the overlay, with fallback, is recorded as open with what it would have to get
right -- a WireGuard interface has no link state to test, and a silent fallback
is this repository's recurring fault in a new place.
0010 says in one line what was the intention throughout: the module system is
the CI/CD. Not a pipeline beside the mesh. Build, test, publish and deploy are
one reconciliation seen at four points, which is why a thing that cannot be a
module cannot be delivered.
Back to 23 records. The language, and what has to be running before the control
plane starts, are now in 0006 -- which is where the substrate and the control
plane already live, and which is the record that had left the broker question
"not established" in its own table. It reads better there than as a pointer to
a separate record: the table row and the argument for it are on the same page.
The store mechanics went into 0008. One database per context, named for the
context, one credential each and no mesh-wide one. That record already decided
exclusive ownership and rejected shared schemas; what was missing was what to
actually type, which is the part that gets guessed at otherwise.
Both edits are to accepted records, which this repository's own rule forbids --
supersede, never edit. Recorded here so it is visible rather than silent. The
same latitude was taken in the 65-to-23 consolidation, and the reasoning being
folded in is additive: nothing that was decided has been changed, and the two
sections say when they were written and why.
mesh-control is built as far as it can honestly go: one context of seven,
inventory, with its schema and the command that applies it. The repos map and
the control plane design say so, and point at ADR 0024 for what it took.
Separately, and more importantly: this repository described the enrolment token
as carrying three things when ADR 0004 says four. The missing one is the
control plane's signing identity -- the reason a node does not have to trust
the broker it dials.
Without it the control plane's authority is transitive through the broker, and
0004 spells out what that costs: a compromised broker could forge declarations,
and since the host applies whatever the link delivers, that is the whole
machine. The record has the argument in full; the design doc had dropped the
conclusion.
Found by reading the two together while deciding what the control plane must
store, which is roughly the only way it would have been found -- both documents
are internally consistent and only disagree with each other.
Two things found by trying to build tier 2.
The substrate design asked whether the message broker has to be running before
the control plane, and framed it as depending on whether the control plane's
own parts talk to each other over it. They do not -- it is one process -- so
under that framing the broker stays out of the bundle.
The framing cannot answer the question. What decides it is how the control
plane reaches a node, and the answer was already decided: only ever over the
link, and the link is the broker. So provisioning the broker would require the
broker. The first node does not escape this by being local, because it enrols
the ordinary way, by dialling the broker at the address in its token -- which
was deliberate, and worth keeping.
The bundle is two images now. The record says what that costs, including a
certificate the broker needs at a moment when there is no mesh to issue one.
The language had never been decided for tier 2. Go, for the same reason the
host is: the bundle pins this image by digest and runs it where nothing can
check it, so the image should hold the program and nothing else.
Also corrects something already built: the bootstrap created one database and
called it 'mesh'. ADR 0008 grants a context only what it exclusively owns and
ADR 0006 says the mesh database names a thing that will not exist. One database
per context, so one today, called inventory.
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.
Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.
Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).
Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.
Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.
The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18,
19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only
the archaeology of what used to be there.
Renumbered contiguously. Renames run in ascending order, so every target number
is already free and no two files ever collide.
The reference rewrite is one simultaneous pass rather than a sequence of
replacements. Numbers moved into slots other numbers were vacating -- the node
host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time
would have cascaded and silently pointed things at the wrong record.
Seven plain-text references survived the merges as prose rather than links,
naming records that no longer existed: the enrolment token, the link boundary,
what a declaration is, reachability, the repository structure. Each mapped to
the consolidated record that now holds it.
Verified rather than assumed: every [ADR NNNN](path) link now has matching text
and target, checked across the whole repository, and the checker passes.
Frontmatter `consolidates:` lists dropped -- they named records that are gone,
and each consolidated record already says in prose what it absorbed.
Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.
the node host 8 -> 1 applies not decides, depends on nothing,
per operating system, root service, the
launcher, episodic, what a declaration is,
actions from the bundle only
a node and how it joins 4 -> 1 what a node is, joining, the link as
security boundary, the enrolment token
modules and the graph 7 -> 1 everything is a module, no domain modules,
three edges, provisioning, the core library
substrate and control 6 -> 1 the test, seven contexts, one control plane,
plane the authority is not a database, the named
products, the pinned bundle
connectivity 3 -> 1 a route is a grant, reachability declared,
filter rules
delivery 5 -> 1 reconciliation not a pipeline, artifacts,
the three silos, a failed step, the verdict
the lab 5 -> 1 (earlier)
how this repository 10 -> 1 (earlier)
works
Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.
The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.
The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
Jochen: a normal application has 3-5 ADRs, maybe 10 for a large one, and we are
at 65. Fair, and the cause is mine -- I recorded every FINDING as a decision
rather than every fork in the road.
Two merges, both cases where one decision had been split across many records
because it was taken over several days rather than at once.
0019 absorbs ten records about how this repository works: what it is and that
it is public, the folder flow, the two design layers, the issue front door,
status in frontmatter, playbooks, the naming rule, the product name. Those were
never ten decisions -- they were one, seen from ten angles as the repository
took shape.
0016 absorbs the five about the lab: a node is a virtual machine, a router is
scenery, a scenario declares the underlay, a scenario is a closed address
space, and the two scenario classes. Same pattern -- one design, split by the
order it was worked out in.
The consolidated 0019 also raises the bar for what earns a record, since that
is what produced 65: a record is warranted when there is a genuine fork -- a
direction reversed, an alternative that will be proposed again, something
contested. A finding is not a decision, and a bug is certainly not. Everything
else belongs in the design document where the reasoning is actually read.
The checker earned its place here. Deleting nine records left 13 dangling links
across the repository and it named every one, including in AGENTS.md. Nothing
was found by reading.
Remaining clusters worth the same treatment: the host (8 records), delivery
(5), modules (6), connectivity (4), substrate and control plane (4). That would
be 52 down to roughly 30.
Jochen: a jungle of specs that slightly contradict or patch each other, and
what matters is a working state rather than history. Both are fair and both are
mine.
Measured rather than assumed. 05-the-node-host and 09-the-node-lifecycle both
covered enrolment, the install commands, the unit file, the launcher and
reconcile -- I wrote 09 without taking anything out of 05, so the same things
were said twice and could drift apart.
Split by what each document IS. 05 is the component: what the host is, its
parts, the declaration vocabulary, the build order, how it is verified. 09 is
what happens to it: install, enrol, run, upgrade, retire. The whole "The
process" section left 05, and the unit file moved to 09 where installing is
described. 05 goes from 338 lines to 245 and now points at 09 rather than
restating it.
09 also carried a 105-line "Resolved" section -- six mechanisms framed as
"these were open and here is the answer". The content is needed; the framing is
history, and history is what makes a document read as a changelog rather than a
description. Renamed to what it actually is and the was-open phrasing removed.
Also added 10-delivery.md, which did not exist: four accepted decisions --
0054, 0063, 0064, 0065 -- had no design document at all, which is the specific
reason the delivery picture felt scattered. It is now one document covering
modules, the three edges, the core library, and how a change becomes a running
thing, with a table of what each property is designed against and what must
exist before it can be built.
0060 named the gap and did not close it: everywhere else an init runs the
launcher at boot, and Android grants neither an init to register with nor
anything worth supervising, because a supervisor would be killed alongside what
it supervises.
Closed by narrowing what is required rather than building something. A host is
resident or episodic, and both are hosts. Being killed by the platform is
disconnection, which 0036 already made ordinary -- and every mechanism an
episodic host needs already exists because it was built for laptops that close.
A partial host can join a mesh and cannot be the first node, since every
bootstrap step is a shape it refuses. Its bundle says so.
Two consequences that are easy to miss: last-heard-from means much less on an
episodic host, so a healthy phone reads as a dead server unless the reader
knows which kind it is; and a declaration may take a long time to land, which
makes 0058's outstanding-versus-failed distinction load-bearing.
Still open, and in that order: what an Android node is FOR, and only then how
it is started.
A sweep for claims overtaken by the last few days. Annotated rather than
rewritten, following the pattern already in 0049 -- what changed and why is the
useful part, and an accepted record should not quietly become something else.
0057's init section was wrong on all three of its claims. It said the host
needs FOUR things from an init; 0061 reduced that to one. It said every machine
the mesh targets already has systemd; Alpine does not, and it is the intended
first node. It said there is no second init to abstract over; there is now, and
the answer is still not an abstraction -- it is a four-line file per system.
What survives is the part that was always right: an init is not a dependency in
0041's sense, because it is not installed, it is what the machine already is.
0048 named Docker as the container runtime. It is now docker or podman,
detected rather than chosen -- because adoption keeps what a machine already
has, so naming one contradicted a rule already decided. That row is the only
one of the five that names two, and the record now says why.
0060 claimed the bundle is portable across operating systems. Its mechanism is;
its contents are not -- package names, unit names, service names all differ, so
an Arch host embeds an Arch bundle. That was my error, and it is the exact
confusion behind the question that found it.
The design layer had the same drift: 07 and 09 said "Docker" where they meant a
container runtime, 09 said systemd restarts the host after an upgrade when the
launcher does, and both install snippets assumed Arch. They now show Alpine and
Arch side by side, which makes the point better than prose did -- step 1
differs per system, step 2 never does.
Checked and NOT changed: 0047's "the vocabulary grows by one shape" is a claim
about the rate, not the count, and is still true. 0037 lists docker among tools
the host manages, which it does. 0041 says nothing about either.
0060 said the container runtime was a separate decision. It is now made, and
the reasoning is worth keeping because it is the opposite answer to the same
question one paragraph earlier.
Abstracting service managers is lossy -- systemd and OpenRC are different
models and LoadState has no equivalent. Container runtimes converged on one CLI
deliberately, so almost nothing is lost: checked against podman 6.1.0, run,
rm -f and docker's own template syntax for state and labels all work unchanged.
Only the probe differs. So: a two-entry lookup, not an interface.
The difference that is NOT in the CLI is the one that would have shipped
silently. Podman accepts --restart unless-stopped, records it, and has no
daemon to act on it -- containers do not return after a reboot unless
podman-restart.service is enabled, which by default it is not. Every command
reports success and the effect does not happen.
That belongs in the declaration rather than the host: a node using podman is
told to enable the unit. Which is what made the service shape's missing 'boot'
field visible, and it is now built.
0060 -- the host is built per operating system. systemd and pacman are the Arch
host's implementation, not abstractions the mesh has to grow. They are not
independent choices: a machine has pacman because it is Arch, and the package
manager, service manager and packaging format arrive together as one decision
somebody made at install time.
Rejected abstracting them, and the reason is correctness rather than effort.
The service applier reads LoadState to tell "not installed" apart from
"stopped", which is what stops it reporting absence as success. An interface
spanning systemd and OpenRC degrades to what both express, and the lowest
common denominator is exactly where that fault lives.
Almost all of it is shared -- the vocabulary, store, apply loop, read-back
discipline, refusal model, bundle and link are portable. Two appliers differ.
And delivery was already per-OS, since a .pkg.tar.zst is an Arch artifact, so
this is the seam that already existed.
Android is the interesting case rather than Debian: no service manager, no
package installation, usually no root. Such a host implements file, directory
and action and refuses the rest -- the same refusal a host already gives an
unknown type, with a different reason. Those three are the portable floor.
The container runtime is deliberately left open: it is not an OS split, since
Arch runs docker or podman.
0061 -- the init is asked for start-at-boot and restart-on-exit, and nothing
else. Both are expressible in OpenRC, runit, s6 and an Android init.rc.
Counting failed starts and rolling back moves into a launcher, because that is
the one piece which must work when the host does not, and a script with a
counter can be tested where OnFailure= can only be hoped for. Supersedes 0059,
keeping its reasoning in full.
The checker found all six places citing 0059 and refused the commit until they
named the replacement.
Left implicit by the previous commit, which said the owning context writes
without saying what does the consuming.
The control plane is the consumer, and there is one of it. Seven contexts but
one deployable, so it is one process dispatching internally rather than seven
consumers racing -- which matters because the as-is records two consumers
accidentally sharing a queue and silently splitting the traffic, each getting
half of what it expected. With one consumer that cannot arise.
The broker is also the buffer while the control plane is down: nodes keep
publishing, messages queue, the control plane drains them on return. That is
what makes a single control plane tolerable -- an outage delays the mesh's
knowledge rather than losing it.
One consequence named because it will otherwise be discovered: an unbounded
queue grows until the broker's disk is full, and the broker is the component
every node depends on. The bound is per queue and undecided -- dropping the
oldest health report is obviously right, dropping the oldest declaration
acknowledgement is not.
Not approved as drafted -- four things came out of checking them against each
other, and one was a bug that would have broken every upgrade.
The bug: 0059 specified Restart=on-failure while 0057 has the host restart onto
a new binary by exiting CLEANLY. on-failure does not restart a process that
exited zero, so every upgraded node would have been left stopped, having
successfully upgraded. Found by reading the two records against each other
rather than by either alone. Now Restart=always in all three places that
mention it.
The host cannot run in a container, and the reason is decisive rather than
stylistic: step 0 of the substrate bootstrap installs the container runtime, so
a host inside a container would need the thing it exists to install. It would
also break 0041 -- copy it onto a machine and run it stops being true when the
machine must already have a runtime. Everything above tier 0 is a container;
the host is not. That split is the tier boundary, not an inconsistency.
systemd is named rather than abstracted. An init is not a dependency in 0041's
sense: 0041 is about what must be installed before the host works, and an init
is not installed, it is what the machine already is. The unit file is the only
systemd-specific artefact and it belongs to the package, so a machine with a
different supervisor ships a different package.
The mesh is a watchdog, and my first draft was half an answer. Recovery must be
local -- nothing dials a node, and a host that cannot start cannot report. But
detection is the mesh's, and a local supervisor structurally cannot do it: it
sees one process failing and cannot tell a broken machine from a broken
release. Only something watching every node can, and that distinction decides
whether the response is "fix this machine" or "stop shipping this version". So
a host rollout is staged -- a few nodes, wait for heartbeats, continue or stop
on silence. Local rollback still needed, because the canary nodes break and
because a node offline during the rollout gets the declaration later with no
batch around it.
The first declaration is the overlay and nothing else. Forced, because a node's
address and peers are assigned rather than chosen. But also the way back in: a
node reachable over the overlay can be fixed by hand if a later declaration
breaks it, and a large first declaration risks a node that is broken and
unreachable at once.
Also stated plainly, because it reads as a contradiction: nodes reach each
other over the overlay and every node consumes from the broker; what 0039
forbids is an inbound CONTROL surface, not reachability.
And in 06: no node holds a credential to any control-plane store, for reads or
writes. Four ADRs already say this separately and none of them said it in one
place. Nodes state over the broker; the owning context writes. With a note that
most high-frequency writes are observability's, not the registry's -- routing
logs into the registry would be the shared-schema mistake arriving through a
door marked performance.
Two corrections and one new decision, all from Jochen catching things.
Pushed, not polled. I described updates as landing "on the next reconcile",
which reads as polling and is not the design. A declaration arrives as a
message on a link that is already open; the host applies it then. Polling over
an existing connection would be slower to land AND constant traffic to learn
nothing.
The timer is for drift and nothing else, and it cannot be replaced by an event
for a definitional reason: drift is change the mesh did not make -- somebody
edited a managed file, a distribution upgrade replaced a config -- so nothing
will ever publish a message about it. Only looking finds it.
Separated the heartbeat from the reconcile timer, which I had been conflating.
They point in opposite directions and answer different questions: the timer
looks at the machine and asks whether it still matches; the heartbeat reports
upward and is what makes silence mean something. A node with nothing to do
sends nothing, and without a heartbeat that is indistinguishable from a node
that stopped.
0059 -- a host that cannot start is rolled back by the service manager. I had
left this open on the grounds that recovery meant the host judging its own
health. That objection does not survive being asked properly: a keepalive is
something else judging the host. The watchdog must be local, because nothing
dials a node and a host that cannot start cannot report -- so it is the service
manager, which is already there.
The failure it prevents is sharper than "the node is down": a host that will
not start looks exactly like a machine somebody switched off, which is the one
condition this design has deliberately decided not to alarm on. So a bad
release reaches every node, each goes quiet, and the mesh reports a fleet of
sleeping laptops.
Confirmed means started and completed one reconcile -- deliberately not "the
link is up", or a laptop on a train would roll itself back. The rollback is a
script shipped by the package, not a host subcommand, because a binary that
will not start cannot be its own recovery. It rolls back once: a second failure
means the machine is the problem, not the binary.
Also refined the records checker, which produced a false positive: a proposed
record may extend another proposed one, because decisions are drafted in chains
and the alternative is marking things accepted to satisfy a check. An accepted
document resting on a proposed record still fails, and that was verified.
0057, 0058 and 0059 are all proposed.
The upgrade question turned out to be a delivery question, so 0058 answers
both.
Today's third silo runs once per node and sends each one a command to install
and start. That is where the as-is records a package install that 404ed from
every mirror while the job went green, an image pull failure that did not fail
the deploy, and a verify stage that was built and never scheduled because it
was missing from a list.
The shape underneath all of those is that the thing reporting success was not
the thing doing the work. Meanwhile ADR 0037 has given every node a component
that applies state, reads back and reports -- so two mechanisms now change a
node and only one checks its work.
0058: a pipeline ends when the declaration is updated. Deploy stops sending
commands to nodes and becomes one write. The host applies it on its next
reconcile, and the host cannot report success it did not verify. The verify
stage disappears as a stage, which is the point -- verification stops being a
step that can be left off a list.
A pipeline result now means "the declaration is updated, and here is which
nodes have applied it". It does not wait for every node, because a node may be
legitimately switched off for a week. Outstanding is reported separately from
failed, since conflating them is how the old system produced a stall with no
error anywhere.
The host is delivered by exactly this path and needs no new resource type: a
`file` writes the package manager's config pointing at the mesh's repository, a
`package` names the version. Added a step I had missed -- before exiting for a
restart, the host runs the new binary once. A package can install something
that does not execute here, and that turns "the node never came back" into "the
apply failed and said why".
Six open items resolved: re-enrolment is decided when the token is issued and
revokes the previous identity; the mesh keeps a recovery copy of what each node
reports it owns, which un-strands the orphans; last-contact is reported with no
threshold, because a laptop off for three weeks is doing nothing wrong;
adoption always completes but a failed line makes a node ineligible for
assignment; a briefing is a structured document whose outcome is computed from
its lines; and the token is printed once and carried by hand, which is the
property that makes it worth anything.
Still open and named: automatic rollback of a host version that will not start.
0057 and 0058 are both proposed.
The host was described as a component and never as something that runs for
years on a machine somebody else also uses. 09 covers every state a machine can
be in and every transition between them.
Four states: unmanaged, hosted, enrolled, disconnected. Only the last two are
nodes, and they are the same node in two situations. `hosted` -- the host
installed but never told which mesh it belongs to -- had no name before and is
where a machine sits between the two adoption commands.
Things that were unclear and now are not:
The first node walks the same path in an unusual order: reconcile from the
bundle, the control plane it just raised issues a token, enrol against it. Its
specialness lasts two commands. A side effect worth having -- enrolment is
exercised on node one, rather than being written and first used on node two.
Enrolment reports profile and inventory BEFORE the control plane decides
anything. The profile is the input to that decision, not a diagnostic; the
control plane cannot decide what a machine should run without knowing what it
can run.
Rebooting mid-apply is safe by construction. The store records each resource
after it worked, so a host that dies half way through comes back and applies
the rest. The rule that stops the host lying about what it did also makes it
crash-safe.
Retiring splits in two. Graceful is a final empty declaration. A node that is
gone will reconcile its last declaration forever -- the honest consequence of
making disconnection ordinary. The answer is not to make the host expire but
that the node holds nothing that outlives revocation: every grant is a per-node
credential revoked at the provider. A lost node keeps running and stops being
able to reach anything. Said plainly rather than implying the mesh can switch a
machine off, which it cannot and should not.
Losing the store is quiet and permanent, so it gets its own section. The host
re-enrols and re-applies fine; what does not come back is removal, because
resources it no longer has a record of become unowned and sit there
indefinitely.
Also corrects 0057, which said the mesh must not upgrade the host at all. That
conflated two acts. Replacing the binary is safe -- Unix keeps the running
inode. Stopping the unit is not. So the host may apply a package naming itself,
and restarts by finishing its apply and exiting cleanly, letting the supervisor
start it on the new binary. It never asks the service manager to restart it.
That makes a fleet-wide host upgrade an ordinary declaration, which the first
draft gave up on.
0057 remains proposed.
The design described what the host does and never what it is at runtime. The
words daemon, long-running, interval, poll and heartbeat appeared nowhere in it
or in the relevant decisions. What exists is a command that runs and exits;
what the design needs is a process holding a link. Nobody had written down that
those differ, so several questions had no answer.
0057 settles them. It runs on every node -- the host is what makes a machine
managed, so a machine without one is not a node. Root, because no useful subset
of the job is unprivileged. A systemd unit, because something must survive a
reboot to hold the link.
It never manages its own unit. The temptation is obvious and it ends with a
host stopping itself half way through an apply, leaving a machine with nothing
running to fix it. The installation owns the host; the host owns everything
else.
Installed as a package, with a tarball as the floor. The package carries the
unit file, the state directory and an upgrade path, which a bare binary does
not. But the mesh's package repository is hosted on the mesh, so any route that
needs the mesh to install the thing that joins the mesh is a circle -- the
tarball is the path that must never acquire a dependency.
Reconciles on start, on a declaration, on a timer and on reconnect. The timer
is the one easy to leave out, and without it `owned` reports what the host
applied rather than what is there -- ADR 0035 violated by omission.
The records checker caught this commit on its first attempt: 05 listed 0057 in
its frontmatter while 0057 is still proposed, and a to-be document may not rest
on an unaccepted record. The section now says so in the body instead.