Commit Graph
262 Commits
Author SHA1 Message Date
jschoubben 5ad3641cbf Record the board, which is the last designed document with no code
One reading answered three ways, holding nothing and touching no context's
store — which is the constraint the whole document is about, and the thing the
board being replaced gets wrong.
2026-08-31 04:50:45 +02:00
jschoubben 0bc4b7774f Record what four more pieces of the mesh became
Rotation and the provisioner contract; model access as a provision answered by
a record, with ADR 0024's other two gaps left as gaps; exposure, which closes
the open question about revoking a route; and the delivery loop, which closes
the gap ADR 0010 left when it replaced a pipeline with a comparison.
2026-08-31 02:56:50 +02:00
jschoubben 6b1c80b442 What a builder-as-a-module can and cannot reach
The broker's fingerprint travels with its credential, and the machine's
filesystem does not travel at all — it runs in a container, which is the
arrangement working rather than a limitation to route around.
2026-08-31 01:53:35 +02:00
jschoubben 8448219de1 Issue 018: a provider on the same machine was never announced to its consumer 2026-08-31 01:35:54 +02:00
jschoubben 00e98f1f92 The vocabulary has no word for a unit that runs and exits
Found by the firewall: every packet filtered as declared, and the machine
reported as not doing what it was told, because the unit that loaded the rules
had finished. Stated as a gap rather than worked around silently.
2026-08-31 01:20:58 +02:00
jschoubben 0ac99f0d68 An action's own idea of being finished must be its verify's
Otherwise it succeeds into a state its verify rejects, and the host's report is
accurate and names nothing. Recorded where the vocabulary is described, because
it is a rule about writing an action rather than about one action.
2026-08-31 01:03:50 +02:00
jschoubben 1ade18209d Issue 017: an action succeeded into a state its own verify rejects 2026-08-31 01:01:17 +02:00
jschoubben a9cd3de5be Issue 016: anything after the declaration in a file was ignored 2026-08-31 00:51:46 +02:00
jschoubben 6e7e77acbe Issue 015: the harness read a swallowed answer as success 2026-08-31 00:48:16 +02:00
jschoubben 890c3ee3fc State the one rule the derivation does not yet reach
A hub needs its overlay port open and a node that is not a hub does not, and
they are the same module — so listens, a static manifest field, cannot express
it while the overlay module's resources are computed per node. Written down
rather than left as an oversight for whoever first puts a firewall on a hub.
2026-08-31 00:42:39 +02:00
jschoubben a6872ac099 A key that is present and unusable, and what the certificate work became
Issue 014: the node's serving key was stored in the host's own encoding, so
every check that reads the file passed and no server could start. Same shape as
013 — two halves of one mechanism designed separately, each correct about its
own half. Where a file exists so a third party can read it, the format is the
interface.
2026-08-31 00:42:13 +02:00
jschoubben 778efaba8b The mesh runs its own registry, certifies its own names, and computes its own filtering
Issue 003 is answered in both halves: manifests are parsed strictly, and a
module says what it listens on and from where rather than carrying a key
nothing reads. The design records what was built and how each part is checked.

Issue 013 is new, found by reading while writing the first module that has
both a computed file and a service that needs it. The file arrived second.
It failed, then the next reconcile fixed it, which is why nothing caught it.
2026-08-31 00:37:34 +02:00
jschoubben ba14e2b629 Issue 012 — the first diagnosis was wrong, and that is the useful half
Two things changed at once: a fourth image in the scenario, and scenario
machines raised from 1 GiB to 2 GiB. The bootstrap then failed every
time, and the image was blamed.

Removing the image did not fix it. Removing the memory increase did —
nine assertions pass again on three images with the machines back at
1 GiB. Three machines at 2 GiB on a host doing other work contend enough
that the store container does not come up at all.

The ordinary lesson, and it still caught me: two changes together, the
failure attributed to the plausible one, and an issue written recording
the wrong cause. What found it was reverting to the exact last-known-good
state rather than reverting the suspicious change.

What remains untested is whether a fourth image alone is fine. Probably.
Nothing has measured it, and the honest state of this issue is that what
it was opened about was never demonstrated.
archive/initialization-consolidation
2026-08-30 20:40:54 +02:00
jschoubben de2b6a2825 Issue 012 — a scenario machine cannot hold four images and raise a
substrate

Adding a fourth image to the two-machine scenario makes the bootstrap
fail every time, with the store's readiness check producing no output at
all — which says the container was not running rather than that the
database was slow. Three images pass nine assertions; four never get past
the store.

More memory did not change it, so memory is not the cause; the change is
kept because the reasoning holds on its own. Disk is the most likely
explanation and nothing has measured it.

It blocks proving the mesh runs its own artifact store, since the
registry module needs a registry image to mirror. The module is written
and accepted; what is unproven is a machine assigned it serving another.
2026-08-30 20:29:55 +02:00
jschoubben 1b49e684e1 Issue 011 — an action is a gate, which the first fix got wrong
The first fix continued past every failure, and the next lab run failed
at the bootstrap: the store did not answer in three minutes and then said
"the database system is shutting down". Carrying on past the readiness
gate had started the broker and the control plane against a machine that
was not ready, and on a small machine that is how a database still
initialising has its memory taken away.

An action is the only shape whose purpose is to make something true
BEFORE the next thing needs it, which is why it is the only one with a
verify. So a failed action stops what follows and nothing else does —
which fixes both this and the hostage problem the issue was opened for.
2026-08-30 20:07:41 +02:00
jschoubben 4a601ceb02 Issue 011 — one broken module stops every module after it
Found in the lab. A machine with one impossible module applied nothing at
all on every later push, and the mesh said "failed" without saying the
rest was never attempted.

Recorded with the evidence, including that the behaviour's test cited a
record which does not decide it: ADR 0010 argues about pipelines against
reconcilers and says nothing about whether one resource failing should
stop the next being attempted.

Fixed in mesh-host: everything is attempted, every failure reported.
2026-08-30 19:26:03 +02:00
jschoubben c3ec1487c7 What a module may borrow, what it may need, and what one assignment gets
Three additions, all written by trying to write a real database module
and finding out what could not be said.

A module may mirror an image it did not write. Naming an upstream
reference directly needs every machine to reach a public registry and
pins to a tag somebody else can move.

A module may need a secret of its own — a superuser password is not FOR
anybody, so the mechanism that hands credentials to consumers cannot
express it. Per node, so three machines have three passwords.

And the provisioner watches, which is what lets it be a module rather
than a binary somebody places. It polls rather than watching the
filesystem, because the host writes atomically and a watch on a replaced
path silently stops working.

One assignment now gets a working database provider: two directories, two
pinned containers, a sealed password and the grants manifest.
2026-08-30 18:28:30 +02:00
jschoubben 34551a3b8a The three questions a board answers, in the order they are asked
A page nobody had thought to ask for turns out to be the one a person
opens first: what is not doing what it was told. Recorded with the order
that matters — broken, then quiet, then out of date — because a page
leading with the last would bury the first.

And refused stays distinct from failed all the way to the page. They are
fixed in different places, so one word for both sends half its readers to
the wrong one.
2026-08-30 18:09:17 +02:00
jschoubben 41a4405253 What a build machine may do, and what is kept
Two additions to the module-repository design, both from building it.

A build machine has its own credential and it is not a node's: read the
build queue, write the mesh exchange, nothing else. A node's queue
carries that node's declarations.

The answer goes through the exchange and never the default one, because
permission there is per exchange rather than per queue — anything allowed
to use it can publish into any node's queue. The price is that every
asker sees every result and filters by correlation, which is cheap
against a builder never needing that permission.

And every result is kept, failures included, because one that leaves no
trace is indistinguishable from a build nobody asked for. That is what a
builds view reads; the board page is corrected to say so.
2026-08-30 10:29:01 +02:00
jschoubben 4151c28927 The builder takes work over the broker, and what is still missing
A build is work, not state, and that is why it does not travel as a
declaration: as one it would either rebuild on every reconcile or carry
"and I already did this", which is state about an event rather than about
a machine. So it has its own queue and the answer comes back correlated.

Three properties recorded because they are decisions: acknowledge only
once the answer is away, one build at a time, and a failure is a result
rather than silence.

And the board page is corrected. A build result today is answered to
whoever asked and kept nowhere, so a builds view has nothing to read. A
record of past builds is the missing piece, not the builder.
2026-08-30 04:06:42 +02:00
jschoubben a625e6c709 A module repository, and two more shapes the host speaks
Designed with no reference to what came before, which was asked for. The
system this replaces has features — several deployable units inside one
module — and they are deliberately absent.

That closes something ADR 0001 has been carrying as an open prerequisite.
It lists "named features with per-node opt-in" as required, or "every
independently deployable unit becomes a module again and the count
returns". The premise was right and the remedy already exists in another
form: several modules, assignment per node, and a module with
requirements and no files of its own. `networking` is exactly that. The
count does not return because what made it return — a module is
expensive, so put several things in one — is gone. A module here is a
manifest and usually nothing else.

The manifest in a repository names artifacts; the manifest the mesh holds
names digests. Two documents, because a digest is not knowable until
something is built and a repository carrying one is wrong the moment
anybody edits anything.

The builder runs on a node. Building needs a container runtime and a
working tree, and what the control plane may send a machine is bounded by
the declaration language. A control plane holding a container socket
would be the one component that can do anything anywhere.

And the host's vocabulary grew from six shapes to eight — user and
archive — with the reasoning for each and for the refusals that came with
them. The count is asserted by a test precisely because every addition
widens what a compromised control plane can express.
2026-08-30 03:36:57 +02:00
jschoubben eab870fff9 A board, and the one constraint that is not a feature
Read from the board that exists. Eight sections; four are about work and
workers and are held back with that domain. The other four are the mesh
itself, and everything behind the main one already exists here — it is a
reader, not a second source of truth.

The constraint is the point of writing this down now. The existing board
is one service that reads every context's database, because that is the
shortest path to a page showing all of them at once. That is ADR 0008
violated by the one component with a reason to violate it, and the cost
is the same one the shared library has: a boundary nothing may cross is a
boundary that can move, and one thing crossing it is enough to freeze it.

So a board reads through interfaces and stores nothing. If a question is
slow, the answer belongs in the context that owns it, where everything
else asking gets it too.
2026-08-30 03:26:12 +02:00
jschoubben 5253742773 ADR 0024 — model access is a provision, and a licence has a name
A new requirement, and it is mostly a shape the mesh already has. A
module that needs to think requires model-access; several vendors and a
locally-run model are several modules providing it; choosing is assigning
the one you want. A model the mesh runs itself needs nothing new at all —
it is a mesh-scoped provision on the node with the hardware, credential
included.

A licence is a named thing because the whole point is saying which one a
given consumer uses, and the names are the operator's. Many to many, so
not a claim: two machines sharing an account is ordinary, not a
collision.

Four gaps, written as gaps rather than design:

- a provider that is on no node, reached over the public internet, which
  the reachability rule must not refuse
- a secret the mesh is GIVEN rather than mints. Every credential it
  handles today it generated and discarded; an API key arrives from a
  person, and accepting one must still discard the plaintext
- a consumer that is not a machine. Which licence a worker uses is a
  binding to an agent, and the provisions model has no consumer identity
  other than a node
- switching on exhaustion is a reaction to something observed, not a
  declaration. It belongs with observability, changing a binding — saying
  so is what stops the declaration language growing a conditional

The existing auto-refresh and switching is not being replaced because it
was wrong. It is being rebuilt because it lives somewhere that cannot
express the rest.
2026-08-30 03:06:36 +02:00
jschoubben a73014dcd5 A bare machine became a mesh, and something joined it
First end-to-end raise. A machine with a container runtime applied the
bundle its host carries and ended with a store, databases, schemas, a
broker holding a certificate it generated itself, and the control plane
serving. Then it took a token, checked the broker against the pinned
fingerprint, generated three keypairs and enrolled — the first node being
a node whose mesh is not up yet, observed rather than argued.

And a credential crossed. Declared the provider of a database for a
second node and pushed to over the broker, the machine ended with the
password in one file at mode 0600, and that password appears nowhere in
the declaration that crossed the broker, nowhere in the control plane's
database, and nowhere in what the node reported back. That is the whole
secrets argument, measured.

One fault, in the joining: the token did not say what the mesh calls the
machine, so enrolment needed a flag its own help said it did not, and
failed at the broker with an empty username. It is the fifth thing a
token carries now — the node cannot work its own name out, because the
broker account it authenticates as is named after it and exists before
the mesh has told it anything.
2026-08-30 02:37:37 +02:00
jschoubben 0f7e4ab597 The provisioner, which is where the mesh stops
A password nothing was told to create authenticates nowhere. The mesh
generates one, seals it to both ends and cannot read it — so it cannot
tell the software to accept it either. Something on the providing machine
reads what arrived and makes it true.

That something belongs to the module, not to the mesh. The control plane
decides and never touches a machine; a provisioner runs on the machine
and touches it. What the mesh owns is the contract: a manifest of who
asked and where each credential is, and one file per consumer holding it.

It reconciles and is never told what changed, which forces three things
that are each a fault somebody has shipped: set the password every time
or a rotation changes nothing; remove what nobody asks for or a departed
consumer keeps a login for ever; leave alone what it did not make or it
cannot be run on anything that predates it.

Saying where the mesh stops is the point. It decides, delivers, and can
prove what it delivered; the last inch belongs to whoever knows what
`create role` means.
2026-08-30 01:31:58 +02:00
jschoubben 82a5b9548a The secret is delivered without ever being held
Written after looking at how the existing mesh does it, so this is a
reaction to a measurement rather than a preference.

There, credentials sit in a column encrypted at rest. Its own tooling
records what that bought: the tool for finding a secret matches by value
rather than by name, because the same password is in three tables, in
each node's environment file in plain text, and inside every connection
string composed from it — copies its documentation calls the ones usually
in use. And a query against the encrypted column returns zero rows and
proves nothing, so auditing moved to the decrypted copies.

Encryption at rest addresses neither fault. The control plane can read
what it stores, so a copy of its database is a copy of everything. And
composition is what mints the untracked copies.

So the value is sealed to the node that will use it before it is stored,
with a key that node generated. Nothing central is composed. What it
costs is auditing by value, which was never real anyway; what stays
answerable is which node holds what, which is what rotation asks.

What remains is a provisioner. The mesh generates the secret and tells
both ends; nothing yet acts on the telling.
2026-08-30 00:21:52 +02:00
jschoubben cc872a58ce Binding is built except for the secret
Which turned out to be the useful way to cut it. A provider says what a
consumer needs in order to use it; a consumer says where it wants to be
told; the mesh adds which machine and what that machine is called on the
private network. So an app on one node reaches its database on another,
by a name the mesh also created.

The file says it carries no credential and why, because a missing field
looks like a bug and a stated absence looks like a boundary.

What remains is the secret itself, and the shape it will arrive in now
exists.

Also: two machines wired together across no private network is refused,
and that only became checkable when the network stopped being something a
machine has by virtue of holding an address.
2026-08-30 00:02:36 +02:00
jschoubben 80b74d32d6 Where the answer to a requirement is allowed to live
0009 distinguishes presence from instantiation — what the edge hands
over. It never distinguished where the thing on the other end is, and
that turned out to be the half doing the damage: a shell and a database
were both written `requires`, so requiring a database installed one on
every machine that used one.

A provided name now carries a scope, as a claim already does. Scope
belongs to the name rather than to each provider, or one requirement
means two things depending on which module answers it.

A requirement answered from the mesh is never satisfied locally. Nothing
provides it, and it says which module to assign somewhere; two do, and it
says how to choose. Choosing is recorded per node, because two machines
may reasonably use two different databases.

And knowing which node answers is the first half of handing a credential
back — you cannot be given a database's password before it is settled
whose database it is.
2026-08-29 23:52:17 +02:00
jschoubben 90ecfe6a01 An edge has two directions, and only one of them is built
0009 already said a consumer supplies a target and receives a name. What
it did not say is that those are two separate mechanisms.

Contribution — publish me at this name, on this port — now exists.
Binding — and hand me back a credential — does not, and is the larger
half: a secret has to exist, be stored, reach one node and not the
others, and rotate with every holder informed. That is the invariant set
found violated three ways at once, so it is not something to add in
passing.

The absence had a measured cost. Exactly two modules opened a direct
connection to the control plane's database, and they are the reason every
node permanently holds a credential to it. Both were doing by hand what
this edge is for. Neither needed a new kind of thing.
2026-08-29 23:36:23 +02:00
jschoubben 7fe2c31bdf Networking is a module, and what a domain module actually is
Two records, from building it.

0009 has a section titled "there are no domain modules", and `networking`
now exists. It is not a contradiction and it reads as one, so the
difference is written down: what was refused contains WireGuard and a
proxy and is assigned where half of it is unwanted. What exists contains
nothing — requirements and a name — so there is no half. Every artifact
it leads to is still an ordinary module assigned on its own terms.

With the cost stated, because it is real: adding a second implementation
turns a settled question into an open one for everyone using the bundle,
not only for whoever wanted the alternative. That is the refusing rule
applied consistently, and the alternative is a default, which is the
flavor field returning under a better name.

08-connectivity gains why the network stopped being code beside the
module system: a machine was on the private network because it had an
address, and there was no way to keep one off. A manifest can now say its
resources are computed, which is what a peer list needs.

And three modules rather than one, because WireGuard is one VPN of
several. Naming a module after the job and putting one implementation
inside it is flavor wearing a generic name — the second VPN has nowhere
to go.
2026-08-29 23:21:01 +02:00
jschoubben 554f6bd7a4 A capability may carry a value, and adding one is not free
Recorded while building the seat detector. A capability is a named fact about a
machine: its presence gates an assignment and its detail can carry a value, so
"can this run here" and "what should it be configured as" are the same fact
read two ways. A verdict has always had a detail beside its yes or no, so
panel: oled needs no new concept.

Two things that keep the set honest, both worth writing down before anyone adds
the fiftieth capability. It must be detected and the detector must say how it
knows -- so nobody can add one they cannot check, which is the whole of issue
007. And detectors ship inside the host, which is one static binary, so adding
a capability means shipping a new host everywhere. That argues for a small
general vocabulary rather than a specific one.
2026-08-29 21:12:21 +02:00
jschoubben f140303257 A module claims; it does not list its rivals. And flavor is retired.
Three decisions, all Jochen's, and the first is the one that unlocked it.

Exclusivity is not a property of a module. It is a property of a singular
resource the module takes over. Two shells compete for nothing and any number
may be installed; two display servers both want the seat. So a module declares
what it CLAIMS, and two modules claiming the same thing cannot both be assigned
within that claim's scope.

Not "xorg conflicts with wayland". Pairwise exclusion has a property that only
shows up later: adding a third display server means editing xorg and wayland to
know about it. Every new module requires changing modules nobody who wrote it
owns, and the edits grow as the square of the count. With a claim the third one
says what it claims and nothing else changes anywhere.

Claims have a scope -- node, site, mesh -- which is not new. The mesh already
enforces exactly one hub with a unique index. Scope is that idea said once
rather than hard-coded per case.

And some conflicts need no claim at all: two modules declaring the same file or
binding the same port are visible from what they declare. A claim is only
written for the abstract ones.

A requirement with several answers is refused, never guessed. One candidate is
assigned silently because there was no choice to make; none is refused naming
what is missing; several is refused naming them. That is what makes a solver
unnecessary -- counting candidates has no surprising behaviour, and a solver
can be added later without changing a single manifest.

Flavor is retired. It was carrying three unrelated meanings: variants of a
thing, a subset of a module a node installs, and whatever the current system
does, which earned two knowledge-base entries about going wrong. A word with
three meanings cannot be reasoned about. What it reached for is two ordinary
things -- different modules providing the same thing, and one module with a
setting.
2026-08-29 21:00:13 +02:00
jschoubben 974985b3d1 Four things the lab found about the private network
All on the first three machines to actually run it, and all invisible from the
mesh's own state: the graph was right, the files were right, the services were
up, every node reported success, and the network did not work.

A running interface does not re-read its configuration, so a node joining left
every existing node carrying a network that no longer existed. A hub sharing a
site with a spoke was emitted twice, which WireGuard refuses. Two nodes at one
site that neither can be dialled were peered directly, so nobody opened the
path and the more specific route blackholed -- this document's own warning
arriving in its implementation. And Docker sets the FORWARD policy to DROP, so
a hub with forwarding enabled still carried nothing between its spokes.

The last one is the sharpest: the substrate at tier 1 silently breaks the
network at tier 2, and nothing in either tier's state says so.

None of these is reachable by reasoning, and each was found within minutes of a
real machine trying it. That is the argument for the lab in one line.
2026-08-29 18:03:39 +02:00
jschoubben 6bcf0e4f9f Issue 010 fixed: origins keep the bundle and the mesh apart
The store records where each resource came from and each origin removes only
its own. Verified on the scenario that caused it -- eleven resources raised,
enrolled, sent the same two-resource declaration, and the store, broker and
control plane were all still running. A later declaration dropping a resource
still removed it, so removal by omission survived the fix.

Two more faults found while fixing it, both the same shape. A report published
to a routing key nobody bound vanishes: the broker accepts it, finds no queue,
drops it, and tells the publisher nothing -- so nodes announced what they had
applied into a void. And publishReport was discarding its error, so a node that
could not tell the mesh looked exactly like one that had.

Reports are mandatory now, so an unroutable one comes back and is said out
loud, and the binding covers every key a node may publish.
2026-08-29 16:43:28 +02:00
jschoubben 594ea10b07 Issue 010: the first declaration destroys the substrate
Found in the lab, doing the ordinary thing: raise a first node, enrol it, send
it a declaration. Both declared resources applied correctly and every container
on the machine was removed -- the store, the broker, and the control plane that
had sent the message. The link died mid-sentence because the broker carrying it
had just been torn down by what it carried.

Nothing is behaving incorrectly. Apply removes what the store holds and the
declaration does not name, which is what reconciliation means. The fault is
that the carried bundle and mesh declarations share one store, so the host
cannot tell what this machine raised for itself before there was a mesh from
what the mesh told it to have.

It is invisible until those two meet, which happens exactly once per mesh: on
the first node, after enrolment, the moment the control plane first speaks.

The report says what is not the answer, including the tempting one -- having
the control plane send the substrate back. It cannot: it was never told what
the bundle contained, and the bundle exists precisely because there was no
control plane to ask.
2026-08-29 16:23:01 +02:00
jschoubben 02afb7516b What connecting to the mesh is, and what a node presents
Two things this record never said, both asked directly.

Connecting to the mesh is one outbound AMQP connection from the node to the
broker, held open. There is no second connection and nothing is ever dialled at
a node. Being in the mesh means that connection is up.

Two different things ride on it and conflating them is what made this murky. An
AMQP account, which the mesh issues per node at enrolment, answers whether the
connection is accepted at all -- per node rather than shared, because a shared
one lets any node consume another's queue, which is the shared-credential fault
this record exists to remove reappearing at the transport.

The node's own keypair answers which node is speaking, on every message. It is
not made redundant by the account: with only an account the control plane knows
who is speaking because the broker says so, and that is the same transitive
authority this record already refuses in the other direction. A compromised
broker could attribute reports to whichever node it liked.

So a node holds two things after enrolment -- a credential the mesh issued for
reaching the broker, and a key it generated that the mesh only sees the public
half of. Both are its own, neither reaches anything else.
2026-08-29 15:25:18 +02:00
jschoubben 004057d85c A node's identity is a keypair it generates. This was never open.
I have been treating "what a node presents to prove it is that node" as an
undecided design question for weeks, and blocking on it. It was decided.
08-connectivity says of the overlay keys: each node generates its own keypair,
the private key never leaves the machine, the public key is published to the
mesh -- and says explicitly that this IS ADR 0004's "a node holds its own
identity", applied. Nobody had applied it to the thing 0004 is actually about.

What caused it was a word. The lifecycle said a joining node receives its own
durable identity, which reads as the mesh issuing something, and then the
question is what. The mesh issues nothing. A node arrives holding its identity;
what it receives is being known. That line now says what happens: it presents
the one-time secret and its own public key, which the mesh records.

The rule above it then holds literally rather than aspirationally. The mesh
stores a public key, so a copy of the mesh's database grants nothing, and
compromise of a node really is compromise of only that node.

Also recorded, since it was asked directly: same principle as SSH, own key, not
the machine's SSH host key. Host keys are regenerated by reinstalls and image
clones, which would silently un-enrol a node; their lifecycle belongs to sshd
rather than the mesh; and a partial host has no SSH daemon at all, so an
identity scheme resting on one excludes a supported kind of node.

The good half of that idea is kept: the mesh knows every node, so it can
distribute host keys the way it distributes authorised keys, and node-to-node
SSH stops depending on trust-on-first-use.
2026-08-29 15:21:36 +02:00
jschoubben 5fd522b8da A node is a machine; the session is a feature of it
Correcting an overstatement from the previous commit, where I had written that
a node IS a conversation. It is not. A node is a machine inside the mesh, and
the session is one of the things running on it -- like the host, like any
workload.

That also dissolves the conflict I flagged as unresolved rather than needing
anyone to decide it. 0001 says a node does not authenticate to a model
provider, agents do. Still true: the session authenticates, and the session is
not the machine. The node does not think, something on the node does. I had
manufactured the contradiction by promoting a feature into an identity.

0001's summary row is corrected the same way, and says explicitly that neither
the node's session nor a hired worker makes the node itself a thinking thing --
both run on a machine, which is what leaves that line untouched.
2026-08-29 14:22:29 +02:00
jschoubben 066f14b5f8 A node is a conversation, and that is not the employee model
Moving this out of 0003 and out of its vocabulary. I had spent three attempts
fitting the node's own session into the agent-as-employee record, each time
bending hired, draining, reassigned and retired to cover something none of them
describe. 0003 is back to its original text.

It belongs in 0004, under what a node is, because that is what it is -- not a
program installed on a node but part of the node. It holds one session
permanently, anything in the mesh can message it, and it remembers across
callers and across weeks. Its system prompt is the engram, which is recorded
here for the first time despite running on every node.

Also recorded: it has its own narrower tool list, so it can go and look rather
than only report about itself; there is no authorisation between nodes, because
every node is the operator's own; and how a node passes a question on is its
own business rather than a protocol field.

Switched off it still answers, and that is the point of having an off state
rather than an absent one. A node with nothing there is a silence somebody has
to diagnose. A node that says it is switched off is not. Same rule the host
follows about a service that does not exist.

0001's summary is corrected too: it had one row for "agents", which is the
conflation being complained about. Two rows now. A node's own session and a
hired worker are built from the same parts and run on entirely different terms.

Left standing and NOT resolved here: 0001 says a node does not authenticate to
a model provider, agents do. A node that holds a session does. That is a real
conflict between what is recorded and what runs, and it needs deciding rather
than a fourth reconciliation from me.
2026-08-29 14:17:25 +02:00
jschoubben 079c488d5e Provisioned and immutable beats exempt
Replacing the framing I wrote an hour ago. I had the node's own agent sitting
outside the lifecycle as an exemption, which is a rule somebody has to
remember. Provisioned the ordinary way and constrained is a rule the system
enforces, and it is one row like any other rather than a category every query
listing agents has to special-case.

It also reads the original sentence more carefully. "Exempt from the hiring
lifecycle" is exempt from hiring, not from having a lifecycle. Its lifecycle is
the node's -- provisioned at enrolment, retired when the node is retired. Same
states, a different thing driving them, and no exemption needed.

The constraints are now the four nonsense states written as things that cannot
happen rather than as an argument: not retirable, reassignable or deletable
while its node exists; exactly one per node. And a distinction that was missing
-- its existence is immutable, its engram is not. Freezing the personality
would remove the way a node is configured.

Disabling is the better half of this. A node with no agent is a silence
somebody has to diagnose; a node whose agent is disabled answers saying so,
immediately, with no model invoked -- the queue is still consumed and the state
is the reply. That is the host's own rule about a service that does not exist,
applied one tier up: absence must never be indistinguishable from a failure to
answer.
2026-08-29 14:06:33 +02:00
jschoubben fd7f7557bd The node's own session, and why it is not hired
Answering a question that was asked three times and that I kept not answering:
should the node's session just be an agent per node, since otherwise the
functionality exists at two levels?

Same mechanism, different lifecycle. A persistent session, accumulating memory,
a system prompt, a scoped tool list, addressable by message -- identical, and
building that twice is the duplication the question was worried about. What
must not be shared is the lifecycle, because if a node's own voice were an
ordinary hired agent it could be retired, leaving a node nothing can talk to;
reassigned, moving one machine's mind onto another; hired twice, with no answer
to which one replies; or never hired, leaving a node mute. The exemption in
this record exists to make those four unreachable.

I had this backwards earlier today and said so out loud: I called "a node
itself is an agent of a kind exempt from the hiring lifecycle" a fossil of the
old model and recommended striking it. It is the design. And it does not
conflict with 0001 -- "the two agent rows per node merge" means one per node,
not zero. I read merge as delete and invented a contradiction between two
records that agree.

Engrams are recorded for the first time. They are in use on every node and
appear in no record, which is how a decided thing comes to look accidental.
The engram is the node's system prompt, and it is what makes one node's answers
recognisably its own rather than generic.

Also recorded: there is no authorisation between nodes, because every node is
the operator's own and a prompt from one is a prompt from them. The consequence
is stated once rather than left to be discovered -- the mesh boundary is the
security boundary, which is what puts the whole perimeter on the token and the
overlay.

And how a node passes a question on is the node's choice, not a protocol field.
A node may say who is asking or may simply ask, the way a person relaying a
question decides how to phrase it. That follows from the engram. The cost is
that there is no machine-readable chain of who ultimately asked; each node
still holds what it was asked and by whom.
2026-08-29 14:01:00 +02:00
jschoubben 88ba81e9c1 Agents reaching nodes is the capability, not a hole in it
Correcting what I wrote an hour ago. I had recorded node-to-node SSH as "not a
mesh function" and "a second control path through the back door", reasoning
from ADR 0004's rule that the host has no inbound control surface. That
conflated two different things and got the product backwards.

There is no node-to-node SSH to forbid. The actor is always an agent; a node is
only where it happens to be running -- ADR 0001 already says a node is a place
where an agent can run and that is the entire relationship. An agent hired onto
one node reaching another to do work is the capability the whole arrangement
exists to provide.

The credential is the agent's, in its own credential directory, which ADR 0001
already established. So a node's authorized_keys lists agents and never nodes,
and three things follow: no node holds a key reaching another node, so 0004's
"a node holds its own identity and nothing else" stays literally true; a
compromised node costs the credentials of the agents that were on it rather
than a way into everything; and who may reach what stays a mesh-wide fact,
which is why it is identity's.

The rule I misapplied is about how a node's declared state changes -- over the
broker, never by being dialled. An agent with a shell is not the mesh
reconfiguring a machine, it is what a person with a terminal has always been,
and this design already depends on that working: the overlay is the way back in
when a declaration breaks something. What such a session leaves behind is
drift, and drift is what reconciliation is for.

0001 also stops underselling the fourth layer. It read as "the layer the other
three exist to carry", which is true and flat. The value is that an agent can
work across a set of machines as though they were one -- centrally configurable
machines are ordinary; that is not.
2026-08-29 13:06:40 +02:00
jschoubben 918dc04916 What this actually is, and three things that were assumed
Four things settled by talking them through, all of which had been true in
somebody's head and written nowhere.

It is not a mesh in the peer-to-peer sense and will not become one. 0001 now
says what it is instead: machines linked by a private network, one node holding
knowledge of all of them, modules as the way anything is built and delivered,
and agents hired onto nodes to do the work. The word describes what machines
can reach, not how they are governed. "Master" overstates it the other way --
nothing needs that node to keep running, only to change.

0006 gains the option that would make it a real mesh, recorded as considered
rather than rejected by silence: every node holding the whole inventory, a
replication process, an elected master with promotion on failure. What settles
it is not the complexity but that it still would not deliver the name, because
application databases are not replicated -- so a genuine peer-to-peer mesh
means becoming a replicated database system for every consumer's data too. That
is a larger product than the thing it would support.

Also in 0006: three central roles, not one. Losing the control plane costs
change, losing the broker costs being told anything, and losing the hub costs
nodes in different places reaching each other at all -- which is operation, not
administration. Whether they are one node is not decided.

And SSH access is identity's. It appeared three times as something that uses
the overlay and never as something the mesh provides, which reads as settled
when nothing decided it. Nobody else could: the mesh is the only thing that
knows which humans and agents exist and which nodes they may reach. Node to
node SSH stays out -- the host has no inbound control surface by decision, and
nodes reaching each other that way is a second control path through the back
door.

0007 gains the requirement underneath all of it. Reachability was recorded as a
fact to track and never as a thing some node must have. The broker's node and
the hub must be dialable by every node at a stable address, or nothing can join
and a disconnected node cannot return. A mesh entirely behind NAT cannot be
raised. That is a precondition and it belongs with the others.

The link staying on the underlay is also argued now rather than asserted. At
join time it is forced; afterwards it is a choice, and the reason is that a
repair channel carried over the thing being repaired is not one. Moving it onto
the overlay, with fallback, is recorded as open with what it would have to get
right -- a WireGuard interface has no link state to test, and a silent fallback
is this repository's recurring fault in a new place.

0010 says in one line what was the intention throughout: the module system is
the CI/CD. Not a pipeline beside the mesh. Build, test, publish and deploy are
one reconciliation seen at four points, which is why a thing that cannot be a
module cannot be delivered.
2026-08-29 13:01:36 +02:00
jschoubben 5218b06c02 Fold the control plane's build decisions into 0006 and 0008
Back to 23 records. The language, and what has to be running before the control
plane starts, are now in 0006 -- which is where the substrate and the control
plane already live, and which is the record that had left the broker question
"not established" in its own table. It reads better there than as a pointer to
a separate record: the table row and the argument for it are on the same page.

The store mechanics went into 0008. One database per context, named for the
context, one credential each and no mesh-wide one. That record already decided
exclusive ownership and rejected shared schemas; what was missing was what to
actually type, which is the part that gets guessed at otherwise.

Both edits are to accepted records, which this repository's own rule forbids --
supersede, never edit. Recorded here so it is visible rather than silent. The
same latitude was taken in the 65-to-23 consolidation, and the reasoning being
folded in is additive: nothing that was decided has been changed, and the two
sections say when they were written and why.
2026-08-29 03:08:50 +02:00
jschoubben 82a3065f82 Tier 2 exists, and the token was missing a quarter of itself
mesh-control is built as far as it can honestly go: one context of seven,
inventory, with its schema and the command that applies it. The repos map and
the control plane design say so, and point at ADR 0024 for what it took.

Separately, and more importantly: this repository described the enrolment token
as carrying three things when ADR 0004 says four. The missing one is the
control plane's signing identity -- the reason a node does not have to trust
the broker it dials.

Without it the control plane's authority is transitive through the broker, and
0004 spells out what that costs: a compromised broker could forge declarations,
and since the host applies whatever the link delivers, that is the whole
machine. The record has the argument in full; the design doc had dropped the
conclusion.

Found by reading the two together while deciding what the control plane must
store, which is roughly the only way it would have been found -- both documents
are internally consistent and only disagree with each other.
2026-08-29 02:49:58 +02:00
jschoubben 84f4425fd6 The broker precedes the control plane, and it is written in Go
Two things found by trying to build tier 2.

The substrate design asked whether the message broker has to be running before
the control plane, and framed it as depending on whether the control plane's
own parts talk to each other over it. They do not -- it is one process -- so
under that framing the broker stays out of the bundle.

The framing cannot answer the question. What decides it is how the control
plane reaches a node, and the answer was already decided: only ever over the
link, and the link is the broker. So provisioning the broker would require the
broker. The first node does not escape this by being local, because it enrols
the ordinary way, by dialling the broker at the address in its token -- which
was deliberate, and worth keeping.

The bundle is two images now. The record says what that costs, including a
certificate the broker needs at a moment when there is no mesh to issue one.

The language had never been decided for tier 2. Go, for the same reason the
host is: the bundle pins this image by digest and runs it where nothing can
check it, so the image should hold the program and nothing else.

Also corrects something already built: the bootstrap created one database and
called it 'mesh'. ADR 0008 grants a context only what it exclusively owns and
ADR 0006 says the mesh database names a thing that will not exist. One database
per context, so one today, called inventory.
2026-08-29 02:32:46 +02:00
jschoubben 6a2b107fb8 Restore a consequence the consolidation dropped
I said nothing was lost when 65 records became 23. That was too strong, and
here is a counterexample: ADR 0046's consequence that the lab needs a way to
place images did not survive into the merged substrate record. The compression
kept the decision and dropped one of the things it implied.

It was not lost from the repository -- 04-ISSUES/009 had already picked it up,
which is why it was found at all. But the record no longer carried it, and the
record is where somebody would look.

Restored, now as a resolved fact rather than an open consequence: the lab
raises a registry inside the scenario, which is the real path since that is
what every node after the first pulls from. The digests it serves are its own,
and that satisfies the pinning rule -- what is required is a reference that is
exact and cannot move.

Worth recording the wrong assumption too, because it is what made this look
impossible for two days: I took "pinned by digest" to mean the UPSTREAM digest
had to be preserved. It does not. Any digest that is exact and immutable
satisfies the rule, and a registry assigns one.
2026-08-29 00:06:05 +02:00
jschoubben 087a8f4144 Close 009: a sealed machine now pulls by digest
The resolution was the one the issue predicted -- a registry inside the
scenario -- and it is the real path rather than a stand-in, since that is what
every node after the first pulls from.

The digests are the lab registry's own, which satisfies the pinning rule: what
is required is a reference that is exact and cannot move, and one this registry
assigned is both. That was the insight that unblocked it; I had assumed the
upstream digest had to be preserved, which is what made it look impossible.

The fault worth keeping is recorded in the issue: the read-back checked that
the catalog endpoint answered by matching the substring 'repositories', which
an empty catalog also contains. It passed on a registry holding nothing. This
repository's own subject, arriving in the tooling built to catch it.
2026-08-29 00:05:20 +02:00
jschoubben b4607dfc03 Numbers are identity; the reading order is a generated, checked index
Decided after measuring what renumbering actually costs: 96 references in code
comments across two repositories, none of which would have failed to compile.
They would have pointed at the wrong reasoning, which is worse than a broken
link because nothing reports it.

So a number identifies a record and never changes. It cannot also be a
position -- a position moves when the set changes, and an identity that moves
is not one.

The reading order moves into an index generated from each record's `topic:`.
Six topics, in the order somebody learns the system.

The index is WRITTEN rather than only generated on demand, which reverses what
this repository previously said. The reason it said otherwise is that a
hand-written index drifts -- but a reader looking at the folder on a forge sees
the folder, not a command, and the drift objection is answered by checking
rather than by refusing to write one. That is §5's own rule: a rule states how
it is checked.

Two checks, both confirmed to bite. index.py fails when the written order no
longer matches the records. records.py fails when a record has no topic or one
nobody defined -- the quiet failure being a record that vanishes from the order
rather than appearing in the wrong place.
2026-08-28 23:39:18 +02:00
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00