Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24

Merged
jschoubben merged 177 commits from reconcile-init-into-main into main 2026-09-05 10:27:11 +00:00
177 Commits
Author SHA1 Message Date
jschoubben 495db6dc89 Restore initialization's issue 003 — the status flip was based on a wrong premise
Init's 003 was already resolved with its own consolidated attribution
(03-DESIGN/01-to-be/08-connectivity.md); the re-homing overwrote it with the
session's ADR-0045 firewall attribution. Init is canonical and the firewall
decision is recorded in the ported ADR 0045 regardless, so 003 is restored
untouched. No initialization record is modified by this reconciliation.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 12:26:23 +02:00
jschoubben e269f9a185 Re-home this session's new ADRs (0039-0049) and issues (032-037) onto the consolidated scheme; flip issue 003; port repos.md sdk line + feature-branches playbook (07); regenerate index
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 12:24:07 +02:00
jschoubben 546caa31c4 Merge initialization into main — adopt its consolidated structure as canonical
The real work lived on initialization (consolidated decisions 0001-0038, the
fuller issue set 001-031, the control-plane/substrate/node-lifecycle/delivery
design, research 011/012, the checks tooling). main had diverged onto a stale
base and only carried this session's genuinely-new work. This merge makes
initialization's tree canonical on main; this session's 11 new ADRs and 6 new
issues are re-homed on top in the following commits. initialization is recorded
as a parent so its history is preserved in main's ancestry.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 11:59:01 +02:00
jschoubben 9638aa02c5 031 — a machine becomes each thing it was told, in turn
Declarations queue and the host applies all of them, oldest first, so a
machine pushed five things in a minute spends five applies becoming the
last one. Correct every step — each declaration is the whole machine —
and wasted in all but the final step.

Only visible since a report names its declaration: the reports arriving
were about ever-older ones, while timestamp comparisons used to happen
to pass. The fix is consumption order, not the queue; the open question
is what a superseded declaration's report should say, because silence
reads as disobedience and "applied" would be a lie.
2026-09-02 01:06:12 +02:00
jschoubben 6bdd3f2bfa 030 — asking what a machine should be re-signed its certificate
Composing a declaration signed the machine's certificate anew each
time, and a signing carries a fresh random serial — so what the mesh
would send differed from what it had sent by one byte, for ever, and
every machine carrying a certificate stood eternally waiting.

Third find of the same rule: issued once and kept. The port had it, the
secret had it, the certificate composed fresh on every asking — and the
keeping column had existed since the serving-key migration, written by
nothing, the same shape ReleasePorts was found in.

Found by keeping the scenario standing and diffing two plans seconds
apart: one line, where four theories had none.
2026-09-01 22:28:26 +02:00
jschoubben 6385d52bc0 029 fixed — the store is added, not built, and the cycle is refused
The registry module names its image by digest, the way the bundle names
the three a first node starts from, and goes in as a manifest. The lab
now walks that path and passes.

A manifest that provides the store and also builds artifacts is refused
at parse. The lab had been passing only because its Docker Hub stand-in
quietly received the push — a prop covering for the thing under test.
2026-09-01 21:43:42 +02:00
jschoubben 013dbce4a9 029 — the artifact store cannot be delivered by the artifact store
Installing the module that provides `artifact-store` requires something
that provides `artifact-store`: its image is mirrored in, mirroring
publishes to the store, and the builder refuses to run without one.

Never seen, because the lab always has a registry standing before the
mesh asks for one, and so does any mesh built on a machine that already
had one.

What it blocks is larger than a registry. The store holds images and
packed archives both — it is the module catalogue in artefact form — so
until it exists a mesh can run only what its bundle already carries.

The substrate record already answers it. It asks of each candidate
whether it can grant itself the thing it provides: the store cannot
create its own database, the broker cannot create its own virtual host,
and the registry cannot grant itself a repository. So the registry
module names its image and is never built.

With a limit worth saying out loud rather than discovering: a module
providing the store may not build artifacts of its own, a UI or a tool
server included, because there is nowhere to put them until it runs.
Such a registry is two modules.
2026-09-01 21:15:05 +02:00
jschoubben e75683c01f 026 reopened — the rule is right about data, wrong about facilities
Enforcing "declare what you mount" refuses the builder, which mounts the
container runtime's socket. That socket is not the builder's data: it
exists already, the machine owns it, and declaring it as one of the
module's directories would be a lie the host would act on.

Two kinds of mount are spelled identically today — the directory my data
lives in, and a machine facility I was granted. Until a manifest can say
which, the fourteen declared mounts are right by coincidence, which is
what this issue was opened about.

Recorded rather than decided: separating them is new vocabulary, and
inventing it to turn a check green is how a mechanism nobody chose ends
up load-bearing.
2026-09-01 19:45:55 +02:00
jschoubben bd8f09d647 026 fixed — and the coincidence turned into a rule
The fourteen undeclared mounts are declared. More to the point, a
manifest that does not declare one is now refused: they were right by
coincidence, and a checklist nothing enforces is a checklist that is
true until the next commit.

Refused in the control plane, because the machine cannot tell the
difference — asked to mount a path that does not exist, it makes the
directory, which is a thing it is perfectly able to do.
2026-09-01 19:36:15 +02:00
jschoubben 0bcfeb4e80 028 fixed — the mesh assigns the port, and knows what it cannot move
A module now says its port once, in `listens`, and the container's
mapping, the rule set and what a consumer is told are all derived from
one assignment. The three hand-written copies that agreed only because
one person wrote them are gone.

The half that made this an issue rather than an inconvenience was that
the substrate is not a module: nothing in the mesh had heard of its own
store, so it handed a database module the port the store already had.
The machine now says what it carries, and the mesh assigns around it.

Ports the protocol fixes became claims, which needed no new mechanism —
the mesh already had one for what is singular on a machine.

Left open, and unchanged by any of this: whether a module should publish
to the machine at all. Assignment makes publishing safe without making
it necessary.
2026-09-01 18:32:58 +02:00
jschoubben c2a37ab7f8 38 — the mesh assigns the port, and a module does not care
Jochen's call, and the right one: a module cannot choose a port well,
because it is written once and assigned anywhere. Any number it picks is
a guess about a machine it has never seen, and two modules guessing the
same number is not a mistake either of them made.

Writing it up turned up something the issue had missed. The same number
appears three times in every module — the rule set, what a consumer is
told, and what the runtime publishes — and nothing checks that they
agree. They agree today because one person wrote all three. A module
whose `serves` said one thing and whose container published another
would resolve, compose, apply, and hand every consumer a port that
answers nothing.

So the decision is one source with the other two derived, and an
assignment made once and kept, as a credential is.

The part that needed thought is ports that cannot move — mail on 25,
submission on 587. Those become claims, which is what the mesh already
has for what is singular on a machine. Two modules wanting 25 is the
same shape as two wanting the seat, and gets refused by name at
assignment rather than by a container runtime at apply. That makes this
mostly a matter of pointing an existing mechanism at ports.

Left open: whether a module should publish to the machine at all.
Assignment makes publishing safe without making it necessary.
2026-09-01 17:42:27 +02:00
jschoubben 6330abce5f 028 — two things want one port, and nothing says so until the machine
Found by fixing 027 and pushing again. The declaration is now accepted
and the database container still cannot start: the mesh's own store
holds 5432 on that machine, and the module publishes 5432.

Nothing catches it because the substrate is not a module. It arrives
from the bundle before there is a mesh to ask, so the control plane has
never heard of the store and does not know it holds a port. Resolution
can compare modules with each other and cannot compare one against what
the mesh is built on.

Nor does it compare modules with each other. A port is exclusive on a
machine in exactly the way a claim is, and the mesh has a mechanism for
that which ports do not use.

It has been met before: the end-to-end test that exercises a real
database publishes 5433 rather than 5432, inline, with nothing saying
why. That is how a constraint becomes folklore.

The open question is bigger than the bug. Whether a module should
publish to the machine at all decides how a consumer reaches it, and
changes what `serves` means.
2026-09-01 17:21:31 +02:00
jschoubben 03266fd4a2 027 — a container cannot follow a file, and a rotated credential is the case
Found by the forge failing to start. I had put `restart-on` on nine
containers so they would pick up a rotated credential; it belongs to a
service, and the host refused the whole declaration.

Removing it fixes the modules and leaves the reason I reached for it.
The mechanism is written against exactly this, in the host's own words:
a running service does not re-read its configuration, so replace the
file, find it already running, do nothing, and the machine keeps
behaving as before while every check passes. Every word of that applies
to a container, and nearly everything the mesh runs is one.

The cost is concrete. Rotation replaces the file and tells the provider
to accept the new credential. A provider reconciles, so it takes it. A
consumer is usually a container, so it does not — and the two ends hold
different passwords, which is the fault ADR 0001 records costing two
days. The test that proves rotation works uses a consumer that reads the
file on each attempt, so it does not meet this.

Two things a fix has to keep: it stays declared state rather than a
command, because the link may not carry an action; and where an env-file
changed, the honest verb is recreate rather than restart, because a
container's environment is fixed at creation.
2026-09-01 17:19:43 +02:00
jschoubben 3f00d3b413 026 filed; 025 corrected; the image store is a module
Three corrections, two of them to things I wrote today.

026 is the serious one. Four modules mounted fourteen host paths nothing
declared — the mail spool, the databases, the object store's data. The
runtime creates those as root, so owner and mode go unapplied, and the
rule that keeps a directory holding data the mesh did not put there is
written in terms of declared directories. It reached the configuration
and missed the data. The cause was carrying compose files across: a
container shape that can express one gets filled in like one.

025 claimed nothing turns a tag into a digest. That is false, and the
answer was designed and built before I wrote it. A module names an
artifact, not an image, and `kind: upstream` mirrors somebody else's
image into the mesh's own registry, pinned by the digest it lands with.
The two-document split the issue described as the shape of a fix is the
design. Pinning twelve images by hand was treating the symptom, and left
them pointing at a public registry rather than the mesh's.

And the image store was written up as something the mesh does. It is an
ordinary module — considered for the substrate and removed, because the
test is whether the control plane needs it before its first instruction,
not whether it can grant itself one. So somebody's own registry is the
same module as the mesh's.
2026-09-01 16:10:13 +02:00
jschoubben befbb1d978 The five modules could not have run, and now the forge does
Recorded against phase 3, because the phase note said running them
needed images stocked and provisioners built — and missed that not one
of them named an image that exists. Sixty-four zeros where a digest
belongs, eighteen times, parsing and resolving perfectly.

The forge now runs: on a database another module provides, with a
password it did not choose and a connection string it could not have
written itself. First of these descriptions to be started rather than
planned, and it exercises the whole of the credential work.

What remains is the mechanism rather than the data. Nothing turns a tag
into a digest as part of the mesh's own work, so it was done by hand —
which is what the issue says a person should not be asked to do. Asking
a registry takes a second and pulls nothing, so the main argument for
leaving it undone is gone.
2026-09-01 15:40:52 +02:00
jschoubben 2bbdc52440 025 — half done: nothing unrunnable reaches a machine now
The refusal landed, and the examples pin images that exist. What is
still missing is the part that makes it unnecessary: nothing in the mesh
turns a tag into a digest, so it was done by hand — which is precisely
what the issue says a person should not be asked to do.

Recorded because the resolution mechanism turned out to be trivial:
asking a registry what a tag points at takes about a second and pulls
nothing. That removes the main argument for leaving this open.

Also records the two faults that fell out of pinning for real. The mail
system named seven repositories that do not exist, because it publishes
to a different registry than the manifest assumed, and one of the seven
had been renamed upstream. Nothing checking only the shape of a
reference could have found either.
2026-09-01 15:13:52 +02:00
jschoubben fc7717f6b1 021 closed — the record was left open after the fix landed
The code and its tests went in hours before the record was touched, so
an issue that read `located` had been fixed all along. That is the exact
failure the frontmatter exists to prevent: status is meant to be
answerable from the record rather than by reading the code.

Closed with the commit that did it, and cross-referenced to 022 and 023,
which came out of the same mistaken instinct — treating the machine as
a boundary, then as an identity, then finding a consumer had a password
and no name to present with it.
2026-09-01 15:04:36 +02:00
jschoubben 0871e6ec11 37 — where a module lives, proposed
The module descriptions sit in `examples/` inside the control plane, and
that name has been doing harm: everything there reads as a sketch, and
one shipped naming a container image nothing builds. A directory called
the catalogue would have made "does this work" the obvious question.

The shape of the answer turns on one measurement. Of the 126 modules in
the system being replaced, 47 are software in their own right — the
largest is 182 source files, and a speech-capture module carries a whole
daemon. Another 44 ship helper scripts. Only 35 are a description and
nothing else.

So a catalogue cannot be a folder of manifests, because two thirds of
modules are programs. That splits them four ways, and only two of the
four belong in a catalogue: things the world made that we describe, and
packages with some files. What the mesh is made of stays in the
repositories that build it. What we wrote keeps its description beside
its code, in the same commit, because nothing else can stop the two
drifting.

The mesh's list of modules is a table, not a repository, and it already
records where each module came from and at which commit. Nothing needs
inventing for modules from anywhere; a repository of ours is just the
source we curate.

The check that a description is valid should move to a command on the
control plane's binary. Today a test reaches into the control plane's
internals to parse manifests, and another reads its build file to check
images exist — two jobs tangled. A command would also give the same
check to somebody describing their own application, which is the case
that matters most and has none.

Left open: how a provisioner's image gets published and pinned, and
whether thirty-five install-a-package modules deserve to be modules at
all.
2026-09-01 14:24:53 +02:00
jschoubben f6b1834ea9 024 fixed — the registry was addressed by hand and nothing else was
The cause was one line. Machines get a systemd-networkd unit with a
static address, so networkd finishes and reports the link configured.
The registry ran `ip addr add` inline, which leaves networkd waiting to
configure something it was never told about — and
systemd-networkd-wait-online has an infinite timeout.

So network-online.target was never reached and everything ordered after
it never started. On these machines that is Docker, so `docker load`
blocked on a socket whose daemon was queued behind a target that would
never come, and three bounded timeouts stacked to thirty-five minutes.
These machines have no DHCP by design, so that wait was never going to
end.

The hypothesis in this record was wrong and the record now says so.
Stocking had just been changed, so stocking looked guilty; stocking
takes 34 seconds and always did, timed directly before changing
anything.

Fixed with two things that made it cost hours instead of minutes: an
image placement now waits for the runtime and refuses after 120s naming
what systemd is waiting on, and the end-to-end test passes onProgress —
the raise reported every step and the test discarded it, which is why
thirty-five minutes and four minutes of silence looked the same.

The suite then ran to completion, 23 of 24, the one failure a check of
its own flagging a path as a credential because `/` is in the base64
alphabet.

Also recorded: a redirected log lags, because Node block-buffers stdout
to a file. Read as a stall twice, the second time right after the real
fix — where a buffering artifact argues the fix did not work.
2026-09-01 09:53:08 +02:00
jschoubben 407416e6d0 024 — a run stalls before the host is placed, and says nothing while it does
Seen twice today. Once mid-run: thirteen passes, then the process ended
with no summary, no failure and no receipt. Once from the start: the
first test ran 35 minutes against a measured 4.5 and was still running
when it was stopped.

Ruled out rather than assumed: not memory (84 GiB free, no OOM), not the
daemon (the stalled machine answered `incus exec` immediately), and not
the changes under test — the anchor VM had no host log and no
containers, so the run never reached placing the host.

What changed just before is that the rebuild went from two artifacts to
six, and every one of them is pushed into the scenario's registry, which
is the step the second stall sat in. Recorded as what changed, not as
the diagnosis.

The reason this is an issue and not a slow test: the suite prints
nothing between starting a scenario and finishing its first test, so
four minutes and thirty-five look identical from outside, and the only
recourse is to guess. That is how a workstation was left unbootable in
August. And a run that ends silently after thirteen passes is a run
somebody may believe.
2026-09-01 03:20:40 +02:00
jschoubben 39916b26e9 What stands between 3.1 and a running identity provider is a program
023 is fixed, so the design faults are gone and one concrete thing is
left: the realm provisioner does not exist. Its manifest named an image
nothing builds and no program backs, which has been removed — a manifest
describing a program nobody wrote is the same mistake as the credential
files that could never be read.

Keycloak's manifest now says what is true today, and the gap is loud: it
no longer claims to provide oidc-client, so a consumer asking for one is
refused by name at plan time instead of resolving cleanly and waiting
for a client nothing will create.

The provisioner should be written against a real Keycloak in the lab
rather than from the API documentation. The object store's took three
corrections that only a running server produced.
2026-09-01 03:13:11 +02:00
jschoubben 0760280bc6 The largest gap in the coverage list is not a gap
Tool servers — 56 modules, over half — were written up as the biggest
missing thing. They are expressible with what exists, and the first
framing was wrong in a way worth keeping: a module provides `tools` and
the session requires them does not work, because a requirement has one
answer and 56 modules offering tools would be 56 answers.

Turned around it fits exactly. The session provides `tool-host`; every
module offering tools requires it and contributes where its tools are.
Many-to-one is what `contributes` has always been, and the session
receives all of them in one file. Verified by resolving it rather than
by reading the code.

It only became possible today: until 022, several modules on one node
requiring the same thing was refused outright. Worth noting because it
means the credential fix bought more than credentials.

What remains is a decision about what a tool server is, which is work
rather than a missing shape.

The entry stays in the list rather than being deleted — a checklist that
quietly loses its biggest item reads as though nobody looked.
2026-09-01 03:09:11 +02:00
jschoubben f583502fc9 023 fixed — the mesh says who a consumer is, and what it is bound to
Both halves had one cause: the mesh knew something and did not say it.

Who a consumer is now comes from one derivation, sent to the provider in
its grant and to the consumer in its binding, so the two agree by
construction. The provisioners use the name they are given and refuse to
invent one, because a name of their own would create a login the
consumer could never guess while everything reported success.

Bound values reach the file that needs them through the symmetric twin
of the sealed placeholder — simpler, because they are not secret, so the
control plane fills them in and the host gains nothing.

The lab run meant to prove this failed in a way that looked like the fix
being wrong: rotation could not authenticate against a real database.
The cause was the suite rebuilding the control plane's image and not the
provisioner's, so an image built that minute ran against a provisioner
built the day before. That is 005's family and is recorded with the
issue, because the misleading part is worth more than the fix.
2026-09-01 03:05:43 +02:00
jschoubben 4adc656b54 Where Phase 3 actually stands
All three modules have manifests, all three parse, resolve and plan, and
none of them can start. Worth writing down before it reads as progress
or as failure, because it is neither.

The vocabulary held. Nothing in 3.1–3.3 needed a new shape — including
the mail system's several containers on a private network, which was the
one expected to break it. That was the question this phase was designed
to answer.

What did not hold was underneath: 022, now fixed, and 023, open. Both
are about credentials rather than about what a module can say.

The third fault was in the manifests, not the design: a secret declared
at a path named .env and read as one, when a sealed file holds a
password and nothing else. That is what a manifest checked only by a
parser buys, and it is why there are now two tests reading the manifests
on disk.

023 is the whole of what remains before the identity provider runs.
2026-09-01 02:56:54 +02:00
jschoubben dea46c3c6c Correct what the coverage document said about actions
I wrote that a module cannot declare an action. It could — the parser
accepted one, and the refusal only came on the machine. The claim was
wrong in the direction that matters: it read as "the design prevents
this", when what prevented it was a check at the far end that nobody
would connect back to the manifest.

Health checks are still the gap most worth closing, but the shape of the
answer is different from what I wrote. An action is not available to a
module at all, so a health check needs a way to say ask this and expect
that without saying run this — closer to a listens entry than to an
action.

Also records the finding itself, because it is a recurring shape here
and not a one-off: a rule enforced only at the far end is enforced and
unusable.
2026-09-01 02:56:05 +02:00
jschoubben 9073d3f2df Say plainly that env-file never points at a sealed secret
The playbook offered `env-file` and `${secret:name}` as alternatives,
and that reading is what produced the bug every example module shipped
with: own-secrets pointing at a path named `.env`, mounted as env-file,
holding a bare password. The container starts with no password set —
which is a service running on the wrong credential, not a failure.

They are not alternatives. A sealed file holds a password and nothing
else, so env-file points at a file the module declares whose content
leaves a hole, and the host fills it on the machine. A provisioner is
the exception, because it reads a password file.

Written out as the three lines a module needs, with the failure it
prevents named, since the abstract version was already there and was
read the other way.
2026-09-01 02:52:55 +02:00
jschoubben 2baf22ac43 A pair is a module and a provider, not two machines
Amends the credentials page, which said "every pair has its own
credential" and meant two machines. Built that way, it was wrong in a
way that only shows on a real node: a machine running several services
against one database server had one credential between them, so the
provider refused to plan at all and the consuming node quietly gave the
first module a credential and the rest nothing.

The page already argues the case against itself — one credential with
many holders is the first of the three faults it was written to remove.
It just drew the boundary at the machine.

Two modules on one node are as separate as two on different nodes, and
one login opening both is what this page exists to prevent. It is also
what makes withdrawal possible: one role per machine cannot say that
this module has lost its login and the others still have theirs.
2026-09-01 02:46:37 +02:00
jschoubben 80f18caf03 022 fixed; 023 filed — a password is not a connection
022 turned out to have a silent half worth recording: the provider
refuses loudly and names the modules, which reads as a decision, while
the consuming node does not refuse at all. Three modules wanting one
database produce one need, so two of them get no credential file and
each starts and fails to authenticate with nothing saying why.

023 is what remained after fixing it. A consumer now gets its own
password, in whatever shape its configuration wants, and still cannot
connect: the user name is invented by the provisioner and recorded
nowhere in the mesh, and the host and port sit in a JSON binding that an
application reading KEY=value cannot use.

The asymmetry is backwards and the coverage document now says so. The
secret is the hard case, because the mesh must not be able to read it,
and the secret is the part that arrives. The host and port are ordinary
facts the mesh holds in the clear, and they are the ones stuck.

Keycloak, Gitea, Mailu and MinIO all parse and resolve and none of them
can start. This is what stands between the module set and a running one.
2026-09-01 02:40:40 +02:00
jschoubben c86adbe3cc 022 — a credential belongs to a node, so a second consumer refuses
Found while checking whether the module vocabulary covers real use
cases. A node running three modules that all want a database cannot be
planned at all:

  anchor has 3 modules asking for "postgres-database" and they would
  share one credential: gitea, keycloak, umami

The refusal is right about what it says and wrong about what it implies.
They would share one credential, and sharing is worse than refusing —
but the arrangement being refused is the ordinary one, and the node this
mesh exists to take over runs eight modules against one database server.

The cause is the key: a credential is keyed by provision, consumer node
and provider node, so `consumer` is a machine. The provisioner inherits
it and names the role `mesh_<node>`. The refusal is not a check that
caught something; it is the only honest thing that function can do with
a key that cannot tell two consumers apart.

It is the same mistake as 021 with a different face. There the machine
was treated as a trust boundary; here it is treated as an identity, as
though "who is asking" is answered by naming a host. Two modules on one
node are as separate as two on different nodes.

Worth stating plainly: without the refusal, gitea's login would have
opened keycloak's database, and nothing would have said so — from the
provisioner's side it created exactly what it was asked to create.

Not a local fix. It crosses the control plane, the grant file naming and
every provisioner that names something after a consumer.
2026-09-01 02:28:39 +02:00
jschoubben 36d342f176 What a module must be able to say, measured against 127 that exist
Every manifest in the system being replaced was read and every key
counted, then set against what the new one can express. Three findings
worth more than the table.

**The most-used key was already covered and I expected a gap.**
Depending on another module — 65 manifests, the commonest thing any of
them says — is a requirement naming a module, which already means that
module rather than anything providing the name.

**The largest real gap is tool servers: 56 modules, over half.** A
module can already run one; what is missing is anything saying it offers
tools. That is plausibly a provision rather than new vocabulary, which
would need nothing added — not yet decided, and recorded as undecided.

**The gap most worth closing is health, at seven modules.** The mesh
knows a container is running, which is not whether it answers, and this
project has paid for that distinction twice. An action with a verify is
exactly the right shape and may not arrive over the link, so a module
cannot declare one.

Two things are missing deliberately and say so: stage hooks, because the
link may not carry an action and a module needing setup ships a program;
and flavours, retired in favour of claims.

Config merging is missing and should stay missing. A mechanism that
understands TOML gets asked for YAML, then INI, which is how the thing
being replaced became unholdable.

Also records what the survey found that is not about coverage: manifests
that had stopped matching what was actually brokered, one fact derived
in two places giving two answers, and a live listing returning
credentials in plaintext.
2026-09-01 02:20:34 +02:00
jschoubben 1ee62392b9 Playbook 06 — writing a module, from doing it once
Written after porting the first real workload end to end. Every step
exists because skipping it cost something, and the ratio is recorded
because it is the lesson: six attempts, one real bug, and the mesh was
right every time.

The rule worth carrying out of it: read the host's log before
theorising. A declaration that was sent and not applied says so there
and nowhere else — it took an hour to look, and the answer was one line.
2026-09-01 01:46:46 +02:00
jschoubben ecfc2c215e Point the plan at the survey, and name the likelier failure
The conversion's detail is operational and names machines, so it lives
in the mesh's knowledge base rather than in this repository:
`migration/where-service-data-lives` for where every service's data
actually sits, and `troubleshooting/db-password-frozen-at-first-init`
for the lockout. This document says the rule; those say the specifics.

The lockout is the finding worth carrying here, because it is worse than
the one this plan was already guarding against and it is likelier. A
database image consumes its password variable only when its data
directory is empty. Everything keeps data on a persistent directory, so
the role holds whatever password it was created with for ever;
regenerate the variable and the application moves on while the database
does not, permanently, because nothing reconciles it.

Eight modules are in that state today and work only because nobody has
regenerated their credential since their data directory was created.
It was already documented in the knowledge base and my survey had missed
it — found by searching, which is the argument for the knowledge base
existing.
2026-08-31 21:46:53 +02:00
jschoubben 122405df8e 020: the server version is not it either
Pinned 2.5.0 rather than latest, on the suspicion that its draft
profiles extension was involved. Identical failure, so that is ruled out
and recorded — two of the three guesses in this issue have now been
tested and both were wrong, which is the useful half.

The scenario keeps the pin regardless; it should have had one from the
start.
2026-08-31 21:42:13 +02:00
jschoubben acd5a0d80b File 020 — a certificate is issued and never collected; close Phase 1
Against a real ACME server the proxy orders, the challenge is answered
at the name on port 80 through the proxy itself, the authorisation goes
valid, finalisation is accepted, and the authority issues a certificate.
The client then posts to an empty URL to collect it, and never does.

Read from the authority's own log rather than inferred. Across one run
it issued two certificates and accepted finalise three times: the client
reaches issuance every attempt and fails at the same step after it.

Ruled out and recorded, so nobody repeats it: the directory is complete;
the authority's API certificate covers the address; the challenge path
works. A hand-written server config was suspected and was wrong —
replacing it with the server's own default, changing only the challenge
port, gives the identical error.

Filed rather than pursued because what remains is interop between two
libraries against a server that exists to be a test server, and may say
nothing about a real authority. What the mesh needed to show, it showed:
a routed name gets a certificate ordered from a configured authority,
and an unrouted one gets nothing — that second assertion passes.

Phase 1 closes with this one item partly open. Two of its four tasks
needed no code at all, the network shape was built, and the next thing
to learn comes from moving a module rather than a fourth lab run.
2026-08-31 21:40:03 +02:00
jschoubben f4347e2f14 Correct an overstatement: a sealed secret is readable on its node
An earlier paragraph implied a secret becomes unrecoverable once
accepted. It does not. It is sealed to the node, which holds the private
half and writes the plaintext into the module's own file at 0600 — the
value is there, on the machine, as an ordinary file.

What does not exist is a way to ask the mesh what a secret is. That is
the property worth having and it is narrower than what was written.

The reason to capture the old system's environment first is simply that
adoption means supplying those values, not that they become
unrecoverable.
2026-08-31 21:34:50 +02:00
jschoubben b68103d198 A module is adopted with the credentials it already has
Nothing is rotated during the conversion. A service keeps the password
it is already using, because minting a new one is how a running service
stops being able to reach its own database mid-migration.

The mesh has both paths already: generate-and-seal for a new module,
accept-and-seal for an adopted one. Adoption needs the second, and it is
built.

Rotation becomes a separate act afterwards, once everything works — the
machinery is proven, and it is a thing to do deliberately rather than as
a side effect of moving a service between systems.

Records the step that has to come first and is easy to miss: read the
current environment out of the old system while it can still be read.
Once accepted, the mesh cannot show a secret back, and once the old
system is gone neither can that. A password nobody wrote down is a
service nobody can adopt.
2026-08-31 21:31:26 +02:00
jschoubben f801b4b3e2 The board is published publicly, which decides everything else about it
Assumed throughout and stated nowhere — the wrong way round for the most
consequential fact about this component.

A board reachable only over the private network would sit inside the
boundary 0004 already calls the security boundary, and a login there
would guard a room whose door is inside the building. This one faces the
internet, so its login is a perimeter rather than defence in depth.

Which makes the identity provider the mesh's outermost gate. The board
presents the control plane, and the control plane's networked surfaces
can change the mesh (0035) — so whoever that provider admits can assign
modules, from anywhere. Written flatly because it is easy to arrive at
one reasonable step at a time and then be surprised by.

What follows is not the board's own design: who may log in is a decision
about the mesh rather than about an application; a public name needs a
certificate from an authority the world trusts, which is why that work
exists; and the provider going wrong in the permissive direction is a
mesh-wide exposure with no local symptom.

The command line is unaffected and is why this is tolerable — it
authenticates through nothing and answers to the machine's own login, so
the mesh stays operable by somebody standing at it whatever happens to
the gate. That is the property to protect if the rest is ever traded
away.
2026-08-31 21:26:06 +02:00
jschoubben a1a10e9ed2 Bootstrap ends at a usable mesh, and the first credential comes from a person
Bootstrap stopped when the control plane started — a mesh that runs and
cannot be used by anybody not standing at the machine, since the
networked surfaces need an identity provider and no module has been
assigned yet. It now runs through the provider and the first login.

The obstacle was not incidental. The mesh has never held a readable
secret: Make generates and seals, keeping no readable copy. An initial
administrator's credential is the first value a person must read.

Generating it and printing it once was the convenient option and is
refused. It would give the control plane a plaintext secret for the
first time — briefly, and to one terminal, but the capability would then
exist, and an exception made for one case does not stay one. The next
awkward credential gets printed too, and "a copy of the database is a
copy of nothing" stops being checkable by reading the code.

So the operator supplies it, on standard input, not echoed — the path
that already exists for a model-access key. What is created is an
account in the identity provider, not a user of the mesh; there is still
no user model.

Unattended bootstrap remains possible and the value still comes from
outside: automation supplying it is the operator supplying it. What is
refused is the mesh inventing one, so an unattended bootstrap with
nothing provided yields a mesh with no administrator — correct rather
than broken.
2026-08-31 21:23:55 +02:00
jschoubben 0ef6d3f574 One implementation, several surfaces, and what that costs
The mesh is operated from a command line and must be operable from a
browser and from a model's tools, without becoming three systems. The
pattern is already in the code and was unnamed: `board` serves HTTP by
calling the same functions the CLI calls, holding nothing.

Takes the decision 0034 said had to be taken deliberately rather than
arrive with a feature: the HTTP surface is not read-only, so a browser
login now carries authority over the mesh.

Names the dependency by protocol — an OAuth2 identity provider — as the
mesh does for AMQP, S3 and OCI. Keycloak is what fills the role; what
the control plane knows is that it validates a token, and replacing the
provider is a migration rather than a redesign.

Says what this must not become, because it is the failure the project
was started over: a kernel every module imports, 155 files of code from
every context. Shared surfaces are not a shared library. Three adapters
calling the same functions is not the same as logic leaving the context
that owns it.

And records the loop it creates. The networked surfaces depend on a
module the control plane assigns, so when identity is down nobody can
authenticate — including whoever is trying to fix it. The way out is the
command line, which authenticates through nothing and is available to
the account that owns the machine. Hence the rule: no capability exists
only behind an authenticated surface, because that is a capability which
disappears exactly when identity does.
2026-08-31 21:17:17 +02:00
jschoubben fccac61e58 The board is a web application, not a category
Supersedes 0032, which decided the right thing and described it wrongly.
The decision is unchanged: the account that installed the host owns the
mesh, and there is no user model.

What was wrong was inventing "a surface that delegates authentication"
for the board. It is a web application with a login, in the way every
web application has a login. That is a fact about an application, not a
property of the mesh.

The cost was not cosmetic. It made the identity module look like part of
the mesh's authority — something the mesh depends on to know who anybody
is — when the mesh knows nothing about people at all and one of the
applications running on it happens to have a login.

Keeps the line that is worth writing down, and states it more plainly:
signing in to an application must not become authority over the mesh.
Today it cannot, because the board reads and does not act. The moment it
can assign a module, whoever it lets in has mesh authority — and it
would arrive as a feature rather than as a decision. So a surface that
can change the mesh is a change to who owns the mesh, and is taken as
one. Not forbidden; just not something that turns up in a pull request
titled "add assign button".
2026-08-31 21:08:39 +02:00
jschoubben 10fa7c76d7 The substrate is a store and a broker
Third correction to one table today, found the same way as the other
two: by asking whether both halves of the test were answered, or only
the easy one.

0006 admits the registry because "it cannot grant itself a repository" —
true, and the second half. Nothing established that the control plane
needs one in order to run. Counted rather than argued: the bundle raises
twelve resources and no registry is among them. The registry arrives
afterwards as an ordinary module, which is exactly what the lab asserts.

0006 half-said this already, calling it "substrate by role and ordinary
by delivery, provisioned once there is a control plane to do it". A
member provisioned by the thing it supposedly precedes is not a member;
that phrase was carrying a contradiction rather than resolving one.

The registry is a closer call than the object store and the difference
is worth keeping: the control plane never touches an object store at
all, but it genuinely uses the registry. So the registry is a real
dependency of the mesh operating and not of the control plane starting —
and it is the second that the word means.

The substrate is now exactly what the bundle raises, which is the
strongest form the list can take: checkable by counting rather than by
reading an argument, and the two cannot drift.

The finding is not about substrates. A test with two conditions is a
test only when both are asked.
2026-08-31 20:30:19 +02:00
jschoubben a028337490 The local account owns the mesh; a surface delegates to a module
Answers what 0031 left open, and a question it did not ask — who owns
the mesh at all. There was no answer, and the absence was invisible
because every operation so far has been run by the person sitting at the
machine, so nothing had to say whether that was the design or the
circumstance.

The account that installed the host owns the mesh on that node. No user
model, no roles, nothing to administer. It follows from 0004 rather than
adding to it: there is no authorisation between nodes because every node
is the operator's own, so a user model inside that boundary would guard
nothing — anyone it could stop could read the node's key off the disk.

The board is different, and the difference is the network. A surface
reachable by a browser has to know who is asking, because those people
are not by construction people with a shell on the machine. So it
delegates to an OAuth provider, which is a module.

That does not make identity substrate. A surface delegating
authentication is not the control plane delegating it: the control plane
runs, applies declarations and reaches nodes with no identity provider
in existence. Only the board needs one.

Records the cost plainly: anybody with a shell on a node has full
authority there, and there is no way to give somebody authority over one
node without giving them a login on it.
2026-08-31 20:23:42 +02:00
jschoubben e3934e4449 The control plane authenticates nobody, so identity is a module
Closes the last open question about what the substrate contains. 0006
left an identity provider conditional — substrate only if the control
plane delegated authentication — and said the decision had not been
taken. It is now: it delegates to nothing.

The conditional was never about machines. A node proves itself with a
keypair it generated over a broker account issued at enrolment, and
declarations are verified by signature; none of that involves an
identity provider. It was only ever about whether a person signing in to
a mesh surface would be authenticated by something else.

So the substrate is three — a relational store, a message bus, an image
registry — and with 0028 having removed the object store, no member is
conditional and every one is there for the same reason.

It does not settle how a person signs in to a surface, deliberately.
What is settled is that whatever answers that is not something which
must exist before the mesh does, so it can be decided late or replaced —
which being substrate would have prevented.
2026-08-31 20:18:03 +02:00
jschoubben c570c687f6 The handover: switch off the old brain, leave the services running
The conversion method, recorded because it decides everything else and
was not written down.

The old control plane is stopped — provisioning, coordinator, syncs, the
pipeline, anything that decides or writes. The workloads it was managing
keep running, because nothing is managing them. The new mesh then takes
ownership one module at a time.

Nothing is ever unassigned in the old system. Unassigning is how it
removes things and removing is how data is lost; it is asked to stop
having opinions, never to take anything away.

Disabled rather than merely stopped, which is the part easy to get
wrong: those units are enabled, so a stop lasts until the next reboot. A
reboot mid-conversion would bring the old control plane back to
regenerate managed files underneath the new one — the one situation
where two systems really would fight over a machine.

A service left running with nothing managing it is the safe state: it
has its data, its configuration is on disk, and nothing will change
either. The risk in a conversion is in the managing, not the running.

Also records why taking ownership piecemeal is safe: the new host's
orphan removal is per-origin, so it only removes what it recorded
itself. Services it was never told about are not orphans to it.
2026-08-31 19:55:35 +02:00
jschoubben d1ab2dc0b4 Data outlives the mesh that declared it, and the conversion starts where it lives
0030, found by asking what the conversion actually needs rather than by
reviewing anything. The host deleted a directory and everything under it
when it stopped being declared — which happens when a module is
unassigned, or when a manifest is edited to move a data folder, which is
the exact operation this plan needs. A database's files, a mail spool.
The report said "removed".

A directory still holding something is now kept and said so. No flag and
nothing to remember: emptiness is the test, and it works because the
removal order was already right — the mesh's own contents are gone by
the time the directory is reached, so what remains is by definition
something nobody declared.

The plan now says data outranks its own ordering: copy, read back
through the service that owns it, and only then point anything at the
new location. Never move and then check.

And it records where this starts — the node holding all the production
data — with what that costs stated rather than argued with. Everything
proven so far was proven on machines that could be destroyed and raised
again. A scenario proves the mechanism, not the state on that machine.
2026-08-31 19:54:41 +02:00
jschoubben 9ad0ec35e0 The conversion is done by hand, and that removes work from this plan
Recorded because it is load-bearing and was not written down: moving
from the current system to this one is a person at a command line, not a
migration program.

What that removes is larger than what it adds. Nothing in this plan
needs an importer, a translation layer, a compatibility shim, or a way
of keeping two systems agreeing while both are live — each of which
somebody would otherwise reasonably build, use once, and maintain for a
year.

It also settles what "safe" means for the system being retired: a fix to
it must be safe on its own, because there is no careful rollout to
sequence it into. A change needing three steps in the right order is a
change that will be half-applied. That reversed a certificate default I
had chosen this morning.
2026-08-31 19:46:21 +02:00
jschoubben e4327a3a5e Phase 1.3 done: ordering was already there, the network was not
Ordering needed no change for the third time running — resources apply
in the order declared and nothing sorts them — and is now asserted,
because sorting them for any sensible reason would have passed every
other test.

Separates ordering from readiness, which the task had run together: a
container started is not a container ready. Nothing waits, and what
needs something usable retries. That is deliberate and more robust than
start ordering, since a dependency can restart long after apply.

The network was the first thing in Phase 1 that genuinely needed
building, and the first that needed a decision: 0029 records why a shape
rather than an action, and the vocabulary is nine.
2026-08-31 18:55:22 +02:00
jschoubben ce486fd5d2 Phase 1.2 done, and it is the same surprise as 1.1
A session as a licence consumer needed no change either: the two
sessions are two modules, so the existing (node, module) binding already
names them apart. 14-model-access.md's "a step toward it and not it" is
true of a worker and not of a session, and the difference is that there
is one session per node rather than many per machine.

Records what stays open: the worker half of that gap is real and
unaffected, and belongs with 0003, which is unbuilt.

Two tasks in a row that were already possible. Both were written from
the design rather than from the code — the review's own finding arriving
in the plan it produced. The remaining Phase 1 items should be checked
against the code before being started rather than after.
2026-08-31 18:41:19 +02:00
jschoubben fdd909ec40 Phase 1.1 done, and it was not the task that was written down
An object-store provision, proven against a real store with seven
assertions.

The finding is worth more than the task: the control plane
special-cases nothing. provides, requires, contributes and grants are
name-agnostic, so asking for a bucket needed no change to the mesh at
all. What was missing was a provider and the last step on the machine —
"add an object-store provision" was never mesh work, and the breakdown
now says so rather than leaving the next person to rediscover it.

Named s3-bucket by 0027: the coupling is to the API, not the product,
because swapping one store for another does not break a consumer. A
database is the other case and names its engine.

Records the assertion a database does not need, because it is the one
that will be forgotten when somebody writes the next provider: one store
holds every bucket behind one endpoint, so isolation is a policy rather
than a property, and a policy granting everything passes every test that
only checks a consumer can reach its own bucket.
2026-08-31 17:53:02 +02:00
jschoubben cb1954e7b5 Accept 0024, and rewrite the work breakdown around what is actually being done
**0024 accepted.** Model access was decided, built, and proven in the
lab, and two design documents rest on it; only the status had never
moved. The gate is green again.

**The work breakdown rewritten.** It planned a decomposition of the
existing system in place — extract contexts, declared features, shrink
the shared library. That is not the work. A replacement is being built
beside it, and only the old Phase 0 survived contact with reality, so
the one document meant to say what happens next was describing a system
being retired.

Now ordered by what "modules move across one at a time until the old
registry is off" actually requires:

- Phase 0 is marked done against the twenty-two lab assertions, **and
  carries its own limitation**: every module exercised was written to
  test the mechanism, so the vocabulary was shaped by its own fixtures.
- Phase 1 is the vocabulary gaps found by asking what real modules
  need — an object-store provision, a session as a licence consumer, a
  network shape with ordering, public certificate issuance.
- Phase 2 is one module, then a week of running it, because the point of
  going first is to find what Phase 1 missed.
- Phase 3 picks modules that each prove something the first did not; the
  mail system is last because it is the one that may send work back into
  the declaration language.
- Phase 4 is switching the registry off, named as a phase so it is not
  mistaken for the goal.

Keeps the rules of engagement unchanged — they were about how work is
done, not what it is — with one addition: stop and ask before anything
that touches a machine outside the lab.

Adds a section on keeping the list true, since the document it replaces
was wrong for weeks and nothing said so. A claim here is counted, not
reasoned, and a phase is done when the lab says so.
2026-08-31 17:25:27 +02:00
jschoubben 1b5308c9cc Review of the to-be layer: check what the documents claim against what runs
First pass of a design review, done by reading documents against code
and against a raised mesh rather than against each other. Every error
below was invisible to a proofread.

**Statuses were stale, and nothing checked them.** Ten to-be documents
said `designed` while naming working, lab-proven code — several with a
*What was built* or *Raised, and observed* section. Added a
`status-vs-code` check: naming a file is a claim that the file
implements this, so a document that points at one has stopped being
merely designed. It failed on all ten before it passed, per the rule
this folder sets for its own checks.

**The bundle carries three images, not two.** 07 reasoned about which
substrate services go in and overlooked that the control plane is in
there too — it is what the substrate exists to start, and there is
nothing to fetch it with yet. Counted, not deduced.

**The bootstrap uses four shapes, not six.** It listed `file` and
`directory`, which substrate-first-node.lock never asks for. The claim
that mattered — nothing is blocked on the host — was true either way,
which is why the wrong count survived.

**The eight capabilities were documented nowhere.** Implemented in
internal/profile/detectors.go and enumerated in no document, including
the one about the host that detects them. A vocabulary modules write
against, readable only by reading the code. Now written down, with the
seat/graphical-session distinction that is wrong in both directions if
collapsed.

**MinIO swept out of the to-be layer** per 0028.

The gate now fails on one thing left deliberately: ADR 0024 is
`proposed` while two documents rest on it and the feature it decides is
built and lab-proven. Accepting a decision is not mine to do.
2026-08-31 17:20:47 +02:00
jschoubben cbcbba8099 A provision names its engine; the substrate supplies only the control plane
**0027 — provisions.** A module written against PostgreSQL could be
matched to a provider of SQL Server, resolve as satisfied, and fail on
its first query. The name said the role, so nothing distinguished
engines. Refusing on ambiguity could not help: with one provider of
each name nothing is ambiguous. Enforced at parse rather than
documented, because the old naming was the documentation.

**0028 — the substrate.** 0006 admits an object store on the grounds
that it cannot grant itself a bucket. That answers the second half of
the test and assumes the first: the control plane does not need one.
Verified — no S3 client in mesh-control, and internal/builder/registry.go
records the deliberate choice to put artifacts in the OCI registry as
content-addressed blobs. The row was inherited from the system being
replaced, where an object store distributed module tarballs, and was
never re-tested against the definition above it.

So an object store is an ordinary module, and a mesh with nothing
needing one runs none. Migrating it is module work, not substrate work.

0028 also states what 0006 left unsaid: a substrate service and a
module of the same product are different instances. The substrate is
raised from the bundle before any mesh exists, so it is not in the
module graph — a workload depending on it would depend on something the
graph cannot see, cannot rotate a credential for, and cannot move, and
would put workload data in the store the control plane keeps its own
state in.

Both records were found by reading code against design rather than
design against itself, which is the review that should have happened
sooner.
2026-08-31 17:13:07 +02:00
jschoubben 046a990198 Several sessions at once is a surface property, not a session one
Answers the question 15 raised: a board showing many sessions leaves
one-per-node untouched, because each is still one conversation. Only
concurrent conversations with the same session would touch 0004.

Records soulstream and herdr as the prior art to draw from, and marks
it explicitly off the provisioning path so it stays a note rather than
becoming the work.
2026-08-31 16:52:15 +02:00
jschoubben fbf5b1d04b A session's memory is its own, and it is not declared
Settles the question 15 left open: the mesh session holds its own
memory in the mesh root, rather than assembling a view over the node
sessions. Memory follows the rule the rest of the design already uses —
the context root is the whole of what makes one session a different
agent, and memory is part of what makes it that agent.

The control-plane node is what makes this load-bearing rather than
tidy. Two sessions share that machine; if memory belonged to the
machine instead of the root they would share it too, and the mesh's
recollection would be indistinguishable from that node's own — the
collision 0026 exists to avoid, arriving through the back door.

Also corrects an error made writing it up: memory is NOT declared
state. The engram and tools are — the mesh says what they are and the
host writes them (0011). Memory is written by the session itself and
declared by nobody, so a mechanism that regenerates the root wholesale
would erase it on the next heartbeat, silently, while reporting
success. The root is not uniformly managed and which parts are has to
be explicit.
2026-08-31 16:41:06 +02:00
jschoubben 3c6c16abdf The mesh has a session of its own, and it is the node session's mechanism
A session for the mesh itself, addressed as the mesh, differing from a
node's in exactly three things: the context it starts in, its engram,
and its licence binding. Not a new kind of agent — the same mechanism
pointed at a different root. Two implementations of one mechanism drift,
and the vocabulary collision 0001 exists to undo began exactly that way.

It runs on the control-plane node, and the reasoning is easy to get
backwards: not "the important agent on the important machine", but that
this node is already the one place excepted from "compromise of a node
is compromise of that node". Placed anywhere else it would create a
second such place.

It is an addition to per-node messaging and never a replacement. 0001
holds that losing the control plane costs change, not operation — and a
mesh whose only conversational surface lived there would lose the
ability to ask anything while every machine kept running perfectly.

Writing it up exposed that the node session's setup was never designed
at all. 0004 gives behaviour and stops: nothing said how a session
starts, where its context lives, or how a broker message becomes a
prompt. That gap was invisible until something had to be built *like* a
node session. 15-the-agent-session.md covers both as one mechanism.

It also makes "a consumer that is not a machine" undeferrable. The
control-plane node now hosts two sessions that must hold different
licences, and a per-machine binding cannot express that at all. Noted in
14-model-access.md against the gap it was already recorded as.

Also completes the to-be index, which stopped at 10 and omitted four
documents. Pre-existing broken ADR references in the older rows are left
alone rather than guessed at.
2026-08-31 16:32:59 +02:00
jschoubben e823cc1cc5 The design record is read where it is written, never copied to be found
Decides the question 006 narrowed to. An agent reads this repository
directly and the search consults it, so these documents surface beside
ordinary results instead of only when somebody already suspects they
exist.

A scheduled sync into the mesh's memory was the option that works with
what exists today, and lost on the ground this repository can least
afford: it makes a second copy, and the copy that is searched quietly
stops matching the copy that is edited. A design record that has
silently diverged from the reasoning it claims to carry is worse than
one that cannot be found — the first misleads, the second merely fails.

Amends what 0019 promised rather than satisfying it: these documents
will not be indexed, they will be read. The commitment that survives is
the one that mattered — that a searcher finds them without already
suspecting they exist.

Gated on an agent that does not exist yet, so 006 stays open on the
build with a decided shape. What closes it is a check that fails today
by design: search the mesh's memory for a phrase that appears only in a
design document here, and require it back.
2026-08-31 15:45:00 +02:00
jschoubben c192810fba 004 and 008 resolved; 006 narrowed to the decision it actually needs
**004 — certificate issuance.** The resolver declared no authority at
all, so the client fell to its built-in production default: there was no
setting set wrongly, there was no setting. It is now a node property
defaulting to staging, which answers the first open question. Staging by
default rather than production-with-an-override, because the alternative
leaves the safe path depending on remembering to opt out of it — 005's
lesson, in a second place. The rollout is ordered and the order is the
dangerous part; recorded, not performed.

**008 — node rescue.** Read back from running nodes as the report asked,
and one of its own claims was wrong in a way that matters: the health
timer does exist and does fire. It simply never calls the rescue script.
A trigger that exists and does not do what the script claims survives a
halfway check, which makes it worse than the absence the report
described. Resolved by making the documentation true, not by
implementing rescue — the replacement host already supervises recovery,
and wiring unattended restart into the fleet being retired is a
deliberate decision rather than a tidy-up. Two "self-healing" claims
narrowed to what they actually do.

**006 — deliberately not closed.** Re-checked today: the indexing still
does not exist. What is gone is the reason it was an issue — the claim
is no longer load-bearing, because the README names the gap and the
decision's reasoning never invoked indexing. A signpost now points here
from the knowledge base, and was measured rather than assumed: it is
reachable, it is not surfacing. Closing it while the indexing does not
exist would be this repository's own named failure, one folder from
where it names it.
2026-08-31 15:20:11 +02:00
jschoubben 345bbe0552 005 resolved: a suite that cannot run on every push says when it last ran
Retired in favour of the lab rather than repaired — that answers the
first open question. The second finding is the one that generalises:
"nothing runs it, and nothing reports that nothing runs it" is not a
fact about that harness, it is a fact about any suite too expensive to
run on every push. The replacement inherited the fault it was replacing.

Records the three rules that now hold, and what the fix taught twice:
the remedy rebuilt the symptom inside itself, and the code that counts
results passed every test while reading nothing.
2026-08-31 15:02:26 +02:00
jschoubben f3ffdae909 Record the resolver as built, and the two things it must not do
A service is reached at <service>.<node>.internal, so what resolves is anything
under a node's name. The mesh writes the data and runs no daemon; two roles,
two claims, because systemd-resolved cannot serve a wildcard at all.

Both prohibitions were found by a machine rather than by reasoning: an address
systemd already held, and reading resolv.conf for upstreams that now point at
itself.
2026-08-31 14:23:02 +02:00
jschoubben 6ecd03694b Issue 019: a comment asserting a fact about a machine, which nothing checked
Twice in one file, a statement about a machine that read as reasoned and was
wrong — and the module's unit tests all passed while the daemon could not
start. That is what a unit test is: it confirms the assertion was made, never
that it is true of any machine.

003 in prose rather than in a manifest key.
2026-08-31 14:11:34 +02:00
jschoubben a036bac47b Record why the module's own secret is named for whose it is
The field was called needs, beside secrets, and both were name-to-path holding
something secret. What separates them is whose, not how secret — so that is
what the name says now.
2026-08-31 13:47:21 +02:00
jschoubben aafeb5c9df Three issues resolved: one closed by evidence, two answered by the replacement
012 named its own closing condition — a scenario with four images coming up —
and the scenario now stocks seven and has raised cleanly many times at the
memory the wrong diagnosis had raised.

001 is answered by the host reading the package database back after installing.
002 was NOT answered and was present here too, so it is a fix rather than a
note: a stale index is now named instead of reported as a failed install.
2026-08-31 13:00:03 +02:00
jschoubben 573a94e102 Correct the record: a limitation that no longer exists, and one that was never written
The connectivity design still said a hub cannot be filtered — a gap recorded in
the morning and closed in the afternoon, left standing as though it were
current. Worse than a stale date: it would send somebody away from something
that works.

`restart-on` was described nowhere, including the part added today that lets a
service reflect a file another module put on the machine. A rule the host
enforces and no document mentions is a rule nobody can rely on.

And nine of fifteen design documents claimed an `updated:` older than their last
change, some by a week. That field is what cross-cutting views are generated
from, so it is not decoration.
2026-08-31 12:33:35 +02:00
jschoubben f4e81074d3 Which resolver is a claim, and was decided before it was asked
A resolver takes over /etc/resolv.conf, which is a singular resource — ADR 0009
lists it in the table beside the seat and pid 1. So choosing between resolved,
dnsmasq and unbound is assigning a module, per machine, and the mesh refuses
two rather than letting them fight over the file.

Recorded because it was treated as an open question two days after being
decided, which is the argument for that table being a table.
2026-08-31 12:15:12 +02:00
jschoubben a3cee17d48 Record what a container can see of the mesh's names, and what it cannot
Found by a container failing to resolve a name every machine could: a container
gets its own hosts file holding only its own hostname, and on the machine it
always worked, which is what made it easy to miss.

Declared containers are given the names. A container somebody starts by hand is
not the mesh's to configure — which is a second, different reason to want a
resolver, recorded beside the first rather than folded into it.
2026-08-31 11:29:25 +02:00
jschoubben 35ce23fb76 Record that a machine coming back is ordinary, and what waking now does
Asked whether a machine that drops off needs re-adopting: it does not, nothing
expires, and the only thing that forces re-enrolment is losing its own key.

The gap was the twenty or thirty seconds after a resume in which a node
believes it is in a mesh it has left — recovering on its own, which made it a
quality gap rather than a fault, and still a machine waiting to be told
something it already knew.
2026-08-31 10:21:50 +02:00
jschoubben 6374c1eb60 Record what "behind" means now
It meant failed-or-refused, so the question this record says must not be lost
was answerable only for the machines that broke. Out of date, never told, and
not worked out are kept apart: the remedy is the same push and they read
differently to whoever is looking.
2026-08-31 05:28:11 +02:00
jschoubben 7c0be6968e Record what a machine says about itself and what the mesh keeps
The yes gates an assignment and the detail carries a value; they are one fact
read two ways, and only one read was being kept.
2026-08-31 04:54:07 +02:00
jschoubben 7f314eb399 Point three design documents at the code that exists for them
Their subject matter has been built and proven for days and their frontmatter
still said code: [] — which is what the cross-cutting view is generated from,
so it was claiming nothing existed for the substrate, the node lifecycle and
delivery.
2026-08-31 04:51:04 +02:00
jschoubben 5ad3641cbf Record the board, which is the last designed document with no code
One reading answered three ways, holding nothing and touching no context's
store — which is the constraint the whole document is about, and the thing the
board being replaced gets wrong.
2026-08-31 04:50:45 +02:00
jschoubben 0bc4b7774f Record what four more pieces of the mesh became
Rotation and the provisioner contract; model access as a provision answered by
a record, with ADR 0024's other two gaps left as gaps; exposure, which closes
the open question about revoking a route; and the delivery loop, which closes
the gap ADR 0010 left when it replaced a pipeline with a comparison.
2026-08-31 02:56:50 +02:00
jschoubben 6b1c80b442 What a builder-as-a-module can and cannot reach
The broker's fingerprint travels with its credential, and the machine's
filesystem does not travel at all — it runs in a container, which is the
arrangement working rather than a limitation to route around.
2026-08-31 01:53:35 +02:00
jschoubben 8448219de1 Issue 018: a provider on the same machine was never announced to its consumer 2026-08-31 01:35:54 +02:00
jschoubben 00e98f1f92 The vocabulary has no word for a unit that runs and exits
Found by the firewall: every packet filtered as declared, and the machine
reported as not doing what it was told, because the unit that loaded the rules
had finished. Stated as a gap rather than worked around silently.
2026-08-31 01:20:58 +02:00
jschoubben 0ac99f0d68 An action's own idea of being finished must be its verify's
Otherwise it succeeds into a state its verify rejects, and the host's report is
accurate and names nothing. Recorded where the vocabulary is described, because
it is a rule about writing an action rather than about one action.
2026-08-31 01:03:50 +02:00
jschoubben 1ade18209d Issue 017: an action succeeded into a state its own verify rejects 2026-08-31 01:01:17 +02:00
jschoubben a9cd3de5be Issue 016: anything after the declaration in a file was ignored 2026-08-31 00:51:46 +02:00
jschoubben 6e7e77acbe Issue 015: the harness read a swallowed answer as success 2026-08-31 00:48:16 +02:00
jschoubben 890c3ee3fc State the one rule the derivation does not yet reach
A hub needs its overlay port open and a node that is not a hub does not, and
they are the same module — so listens, a static manifest field, cannot express
it while the overlay module's resources are computed per node. Written down
rather than left as an oversight for whoever first puts a firewall on a hub.
2026-08-31 00:42:39 +02:00
jschoubben a6872ac099 A key that is present and unusable, and what the certificate work became
Issue 014: the node's serving key was stored in the host's own encoding, so
every check that reads the file passed and no server could start. Same shape as
013 — two halves of one mechanism designed separately, each correct about its
own half. Where a file exists so a third party can read it, the format is the
interface.
2026-08-31 00:42:13 +02:00
jschoubben 778efaba8b The mesh runs its own registry, certifies its own names, and computes its own filtering
Issue 003 is answered in both halves: manifests are parsed strictly, and a
module says what it listens on and from where rather than carrying a key
nothing reads. The design records what was built and how each part is checked.

Issue 013 is new, found by reading while writing the first module that has
both a computed file and a service that needs it. The file arrived second.
It failed, then the next reconcile fixed it, which is why nothing caught it.
2026-08-31 00:37:34 +02:00
jschoubben ba14e2b629 Issue 012 — the first diagnosis was wrong, and that is the useful half
Two things changed at once: a fourth image in the scenario, and scenario
machines raised from 1 GiB to 2 GiB. The bootstrap then failed every
time, and the image was blamed.

Removing the image did not fix it. Removing the memory increase did —
nine assertions pass again on three images with the machines back at
1 GiB. Three machines at 2 GiB on a host doing other work contend enough
that the store container does not come up at all.

The ordinary lesson, and it still caught me: two changes together, the
failure attributed to the plausible one, and an issue written recording
the wrong cause. What found it was reverting to the exact last-known-good
state rather than reverting the suspicious change.

What remains untested is whether a fourth image alone is fine. Probably.
Nothing has measured it, and the honest state of this issue is that what
it was opened about was never demonstrated.
2026-08-30 20:40:54 +02:00
jschoubben de2b6a2825 Issue 012 — a scenario machine cannot hold four images and raise a
substrate

Adding a fourth image to the two-machine scenario makes the bootstrap
fail every time, with the store's readiness check producing no output at
all — which says the container was not running rather than that the
database was slow. Three images pass nine assertions; four never get past
the store.

More memory did not change it, so memory is not the cause; the change is
kept because the reasoning holds on its own. Disk is the most likely
explanation and nothing has measured it.

It blocks proving the mesh runs its own artifact store, since the
registry module needs a registry image to mirror. The module is written
and accepted; what is unproven is a machine assigned it serving another.
2026-08-30 20:29:55 +02:00
jschoubben 1b49e684e1 Issue 011 — an action is a gate, which the first fix got wrong
The first fix continued past every failure, and the next lab run failed
at the bootstrap: the store did not answer in three minutes and then said
"the database system is shutting down". Carrying on past the readiness
gate had started the broker and the control plane against a machine that
was not ready, and on a small machine that is how a database still
initialising has its memory taken away.

An action is the only shape whose purpose is to make something true
BEFORE the next thing needs it, which is why it is the only one with a
verify. So a failed action stops what follows and nothing else does —
which fixes both this and the hostage problem the issue was opened for.
2026-08-30 20:07:41 +02:00
jschoubben 4a601ceb02 Issue 011 — one broken module stops every module after it
Found in the lab. A machine with one impossible module applied nothing at
all on every later push, and the mesh said "failed" without saying the
rest was never attempted.

Recorded with the evidence, including that the behaviour's test cited a
record which does not decide it: ADR 0010 argues about pipelines against
reconcilers and says nothing about whether one resource failing should
stop the next being attempted.

Fixed in mesh-host: everything is attempted, every failure reported.
2026-08-30 19:26:03 +02:00
jschoubben c3ec1487c7 What a module may borrow, what it may need, and what one assignment gets
Three additions, all written by trying to write a real database module
and finding out what could not be said.

A module may mirror an image it did not write. Naming an upstream
reference directly needs every machine to reach a public registry and
pins to a tag somebody else can move.

A module may need a secret of its own — a superuser password is not FOR
anybody, so the mechanism that hands credentials to consumers cannot
express it. Per node, so three machines have three passwords.

And the provisioner watches, which is what lets it be a module rather
than a binary somebody places. It polls rather than watching the
filesystem, because the host writes atomically and a watch on a replaced
path silently stops working.

One assignment now gets a working database provider: two directories, two
pinned containers, a sealed password and the grants manifest.
2026-08-30 18:28:30 +02:00
jschoubben 34551a3b8a The three questions a board answers, in the order they are asked
A page nobody had thought to ask for turns out to be the one a person
opens first: what is not doing what it was told. Recorded with the order
that matters — broken, then quiet, then out of date — because a page
leading with the last would bury the first.

And refused stays distinct from failed all the way to the page. They are
fixed in different places, so one word for both sends half its readers to
the wrong one.
2026-08-30 18:09:17 +02:00
jschoubben 41a4405253 What a build machine may do, and what is kept
Two additions to the module-repository design, both from building it.

A build machine has its own credential and it is not a node's: read the
build queue, write the mesh exchange, nothing else. A node's queue
carries that node's declarations.

The answer goes through the exchange and never the default one, because
permission there is per exchange rather than per queue — anything allowed
to use it can publish into any node's queue. The price is that every
asker sees every result and filters by correlation, which is cheap
against a builder never needing that permission.

And every result is kept, failures included, because one that leaves no
trace is indistinguishable from a build nobody asked for. That is what a
builds view reads; the board page is corrected to say so.
2026-08-30 10:29:01 +02:00
jschoubben 4151c28927 The builder takes work over the broker, and what is still missing
A build is work, not state, and that is why it does not travel as a
declaration: as one it would either rebuild on every reconcile or carry
"and I already did this", which is state about an event rather than about
a machine. So it has its own queue and the answer comes back correlated.

Three properties recorded because they are decisions: acknowledge only
once the answer is away, one build at a time, and a failure is a result
rather than silence.

And the board page is corrected. A build result today is answered to
whoever asked and kept nowhere, so a builds view has nothing to read. A
record of past builds is the missing piece, not the builder.
2026-08-30 04:06:42 +02:00
jschoubben a625e6c709 A module repository, and two more shapes the host speaks
Designed with no reference to what came before, which was asked for. The
system this replaces has features — several deployable units inside one
module — and they are deliberately absent.

That closes something ADR 0001 has been carrying as an open prerequisite.
It lists "named features with per-node opt-in" as required, or "every
independently deployable unit becomes a module again and the count
returns". The premise was right and the remedy already exists in another
form: several modules, assignment per node, and a module with
requirements and no files of its own. `networking` is exactly that. The
count does not return because what made it return — a module is
expensive, so put several things in one — is gone. A module here is a
manifest and usually nothing else.

The manifest in a repository names artifacts; the manifest the mesh holds
names digests. Two documents, because a digest is not knowable until
something is built and a repository carrying one is wrong the moment
anybody edits anything.

The builder runs on a node. Building needs a container runtime and a
working tree, and what the control plane may send a machine is bounded by
the declaration language. A control plane holding a container socket
would be the one component that can do anything anywhere.

And the host's vocabulary grew from six shapes to eight — user and
archive — with the reasoning for each and for the refusals that came with
them. The count is asserted by a test precisely because every addition
widens what a compromised control plane can express.
2026-08-30 03:36:57 +02:00
jschoubben eab870fff9 A board, and the one constraint that is not a feature
Read from the board that exists. Eight sections; four are about work and
workers and are held back with that domain. The other four are the mesh
itself, and everything behind the main one already exists here — it is a
reader, not a second source of truth.

The constraint is the point of writing this down now. The existing board
is one service that reads every context's database, because that is the
shortest path to a page showing all of them at once. That is ADR 0008
violated by the one component with a reason to violate it, and the cost
is the same one the shared library has: a boundary nothing may cross is a
boundary that can move, and one thing crossing it is enough to freeze it.

So a board reads through interfaces and stores nothing. If a question is
slow, the answer belongs in the context that owns it, where everything
else asking gets it too.
2026-08-30 03:26:12 +02:00
jschoubben 5253742773 ADR 0024 — model access is a provision, and a licence has a name
A new requirement, and it is mostly a shape the mesh already has. A
module that needs to think requires model-access; several vendors and a
locally-run model are several modules providing it; choosing is assigning
the one you want. A model the mesh runs itself needs nothing new at all —
it is a mesh-scoped provision on the node with the hardware, credential
included.

A licence is a named thing because the whole point is saying which one a
given consumer uses, and the names are the operator's. Many to many, so
not a claim: two machines sharing an account is ordinary, not a
collision.

Four gaps, written as gaps rather than design:

- a provider that is on no node, reached over the public internet, which
  the reachability rule must not refuse
- a secret the mesh is GIVEN rather than mints. Every credential it
  handles today it generated and discarded; an API key arrives from a
  person, and accepting one must still discard the plaintext
- a consumer that is not a machine. Which licence a worker uses is a
  binding to an agent, and the provisions model has no consumer identity
  other than a node
- switching on exhaustion is a reaction to something observed, not a
  declaration. It belongs with observability, changing a binding — saying
  so is what stops the declaration language growing a conditional

The existing auto-refresh and switching is not being replaced because it
was wrong. It is being rebuilt because it lives somewhere that cannot
express the rest.
2026-08-30 03:06:36 +02:00
jschoubben a73014dcd5 A bare machine became a mesh, and something joined it
First end-to-end raise. A machine with a container runtime applied the
bundle its host carries and ended with a store, databases, schemas, a
broker holding a certificate it generated itself, and the control plane
serving. Then it took a token, checked the broker against the pinned
fingerprint, generated three keypairs and enrolled — the first node being
a node whose mesh is not up yet, observed rather than argued.

And a credential crossed. Declared the provider of a database for a
second node and pushed to over the broker, the machine ended with the
password in one file at mode 0600, and that password appears nowhere in
the declaration that crossed the broker, nowhere in the control plane's
database, and nowhere in what the node reported back. That is the whole
secrets argument, measured.

One fault, in the joining: the token did not say what the mesh calls the
machine, so enrolment needed a flag its own help said it did not, and
failed at the broker with an empty username. It is the fifth thing a
token carries now — the node cannot work its own name out, because the
broker account it authenticates as is named after it and exists before
the mesh has told it anything.
2026-08-30 02:37:37 +02:00
jschoubben 0f7e4ab597 The provisioner, which is where the mesh stops
A password nothing was told to create authenticates nowhere. The mesh
generates one, seals it to both ends and cannot read it — so it cannot
tell the software to accept it either. Something on the providing machine
reads what arrived and makes it true.

That something belongs to the module, not to the mesh. The control plane
decides and never touches a machine; a provisioner runs on the machine
and touches it. What the mesh owns is the contract: a manifest of who
asked and where each credential is, and one file per consumer holding it.

It reconciles and is never told what changed, which forces three things
that are each a fault somebody has shipped: set the password every time
or a rotation changes nothing; remove what nobody asks for or a departed
consumer keeps a login for ever; leave alone what it did not make or it
cannot be run on anything that predates it.

Saying where the mesh stops is the point. It decides, delivers, and can
prove what it delivered; the last inch belongs to whoever knows what
`create role` means.
2026-08-30 01:31:58 +02:00
jschoubben 82a5b9548a The secret is delivered without ever being held
Written after looking at how the existing mesh does it, so this is a
reaction to a measurement rather than a preference.

There, credentials sit in a column encrypted at rest. Its own tooling
records what that bought: the tool for finding a secret matches by value
rather than by name, because the same password is in three tables, in
each node's environment file in plain text, and inside every connection
string composed from it — copies its documentation calls the ones usually
in use. And a query against the encrypted column returns zero rows and
proves nothing, so auditing moved to the decrypted copies.

Encryption at rest addresses neither fault. The control plane can read
what it stores, so a copy of its database is a copy of everything. And
composition is what mints the untracked copies.

So the value is sealed to the node that will use it before it is stored,
with a key that node generated. Nothing central is composed. What it
costs is auditing by value, which was never real anyway; what stays
answerable is which node holds what, which is what rotation asks.

What remains is a provisioner. The mesh generates the secret and tells
both ends; nothing yet acts on the telling.
2026-08-30 00:21:52 +02:00
jschoubben cc872a58ce Binding is built except for the secret
Which turned out to be the useful way to cut it. A provider says what a
consumer needs in order to use it; a consumer says where it wants to be
told; the mesh adds which machine and what that machine is called on the
private network. So an app on one node reaches its database on another,
by a name the mesh also created.

The file says it carries no credential and why, because a missing field
looks like a bug and a stated absence looks like a boundary.

What remains is the secret itself, and the shape it will arrive in now
exists.

Also: two machines wired together across no private network is refused,
and that only became checkable when the network stopped being something a
machine has by virtue of holding an address.
2026-08-30 00:02:36 +02:00
jschoubben 80b74d32d6 Where the answer to a requirement is allowed to live
0009 distinguishes presence from instantiation — what the edge hands
over. It never distinguished where the thing on the other end is, and
that turned out to be the half doing the damage: a shell and a database
were both written `requires`, so requiring a database installed one on
every machine that used one.

A provided name now carries a scope, as a claim already does. Scope
belongs to the name rather than to each provider, or one requirement
means two things depending on which module answers it.

A requirement answered from the mesh is never satisfied locally. Nothing
provides it, and it says which module to assign somewhere; two do, and it
says how to choose. Choosing is recorded per node, because two machines
may reasonably use two different databases.

And knowing which node answers is the first half of handing a credential
back — you cannot be given a database's password before it is settled
whose database it is.
2026-08-29 23:52:17 +02:00
jschoubben 90ecfe6a01 An edge has two directions, and only one of them is built
0009 already said a consumer supplies a target and receives a name. What
it did not say is that those are two separate mechanisms.

Contribution — publish me at this name, on this port — now exists.
Binding — and hand me back a credential — does not, and is the larger
half: a secret has to exist, be stored, reach one node and not the
others, and rotate with every holder informed. That is the invariant set
found violated three ways at once, so it is not something to add in
passing.

The absence had a measured cost. Exactly two modules opened a direct
connection to the control plane's database, and they are the reason every
node permanently holds a credential to it. Both were doing by hand what
this edge is for. Neither needed a new kind of thing.
2026-08-29 23:36:23 +02:00
jschoubben 7fe2c31bdf Networking is a module, and what a domain module actually is
Two records, from building it.

0009 has a section titled "there are no domain modules", and `networking`
now exists. It is not a contradiction and it reads as one, so the
difference is written down: what was refused contains WireGuard and a
proxy and is assigned where half of it is unwanted. What exists contains
nothing — requirements and a name — so there is no half. Every artifact
it leads to is still an ordinary module assigned on its own terms.

With the cost stated, because it is real: adding a second implementation
turns a settled question into an open one for everyone using the bundle,
not only for whoever wanted the alternative. That is the refusing rule
applied consistently, and the alternative is a default, which is the
flavor field returning under a better name.

08-connectivity gains why the network stopped being code beside the
module system: a machine was on the private network because it had an
address, and there was no way to keep one off. A manifest can now say its
resources are computed, which is what a peer list needs.

And three modules rather than one, because WireGuard is one VPN of
several. Naming a module after the job and putting one implementation
inside it is flavor wearing a generic name — the second VPN has nowhere
to go.
2026-08-29 23:21:01 +02:00
jschoubben 554f6bd7a4 A capability may carry a value, and adding one is not free
Recorded while building the seat detector. A capability is a named fact about a
machine: its presence gates an assignment and its detail can carry a value, so
"can this run here" and "what should it be configured as" are the same fact
read two ways. A verdict has always had a detail beside its yes or no, so
panel: oled needs no new concept.

Two things that keep the set honest, both worth writing down before anyone adds
the fiftieth capability. It must be detected and the detector must say how it
knows -- so nobody can add one they cannot check, which is the whole of issue
007. And detectors ship inside the host, which is one static binary, so adding
a capability means shipping a new host everywhere. That argues for a small
general vocabulary rather than a specific one.
2026-08-29 21:12:21 +02:00
jschoubben f140303257 A module claims; it does not list its rivals. And flavor is retired.
Three decisions, all Jochen's, and the first is the one that unlocked it.

Exclusivity is not a property of a module. It is a property of a singular
resource the module takes over. Two shells compete for nothing and any number
may be installed; two display servers both want the seat. So a module declares
what it CLAIMS, and two modules claiming the same thing cannot both be assigned
within that claim's scope.

Not "xorg conflicts with wayland". Pairwise exclusion has a property that only
shows up later: adding a third display server means editing xorg and wayland to
know about it. Every new module requires changing modules nobody who wrote it
owns, and the edits grow as the square of the count. With a claim the third one
says what it claims and nothing else changes anywhere.

Claims have a scope -- node, site, mesh -- which is not new. The mesh already
enforces exactly one hub with a unique index. Scope is that idea said once
rather than hard-coded per case.

And some conflicts need no claim at all: two modules declaring the same file or
binding the same port are visible from what they declare. A claim is only
written for the abstract ones.

A requirement with several answers is refused, never guessed. One candidate is
assigned silently because there was no choice to make; none is refused naming
what is missing; several is refused naming them. That is what makes a solver
unnecessary -- counting candidates has no surprising behaviour, and a solver
can be added later without changing a single manifest.

Flavor is retired. It was carrying three unrelated meanings: variants of a
thing, a subset of a module a node installs, and whatever the current system
does, which earned two knowledge-base entries about going wrong. A word with
three meanings cannot be reasoned about. What it reached for is two ordinary
things -- different modules providing the same thing, and one module with a
setting.
2026-08-29 21:00:13 +02:00
jschoubben 974985b3d1 Four things the lab found about the private network
All on the first three machines to actually run it, and all invisible from the
mesh's own state: the graph was right, the files were right, the services were
up, every node reported success, and the network did not work.

A running interface does not re-read its configuration, so a node joining left
every existing node carrying a network that no longer existed. A hub sharing a
site with a spoke was emitted twice, which WireGuard refuses. Two nodes at one
site that neither can be dialled were peered directly, so nobody opened the
path and the more specific route blackholed -- this document's own warning
arriving in its implementation. And Docker sets the FORWARD policy to DROP, so
a hub with forwarding enabled still carried nothing between its spokes.

The last one is the sharpest: the substrate at tier 1 silently breaks the
network at tier 2, and nothing in either tier's state says so.

None of these is reachable by reasoning, and each was found within minutes of a
real machine trying it. That is the argument for the lab in one line.
2026-08-29 18:03:39 +02:00
jschoubben 6bcf0e4f9f Issue 010 fixed: origins keep the bundle and the mesh apart
The store records where each resource came from and each origin removes only
its own. Verified on the scenario that caused it -- eleven resources raised,
enrolled, sent the same two-resource declaration, and the store, broker and
control plane were all still running. A later declaration dropping a resource
still removed it, so removal by omission survived the fix.

Two more faults found while fixing it, both the same shape. A report published
to a routing key nobody bound vanishes: the broker accepts it, finds no queue,
drops it, and tells the publisher nothing -- so nodes announced what they had
applied into a void. And publishReport was discarding its error, so a node that
could not tell the mesh looked exactly like one that had.

Reports are mandatory now, so an unroutable one comes back and is said out
loud, and the binding covers every key a node may publish.
2026-08-29 16:43:28 +02:00
jschoubben 594ea10b07 Issue 010: the first declaration destroys the substrate
Found in the lab, doing the ordinary thing: raise a first node, enrol it, send
it a declaration. Both declared resources applied correctly and every container
on the machine was removed -- the store, the broker, and the control plane that
had sent the message. The link died mid-sentence because the broker carrying it
had just been torn down by what it carried.

Nothing is behaving incorrectly. Apply removes what the store holds and the
declaration does not name, which is what reconciliation means. The fault is
that the carried bundle and mesh declarations share one store, so the host
cannot tell what this machine raised for itself before there was a mesh from
what the mesh told it to have.

It is invisible until those two meet, which happens exactly once per mesh: on
the first node, after enrolment, the moment the control plane first speaks.

The report says what is not the answer, including the tempting one -- having
the control plane send the substrate back. It cannot: it was never told what
the bundle contained, and the bundle exists precisely because there was no
control plane to ask.
2026-08-29 16:23:01 +02:00
jschoubben 02afb7516b What connecting to the mesh is, and what a node presents
Two things this record never said, both asked directly.

Connecting to the mesh is one outbound AMQP connection from the node to the
broker, held open. There is no second connection and nothing is ever dialled at
a node. Being in the mesh means that connection is up.

Two different things ride on it and conflating them is what made this murky. An
AMQP account, which the mesh issues per node at enrolment, answers whether the
connection is accepted at all -- per node rather than shared, because a shared
one lets any node consume another's queue, which is the shared-credential fault
this record exists to remove reappearing at the transport.

The node's own keypair answers which node is speaking, on every message. It is
not made redundant by the account: with only an account the control plane knows
who is speaking because the broker says so, and that is the same transitive
authority this record already refuses in the other direction. A compromised
broker could attribute reports to whichever node it liked.

So a node holds two things after enrolment -- a credential the mesh issued for
reaching the broker, and a key it generated that the mesh only sees the public
half of. Both are its own, neither reaches anything else.
2026-08-29 15:25:18 +02:00
jschoubben 004057d85c A node's identity is a keypair it generates. This was never open.
I have been treating "what a node presents to prove it is that node" as an
undecided design question for weeks, and blocking on it. It was decided.
08-connectivity says of the overlay keys: each node generates its own keypair,
the private key never leaves the machine, the public key is published to the
mesh -- and says explicitly that this IS ADR 0004's "a node holds its own
identity", applied. Nobody had applied it to the thing 0004 is actually about.

What caused it was a word. The lifecycle said a joining node receives its own
durable identity, which reads as the mesh issuing something, and then the
question is what. The mesh issues nothing. A node arrives holding its identity;
what it receives is being known. That line now says what happens: it presents
the one-time secret and its own public key, which the mesh records.

The rule above it then holds literally rather than aspirationally. The mesh
stores a public key, so a copy of the mesh's database grants nothing, and
compromise of a node really is compromise of only that node.

Also recorded, since it was asked directly: same principle as SSH, own key, not
the machine's SSH host key. Host keys are regenerated by reinstalls and image
clones, which would silently un-enrol a node; their lifecycle belongs to sshd
rather than the mesh; and a partial host has no SSH daemon at all, so an
identity scheme resting on one excludes a supported kind of node.

The good half of that idea is kept: the mesh knows every node, so it can
distribute host keys the way it distributes authorised keys, and node-to-node
SSH stops depending on trust-on-first-use.
2026-08-29 15:21:36 +02:00
jschoubben 5fd522b8da A node is a machine; the session is a feature of it
Correcting an overstatement from the previous commit, where I had written that
a node IS a conversation. It is not. A node is a machine inside the mesh, and
the session is one of the things running on it -- like the host, like any
workload.

That also dissolves the conflict I flagged as unresolved rather than needing
anyone to decide it. 0001 says a node does not authenticate to a model
provider, agents do. Still true: the session authenticates, and the session is
not the machine. The node does not think, something on the node does. I had
manufactured the contradiction by promoting a feature into an identity.

0001's summary row is corrected the same way, and says explicitly that neither
the node's session nor a hired worker makes the node itself a thinking thing --
both run on a machine, which is what leaves that line untouched.
2026-08-29 14:22:29 +02:00
jschoubben 066f14b5f8 A node is a conversation, and that is not the employee model
Moving this out of 0003 and out of its vocabulary. I had spent three attempts
fitting the node's own session into the agent-as-employee record, each time
bending hired, draining, reassigned and retired to cover something none of them
describe. 0003 is back to its original text.

It belongs in 0004, under what a node is, because that is what it is -- not a
program installed on a node but part of the node. It holds one session
permanently, anything in the mesh can message it, and it remembers across
callers and across weeks. Its system prompt is the engram, which is recorded
here for the first time despite running on every node.

Also recorded: it has its own narrower tool list, so it can go and look rather
than only report about itself; there is no authorisation between nodes, because
every node is the operator's own; and how a node passes a question on is its
own business rather than a protocol field.

Switched off it still answers, and that is the point of having an off state
rather than an absent one. A node with nothing there is a silence somebody has
to diagnose. A node that says it is switched off is not. Same rule the host
follows about a service that does not exist.

0001's summary is corrected too: it had one row for "agents", which is the
conflation being complained about. Two rows now. A node's own session and a
hired worker are built from the same parts and run on entirely different terms.

Left standing and NOT resolved here: 0001 says a node does not authenticate to
a model provider, agents do. A node that holds a session does. That is a real
conflict between what is recorded and what runs, and it needs deciding rather
than a fourth reconciliation from me.
2026-08-29 14:17:25 +02:00
jschoubben 079c488d5e Provisioned and immutable beats exempt
Replacing the framing I wrote an hour ago. I had the node's own agent sitting
outside the lifecycle as an exemption, which is a rule somebody has to
remember. Provisioned the ordinary way and constrained is a rule the system
enforces, and it is one row like any other rather than a category every query
listing agents has to special-case.

It also reads the original sentence more carefully. "Exempt from the hiring
lifecycle" is exempt from hiring, not from having a lifecycle. Its lifecycle is
the node's -- provisioned at enrolment, retired when the node is retired. Same
states, a different thing driving them, and no exemption needed.

The constraints are now the four nonsense states written as things that cannot
happen rather than as an argument: not retirable, reassignable or deletable
while its node exists; exactly one per node. And a distinction that was missing
-- its existence is immutable, its engram is not. Freezing the personality
would remove the way a node is configured.

Disabling is the better half of this. A node with no agent is a silence
somebody has to diagnose; a node whose agent is disabled answers saying so,
immediately, with no model invoked -- the queue is still consumed and the state
is the reply. That is the host's own rule about a service that does not exist,
applied one tier up: absence must never be indistinguishable from a failure to
answer.
2026-08-29 14:06:33 +02:00
jschoubben fd7f7557bd The node's own session, and why it is not hired
Answering a question that was asked three times and that I kept not answering:
should the node's session just be an agent per node, since otherwise the
functionality exists at two levels?

Same mechanism, different lifecycle. A persistent session, accumulating memory,
a system prompt, a scoped tool list, addressable by message -- identical, and
building that twice is the duplication the question was worried about. What
must not be shared is the lifecycle, because if a node's own voice were an
ordinary hired agent it could be retired, leaving a node nothing can talk to;
reassigned, moving one machine's mind onto another; hired twice, with no answer
to which one replies; or never hired, leaving a node mute. The exemption in
this record exists to make those four unreachable.

I had this backwards earlier today and said so out loud: I called "a node
itself is an agent of a kind exempt from the hiring lifecycle" a fossil of the
old model and recommended striking it. It is the design. And it does not
conflict with 0001 -- "the two agent rows per node merge" means one per node,
not zero. I read merge as delete and invented a contradiction between two
records that agree.

Engrams are recorded for the first time. They are in use on every node and
appear in no record, which is how a decided thing comes to look accidental.
The engram is the node's system prompt, and it is what makes one node's answers
recognisably its own rather than generic.

Also recorded: there is no authorisation between nodes, because every node is
the operator's own and a prompt from one is a prompt from them. The consequence
is stated once rather than left to be discovered -- the mesh boundary is the
security boundary, which is what puts the whole perimeter on the token and the
overlay.

And how a node passes a question on is the node's choice, not a protocol field.
A node may say who is asking or may simply ask, the way a person relaying a
question decides how to phrase it. That follows from the engram. The cost is
that there is no machine-readable chain of who ultimately asked; each node
still holds what it was asked and by whom.
2026-08-29 14:01:00 +02:00
jschoubben 88ba81e9c1 Agents reaching nodes is the capability, not a hole in it
Correcting what I wrote an hour ago. I had recorded node-to-node SSH as "not a
mesh function" and "a second control path through the back door", reasoning
from ADR 0004's rule that the host has no inbound control surface. That
conflated two different things and got the product backwards.

There is no node-to-node SSH to forbid. The actor is always an agent; a node is
only where it happens to be running -- ADR 0001 already says a node is a place
where an agent can run and that is the entire relationship. An agent hired onto
one node reaching another to do work is the capability the whole arrangement
exists to provide.

The credential is the agent's, in its own credential directory, which ADR 0001
already established. So a node's authorized_keys lists agents and never nodes,
and three things follow: no node holds a key reaching another node, so 0004's
"a node holds its own identity and nothing else" stays literally true; a
compromised node costs the credentials of the agents that were on it rather
than a way into everything; and who may reach what stays a mesh-wide fact,
which is why it is identity's.

The rule I misapplied is about how a node's declared state changes -- over the
broker, never by being dialled. An agent with a shell is not the mesh
reconfiguring a machine, it is what a person with a terminal has always been,
and this design already depends on that working: the overlay is the way back in
when a declaration breaks something. What such a session leaves behind is
drift, and drift is what reconciliation is for.

0001 also stops underselling the fourth layer. It read as "the layer the other
three exist to carry", which is true and flat. The value is that an agent can
work across a set of machines as though they were one -- centrally configurable
machines are ordinary; that is not.
2026-08-29 13:06:40 +02:00
jschoubben 918dc04916 What this actually is, and three things that were assumed
Four things settled by talking them through, all of which had been true in
somebody's head and written nowhere.

It is not a mesh in the peer-to-peer sense and will not become one. 0001 now
says what it is instead: machines linked by a private network, one node holding
knowledge of all of them, modules as the way anything is built and delivered,
and agents hired onto nodes to do the work. The word describes what machines
can reach, not how they are governed. "Master" overstates it the other way --
nothing needs that node to keep running, only to change.

0006 gains the option that would make it a real mesh, recorded as considered
rather than rejected by silence: every node holding the whole inventory, a
replication process, an elected master with promotion on failure. What settles
it is not the complexity but that it still would not deliver the name, because
application databases are not replicated -- so a genuine peer-to-peer mesh
means becoming a replicated database system for every consumer's data too. That
is a larger product than the thing it would support.

Also in 0006: three central roles, not one. Losing the control plane costs
change, losing the broker costs being told anything, and losing the hub costs
nodes in different places reaching each other at all -- which is operation, not
administration. Whether they are one node is not decided.

And SSH access is identity's. It appeared three times as something that uses
the overlay and never as something the mesh provides, which reads as settled
when nothing decided it. Nobody else could: the mesh is the only thing that
knows which humans and agents exist and which nodes they may reach. Node to
node SSH stays out -- the host has no inbound control surface by decision, and
nodes reaching each other that way is a second control path through the back
door.

0007 gains the requirement underneath all of it. Reachability was recorded as a
fact to track and never as a thing some node must have. The broker's node and
the hub must be dialable by every node at a stable address, or nothing can join
and a disconnected node cannot return. A mesh entirely behind NAT cannot be
raised. That is a precondition and it belongs with the others.

The link staying on the underlay is also argued now rather than asserted. At
join time it is forced; afterwards it is a choice, and the reason is that a
repair channel carried over the thing being repaired is not one. Moving it onto
the overlay, with fallback, is recorded as open with what it would have to get
right -- a WireGuard interface has no link state to test, and a silent fallback
is this repository's recurring fault in a new place.

0010 says in one line what was the intention throughout: the module system is
the CI/CD. Not a pipeline beside the mesh. Build, test, publish and deploy are
one reconciliation seen at four points, which is why a thing that cannot be a
module cannot be delivered.
2026-08-29 13:01:36 +02:00
jschoubben 5218b06c02 Fold the control plane's build decisions into 0006 and 0008
Back to 23 records. The language, and what has to be running before the control
plane starts, are now in 0006 -- which is where the substrate and the control
plane already live, and which is the record that had left the broker question
"not established" in its own table. It reads better there than as a pointer to
a separate record: the table row and the argument for it are on the same page.

The store mechanics went into 0008. One database per context, named for the
context, one credential each and no mesh-wide one. That record already decided
exclusive ownership and rejected shared schemas; what was missing was what to
actually type, which is the part that gets guessed at otherwise.

Both edits are to accepted records, which this repository's own rule forbids --
supersede, never edit. Recorded here so it is visible rather than silent. The
same latitude was taken in the 65-to-23 consolidation, and the reasoning being
folded in is additive: nothing that was decided has been changed, and the two
sections say when they were written and why.
2026-08-29 03:08:50 +02:00
jschoubben 82a3065f82 Tier 2 exists, and the token was missing a quarter of itself
mesh-control is built as far as it can honestly go: one context of seven,
inventory, with its schema and the command that applies it. The repos map and
the control plane design say so, and point at ADR 0024 for what it took.

Separately, and more importantly: this repository described the enrolment token
as carrying three things when ADR 0004 says four. The missing one is the
control plane's signing identity -- the reason a node does not have to trust
the broker it dials.

Without it the control plane's authority is transitive through the broker, and
0004 spells out what that costs: a compromised broker could forge declarations,
and since the host applies whatever the link delivers, that is the whole
machine. The record has the argument in full; the design doc had dropped the
conclusion.

Found by reading the two together while deciding what the control plane must
store, which is roughly the only way it would have been found -- both documents
are internally consistent and only disagree with each other.
2026-08-29 02:49:58 +02:00
jschoubben 84f4425fd6 The broker precedes the control plane, and it is written in Go
Two things found by trying to build tier 2.

The substrate design asked whether the message broker has to be running before
the control plane, and framed it as depending on whether the control plane's
own parts talk to each other over it. They do not -- it is one process -- so
under that framing the broker stays out of the bundle.

The framing cannot answer the question. What decides it is how the control
plane reaches a node, and the answer was already decided: only ever over the
link, and the link is the broker. So provisioning the broker would require the
broker. The first node does not escape this by being local, because it enrols
the ordinary way, by dialling the broker at the address in its token -- which
was deliberate, and worth keeping.

The bundle is two images now. The record says what that costs, including a
certificate the broker needs at a moment when there is no mesh to issue one.

The language had never been decided for tier 2. Go, for the same reason the
host is: the bundle pins this image by digest and runs it where nothing can
check it, so the image should hold the program and nothing else.

Also corrects something already built: the bootstrap created one database and
called it 'mesh'. ADR 0008 grants a context only what it exclusively owns and
ADR 0006 says the mesh database names a thing that will not exist. One database
per context, so one today, called inventory.
2026-08-29 02:32:46 +02:00
jschoubben 6a2b107fb8 Restore a consequence the consolidation dropped
I said nothing was lost when 65 records became 23. That was too strong, and
here is a counterexample: ADR 0046's consequence that the lab needs a way to
place images did not survive into the merged substrate record. The compression
kept the decision and dropped one of the things it implied.

It was not lost from the repository -- 04-ISSUES/009 had already picked it up,
which is why it was found at all. But the record no longer carried it, and the
record is where somebody would look.

Restored, now as a resolved fact rather than an open consequence: the lab
raises a registry inside the scenario, which is the real path since that is
what every node after the first pulls from. The digests it serves are its own,
and that satisfies the pinning rule -- what is required is a reference that is
exact and cannot move.

Worth recording the wrong assumption too, because it is what made this look
impossible for two days: I took "pinned by digest" to mean the UPSTREAM digest
had to be preserved. It does not. Any digest that is exact and immutable
satisfies the rule, and a registry assigns one.
2026-08-29 00:06:05 +02:00
jschoubben 087a8f4144 Close 009: a sealed machine now pulls by digest
The resolution was the one the issue predicted -- a registry inside the
scenario -- and it is the real path rather than a stand-in, since that is what
every node after the first pulls from.

The digests are the lab registry's own, which satisfies the pinning rule: what
is required is a reference that is exact and cannot move, and one this registry
assigned is both. That was the insight that unblocked it; I had assumed the
upstream digest had to be preserved, which is what made it look impossible.

The fault worth keeping is recorded in the issue: the read-back checked that
the catalog endpoint answered by matching the substring 'repositories', which
an empty catalog also contains. It passed on a registry holding nothing. This
repository's own subject, arriving in the tooling built to catch it.
2026-08-29 00:05:20 +02:00
jschoubben b4607dfc03 Numbers are identity; the reading order is a generated, checked index
Decided after measuring what renumbering actually costs: 96 references in code
comments across two repositories, none of which would have failed to compile.
They would have pointed at the wrong reasoning, which is worse than a broken
link because nothing reports it.

So a number identifies a record and never changes. It cannot also be a
position -- a position moves when the set changes, and an identity that moves
is not one.

The reading order moves into an index generated from each record's `topic:`.
Six topics, in the order somebody learns the system.

The index is WRITTEN rather than only generated on demand, which reverses what
this repository previously said. The reason it said otherwise is that a
hand-written index drifts -- but a reader looking at the folder on a forge sees
the folder, not a command, and the drift objection is answered by checking
rather than by refusing to write one. That is §5's own rule: a rule states how
it is checked.

Two checks, both confirmed to bite. index.py fails when the written order no
longer matches the records. records.py fails when a record has no topic or one
nobody defined -- the quiet failure being a record that vanishes from the order
rather than appearing in the wrong place.
2026-08-28 23:39:18 +02:00
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00
jschoubben e1febe8e0f Renumber the records 1 to 23
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18,
19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only
the archaeology of what used to be there.

Renumbered contiguously. Renames run in ascending order, so every target number
is already free and no two files ever collide.

The reference rewrite is one simultaneous pass rather than a sequence of
replacements. Numbers moved into slots other numbers were vacating -- the node
host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time
would have cascaded and silently pointed things at the wrong record.

Seven plain-text references survived the merges as prose rather than links,
naming records that no longer existed: the enrolment token, the link boundary,
what a declaration is, reachability, the repository structure. Each mapped to
the consolidated record that now holds it.

Verified rather than assumed: every [ADR NNNN](path) link now has matching text
and target, checked across the whole repository, and the checker passes.

Frontmatter `consolidates:` lists dropped -- they named records that are gone,
and each consolidated record already says in prose what it absorbed.
2026-08-28 23:28:34 +02:00
jschoubben 77f3a4cea7 Consolidate: 65 decision records to 23
Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.

  the node host          8 -> 1    applies not decides, depends on nothing,
                                   per operating system, root service, the
                                   launcher, episodic, what a declaration is,
                                   actions from the bundle only
  a node and how it joins 4 -> 1   what a node is, joining, the link as
                                   security boundary, the enrolment token
  modules and the graph   7 -> 1   everything is a module, no domain modules,
                                   three edges, provisioning, the core library
  substrate and control   6 -> 1   the test, seven contexts, one control plane,
    plane                          the authority is not a database, the named
                                   products, the pinned bundle
  connectivity            3 -> 1   a route is a grant, reachability declared,
                                   filter rules
  delivery                5 -> 1   reconciliation not a pipeline, artifacts,
                                   the three silos, a failed step, the verdict
  the lab                 5 -> 1   (earlier)
  how this repository     10 -> 1  (earlier)
    works

Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.

The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.

The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
2026-08-28 20:03:24 +02:00
jschoubben 5e83ac2c22 Consolidate: 65 decision records to 52
Jochen: a normal application has 3-5 ADRs, maybe 10 for a large one, and we are
at 65. Fair, and the cause is mine -- I recorded every FINDING as a decision
rather than every fork in the road.

Two merges, both cases where one decision had been split across many records
because it was taken over several days rather than at once.

0019 absorbs ten records about how this repository works: what it is and that
it is public, the folder flow, the two design layers, the issue front door,
status in frontmatter, playbooks, the naming rule, the product name. Those were
never ten decisions -- they were one, seen from ten angles as the repository
took shape.

0016 absorbs the five about the lab: a node is a virtual machine, a router is
scenery, a scenario declares the underlay, a scenario is a closed address
space, and the two scenario classes. Same pattern -- one design, split by the
order it was worked out in.

The consolidated 0019 also raises the bar for what earns a record, since that
is what produced 65: a record is warranted when there is a genuine fork -- a
direction reversed, an alternative that will be proposed again, something
contested. A finding is not a decision, and a bug is certainly not. Everything
else belongs in the design document where the reasoning is actually read.

The checker earned its place here. Deleting nine records left 13 dangling links
across the repository and it named every one, including in AGENTS.md. Nothing
was found by reading.

Remaining clusters worth the same treatment: the host (8 records), delivery
(5), modules (6), connectivity (4), substrate and control plane (4). That would
be 52 down to roughly 30.
2026-08-28 18:53:19 +02:00
jschoubben 10365f2eae Consolidate the design layer: one place per topic
Jochen: a jungle of specs that slightly contradict or patch each other, and
what matters is a working state rather than history. Both are fair and both are
mine.

Measured rather than assumed. 05-the-node-host and 09-the-node-lifecycle both
covered enrolment, the install commands, the unit file, the launcher and
reconcile -- I wrote 09 without taking anything out of 05, so the same things
were said twice and could drift apart.

Split by what each document IS. 05 is the component: what the host is, its
parts, the declaration vocabulary, the build order, how it is verified. 09 is
what happens to it: install, enrol, run, upgrade, retire. The whole "The
process" section left 05, and the unit file moved to 09 where installing is
described. 05 goes from 338 lines to 245 and now points at 09 rather than
restating it.

09 also carried a 105-line "Resolved" section -- six mechanisms framed as
"these were open and here is the answer". The content is needed; the framing is
history, and history is what makes a document read as a changelog rather than a
description. Renamed to what it actually is and the was-open phrasing removed.

Also added 10-delivery.md, which did not exist: four accepted decisions --
0054, 0063, 0064, 0065 -- had no design document at all, which is the specific
reason the delivery picture felt scattered. It is now one document covering
modules, the three edges, the core library, and how a change becomes a running
thing, with a table of what each property is designed against and what must
exist before it can be built.
2026-08-28 18:40:46 +02:00
jschoubben 9d091c81e0 A build edge, a core library that is a domain, and 0063 corrected
Three things from walking a real dev cycle through 0063, all of which Jochen
caught by pushing on where I had glossed.

0064 -- a build edge is a third kind. Research 011 established presence and
instantiation, and both are RUNTIME edges: they answer what a module needs in
order to run. Delivery needs a different question -- what has to be rebuilt when
this changes -- and that relationship is fixed inside an artifact rather than
negotiated when it runs. So the graph as designed could not drive delivery,
which is the real reason 0063 was not approvable.

It is derived rather than declared, read from what a module actually imports,
because a declared list and the imports it describes drift and the imports are
the true ones. The runtime edges stay declared, and that asymmetry is not an
inconsistency: a runtime edge is an intention somebody has, a build edge is a
fact about code that exists.

It also makes design quality measurable. A module with many inbound build edges
is one whose every change is expensive, and the current shared library is
exactly that -- nobody could see it because nothing drew the edges.

0065 -- the core library is the mesh's domain. Jochen disagreed with 0030's
"types, not behaviour" and was right: that guard is aimed at the wrong thing. A
library everything depends on is a hub whether it holds types or code, and the
fan-in is what makes a change expensive. So types ship with the module that
owns them -- trading one wide edge for several narrow ones -- and the core
library holds what is true of the mesh regardless of context, which research
011 already found: a module, a node, an assignment.

The test is "would this still mean the same thing in a context that had never
heard of the one it came from". A node does; a pipeline stage does not.
Domain-driven is the point rather than the label: "who else might want this"
always answers yes, which is how the current one grew.

And it changes the check for the better. "The build output contains no runtime
code" would have enforced a rule now withdrawn. Inbound build edges is a
measurement rather than a prohibition, and it is visible while a hub is forming
rather than after.

0063 revised on both counts, plus a third: I had written "the lab judges it" as
though that were a step. A lab run takes tens of seconds, occupies a VM, and
fails for environmental reasons -- and a shared-library change produces dozens.
One expensive non-deterministic gate fails both ways, and neither failure looks
like itself. Verdicts are now tiered, and a run that failed environmentally is
explicitly not a verdict.

0063 also now carries what must exist before it can be implemented, rather than
leaving that to be discovered.
2026-08-28 18:25:28 +02:00
jschoubben 4ab8a0507f Delivery is reconciliation, not a pipeline; research 008 closes
Jochen: don't rebuild the current coordinator, use it as a pitfall list. That
reframed the last open question rather than answering it.

0058 stopped deploy being a stage that pushes to nodes, and said plainly what
it did not fix: detection. A merge that created no pipeline, and nothing said
so. That is not a defect in the detector -- it is what happens when correctness
depends on an event ARRIVING.

0063 applies 0058's move one level up. The control plane holds what source
exists and what has been built from it, and builds the difference. A change
becomes a build because source is ahead of artifacts, which is a comparison
answerable at any moment. An event makes it fast; nothing makes it necessary,
so a missed webhook costs latency and cannot cost correctness.

The mesh becomes one idea at two layers: the control plane reconciles artifacts
against source, the host reconciles machine state against declarations. The
pipeline as a state machine disappears, and with it the stage list that a
verify step was once omitted from.

That reframing answered the three questions still open in 008, so it graduates
with all six closed. A deployed state is two comparisons rather than an event.
A verdict is about an ARTIFACT and gates whether it may be declared -- sharper
than the question expected. And "before self-hosting" mostly dissolves, because
a reconciler needs source and artifacts as bindings where a pipeline's stages
name their targets.

Four costs recorded, and one is a real risk rather than a trade: a reconciler
that cannot reach its target retries forever, and without something noticing,
the failure is silence -- the exact fault this removes, reintroduced elsewhere.
Also named: the run identity people actually use is lost, and "did my change go
out?" needs a replacement or this will be worse to live with than what it
replaces, whatever its properties.
2026-08-28 02:59:42 +02:00
jschoubben f728c3fd98 File 009: a digest-pinned image cannot be placed in the lab
Two accepted decisions collide, and testing found it rather than review.

0046 pins images by digest and has the host refuse anything unpinned. The lab
places images by exporting them from the workstation, because a sealed scenario
cannot reach a registry -- and that loses the digest, since a repo digest only
exists for an image a registry served. Measured: the load says 'Loaded image
ID:' rather than 'Loaded image:', and the image lands dangling.

So a tag is refused by the host and a digest is unusable in the lab. There is
currently no declaration the lab can raise that exercises the container shape,
which matters because the container shape IS the substrate -- every bootstrap
step past the runtime is one.

The resolution is a registry inside the scenario, and that is not a workaround:
0048 already names an OCI registry as substrate and every node after the first
pulls from the mesh's own. It also removes the lab's export-and-push mechanism
rather than repairing it.

0046 now carries a pointer, since its own consequence is where the collision
was predicted -- half of it is closed and the other half turned out to be
harder than 'not solved here' suggested.
2026-08-28 01:56:49 +02:00
jschoubben 9dc57b4712 Graduate 005; record what 0058 answered in 008
Continuing the sweep. Both were answered by records that did not cite them,
which is the same pattern 003 showed -- an effort stays active because the
decision that resolved it was reached from another direction.

005 graduates. Three of its four questions are answered: provider modules do
not group (0044), the ~50 modules that co-change with nothing stay as they are,
and 'group or leave' was never the right pair -- 0054 reframes it as authority
versus package. Worth noting the debt runs the other way too: this effort's
measurement, that reachability is the ONLY place modules genuinely co-change,
is what 0054 rests on and why connectivity is a context while nothing else
needed one.

Its fourth question moves rather than closes. Whether applications leave the
monorepo before or after they group is a sequencing question, so it belongs to
009-migration.

008 stays active, with its central question marked answered: the coordinator
converges nodes on a declaration rather than dispatching stages (0058). The
three-silo split survives with the third redefined. What 0058 explicitly does
NOT answer is how a change becomes a pipeline reliably -- detection is upstream
of everything it changed and remains the fragile input.
2026-08-28 01:40:02 +02:00
jschoubben 0a37d751e2 Graduate research 003; file the rescue that does not exist
A sweep of the nine active research efforts. 003 was answered five days ago and
nobody closed it -- the decision it asked for was taken without citing it, which
is how an effort stays `active` after being resolved.

Its recommendation is what the mesh adopted, and the match is exact rather than
approximate. "Run the daemons as containers, making Docker the supervisor for
everything" is ADR 0057. Its warning that a mesh-native supervisor inherits
fate-sharing "unless it sits outside the mesh's own process tree" is where ADR
0061 put the launcher. And its insistence that it cannot be all-or-nothing is
why the host itself is the one thing an init starts.

Its incidental finding does not graduate with it, so it is now issue 008: the
automatic node rescue the documentation describes does not exist. No unit
declares OnFailure=, nothing calls the rescue script on a timer.

That is worse than having no rescue. A rescue nobody wrote is a gap somebody
can see; a documented one that is absent is a gap nobody looks for, and the
documentation is read exactly when a node has failed and somebody is deciding
whether to intervene.

The issue names two honest resolutions -- implement it, or delete the
documentation and say a failed node needs a person -- and says the choice is
scheduling rather than technical, since the new host's recovery is built and
tested. It also says what would make the finding certain: it came from reading
the repository, and confirming it on a running node is the difference between
"no unit declares this" and "no unit in the source declares this".
2026-08-28 01:39:10 +02:00
jschoubben ba0d01788e 0062: a host may be episodic; 0060's Android gap closed
0060 named the gap and did not close it: everywhere else an init runs the
launcher at boot, and Android grants neither an init to register with nor
anything worth supervising, because a supervisor would be killed alongside what
it supervises.

Closed by narrowing what is required rather than building something. A host is
resident or episodic, and both are hosts. Being killed by the platform is
disconnection, which 0036 already made ordinary -- and every mechanism an
episodic host needs already exists because it was built for laptops that close.

A partial host can join a mesh and cannot be the first node, since every
bootstrap step is a shape it refuses. Its bundle says so.

Two consequences that are easy to miss: last-heard-from means much less on an
episodic host, so a healthy phone reads as a dead server unless the reader
knows which kind it is; and a declaration may take a long time to land, which
makes 0058's outstanding-versus-failed distinction load-bearing.

Still open, and in that order: what an Android node is FOR, and only then how
it is started.
2026-08-28 01:24:07 +02:00
jschoubben f1b1cd9aa0 Review: three ADRs no longer said what we had concluded
A sweep for claims overtaken by the last few days. Annotated rather than
rewritten, following the pattern already in 0049 -- what changed and why is the
useful part, and an accepted record should not quietly become something else.

0057's init section was wrong on all three of its claims. It said the host
needs FOUR things from an init; 0061 reduced that to one. It said every machine
the mesh targets already has systemd; Alpine does not, and it is the intended
first node. It said there is no second init to abstract over; there is now, and
the answer is still not an abstraction -- it is a four-line file per system.
What survives is the part that was always right: an init is not a dependency in
0041's sense, because it is not installed, it is what the machine already is.

0048 named Docker as the container runtime. It is now docker or podman,
detected rather than chosen -- because adoption keeps what a machine already
has, so naming one contradicted a rule already decided. That row is the only
one of the five that names two, and the record now says why.

0060 claimed the bundle is portable across operating systems. Its mechanism is;
its contents are not -- package names, unit names, service names all differ, so
an Arch host embeds an Arch bundle. That was my error, and it is the exact
confusion behind the question that found it.

The design layer had the same drift: 07 and 09 said "Docker" where they meant a
container runtime, 09 said systemd restarts the host after an upgrade when the
launcher does, and both install snippets assumed Arch. They now show Alpine and
Arch side by side, which makes the point better than prose did -- step 1
differs per system, step 2 never does.

Checked and NOT changed: 0047's "the vocabulary grows by one shape" is a claim
about the rate, not the count, and is still true. 0037 lists docker among tools
the host manages, which it does. 0041 says nothing about either.
2026-08-28 00:43:47 +02:00
jschoubben 66df0eb53e 0061: the launcher supervises; init is asked for one thing
The record said an init is asked for two things -- start at boot and restart on
exit -- which was half a change. It moved the give-up logic out of unit files
and left the restart in one, so init still decided when the host came back.

The launcher no longer execs the host. It supervises it, so restarting is ours
too, and init is asked only to run it at boot. There is an OpenRC script beside
the systemd unit now.

Records the cost honestly: not exec'ing means the launcher must trap the
shutdown signal and pass it down, because a supervisor that exits while its
child runs leaves the host to be killed rather than to stop.

And records what the implementation found: the counter counts consecutive
FAILURES, not starts. Counting starts meant a host that upgraded itself three
times rolled itself back, having worked perfectly every time -- because a clean
exit IS the upgrade path. That is now the second time a clean exit has been
mishandled, so it is called out as the thing to check.
2026-08-28 00:38:18 +02:00
jschoubben c557f99cba Record what testing podman actually showed
0060 said the container runtime was a separate decision. It is now made, and
the reasoning is worth keeping because it is the opposite answer to the same
question one paragraph earlier.

Abstracting service managers is lossy -- systemd and OpenRC are different
models and LoadState has no equivalent. Container runtimes converged on one CLI
deliberately, so almost nothing is lost: checked against podman 6.1.0, run,
rm -f and docker's own template syntax for state and labels all work unchanged.
Only the probe differs. So: a two-entry lookup, not an interface.

The difference that is NOT in the CLI is the one that would have shipped
silently. Podman accepts --restart unless-stopped, records it, and has no
daemon to act on it -- containers do not return after a reboot unless
podman-restart.service is enabled, which by default it is not. Every command
reports success and the effect does not happen.

That belongs in the declaration rather than the host: a node using podman is
told to enable the unit. Which is what made the service shape's missing 'boot'
field visible, and it is now built.
2026-08-27 23:59:05 +02:00
jschoubben e1ad39b500 Per-OS hosts, and an init asked for only start and restart
0060 -- the host is built per operating system. systemd and pacman are the Arch
host's implementation, not abstractions the mesh has to grow. They are not
independent choices: a machine has pacman because it is Arch, and the package
manager, service manager and packaging format arrive together as one decision
somebody made at install time.

Rejected abstracting them, and the reason is correctness rather than effort.
The service applier reads LoadState to tell "not installed" apart from
"stopped", which is what stops it reporting absence as success. An interface
spanning systemd and OpenRC degrades to what both express, and the lowest
common denominator is exactly where that fault lives.

Almost all of it is shared -- the vocabulary, store, apply loop, read-back
discipline, refusal model, bundle and link are portable. Two appliers differ.
And delivery was already per-OS, since a .pkg.tar.zst is an Arch artifact, so
this is the seam that already existed.

Android is the interesting case rather than Debian: no service manager, no
package installation, usually no root. Such a host implements file, directory
and action and refuses the rest -- the same refusal a host already gives an
unknown type, with a different reason. Those three are the portable floor.

The container runtime is deliberately left open: it is not an OS split, since
Arch runs docker or podman.

0061 -- the init is asked for start-at-boot and restart-on-exit, and nothing
else. Both are expressible in OpenRC, runit, s6 and an Android init.rc.
Counting failed starts and rolling back moves into a launcher, because that is
the one piece which must work when the host does not, and a script with a
counter can be tested where OnFailure= can only be hoped for. Supersedes 0059,
keeping its reasoning in full.

The checker found all six places citing 0059 and refused the commit until they
named the replacement.
2026-08-27 23:46:11 +02:00
jschoubben dcc4b8339c Say who consumes the broker and who writes the registry
Left implicit by the previous commit, which said the owning context writes
without saying what does the consuming.

The control plane is the consumer, and there is one of it. Seven contexts but
one deployable, so it is one process dispatching internally rather than seven
consumers racing -- which matters because the as-is records two consumers
accidentally sharing a queue and silently splitting the traffic, each getting
half of what it expected. With one consumer that cannot arise.

The broker is also the buffer while the control plane is down: nodes keep
publishing, messages queue, the control plane drains them on return. That is
what makes a single control plane tolerable -- an outage delays the mesh's
knowledge rather than losing it.

One consequence named because it will otherwise be discovered: an unbounded
queue grows until the broker's disk is full, and the broker is the component
every node depends on. The bound is per queue and undecided -- dropping the
oldest health report is obviously right, dropping the oldest declaration
acknowledgement is not.
2026-08-27 22:20:03 +02:00
jschoubben 19997d56c3 Approve 0057-0059, with four corrections from review
Not approved as drafted -- four things came out of checking them against each
other, and one was a bug that would have broken every upgrade.

The bug: 0059 specified Restart=on-failure while 0057 has the host restart onto
a new binary by exiting CLEANLY. on-failure does not restart a process that
exited zero, so every upgraded node would have been left stopped, having
successfully upgraded. Found by reading the two records against each other
rather than by either alone. Now Restart=always in all three places that
mention it.

The host cannot run in a container, and the reason is decisive rather than
stylistic: step 0 of the substrate bootstrap installs the container runtime, so
a host inside a container would need the thing it exists to install. It would
also break 0041 -- copy it onto a machine and run it stops being true when the
machine must already have a runtime. Everything above tier 0 is a container;
the host is not. That split is the tier boundary, not an inconsistency.

systemd is named rather than abstracted. An init is not a dependency in 0041's
sense: 0041 is about what must be installed before the host works, and an init
is not installed, it is what the machine already is. The unit file is the only
systemd-specific artefact and it belongs to the package, so a machine with a
different supervisor ships a different package.

The mesh is a watchdog, and my first draft was half an answer. Recovery must be
local -- nothing dials a node, and a host that cannot start cannot report. But
detection is the mesh's, and a local supervisor structurally cannot do it: it
sees one process failing and cannot tell a broken machine from a broken
release. Only something watching every node can, and that distinction decides
whether the response is "fix this machine" or "stop shipping this version". So
a host rollout is staged -- a few nodes, wait for heartbeats, continue or stop
on silence. Local rollback still needed, because the canary nodes break and
because a node offline during the rollout gets the declaration later with no
batch around it.

The first declaration is the overlay and nothing else. Forced, because a node's
address and peers are assigned rather than chosen. But also the way back in: a
node reachable over the overlay can be fixed by hand if a later declaration
breaks it, and a large first declaration risks a node that is broken and
unreachable at once.

Also stated plainly, because it reads as a contradiction: nodes reach each
other over the overlay and every node consumes from the broker; what 0039
forbids is an inbound CONTROL surface, not reachability.

And in 06: no node holds a credential to any control-plane store, for reads or
writes. Four ADRs already say this separately and none of them said it in one
place. Nodes state over the broker; the owning context writes. With a note that
most high-frequency writes are observability's, not the registry's -- routing
logs into the registry would be the shared-schema mistake arriving through a
door marked performance.
2026-08-27 22:12:12 +02:00
jschoubben 605c9fd441 Changes are pushed, not polled; and a stuck host rolls itself back
Two corrections and one new decision, all from Jochen catching things.

Pushed, not polled. I described updates as landing "on the next reconcile",
which reads as polling and is not the design. A declaration arrives as a
message on a link that is already open; the host applies it then. Polling over
an existing connection would be slower to land AND constant traffic to learn
nothing.

The timer is for drift and nothing else, and it cannot be replaced by an event
for a definitional reason: drift is change the mesh did not make -- somebody
edited a managed file, a distribution upgrade replaced a config -- so nothing
will ever publish a message about it. Only looking finds it.

Separated the heartbeat from the reconcile timer, which I had been conflating.
They point in opposite directions and answer different questions: the timer
looks at the machine and asks whether it still matches; the heartbeat reports
upward and is what makes silence mean something. A node with nothing to do
sends nothing, and without a heartbeat that is indistinguishable from a node
that stopped.

0059 -- a host that cannot start is rolled back by the service manager. I had
left this open on the grounds that recovery meant the host judging its own
health. That objection does not survive being asked properly: a keepalive is
something else judging the host. The watchdog must be local, because nothing
dials a node and a host that cannot start cannot report -- so it is the service
manager, which is already there.

The failure it prevents is sharper than "the node is down": a host that will
not start looks exactly like a machine somebody switched off, which is the one
condition this design has deliberately decided not to alarm on. So a bad
release reaches every node, each goes quiet, and the mesh reports a fleet of
sleeping laptops.

Confirmed means started and completed one reconcile -- deliberately not "the
link is up", or a laptop on a train would roll itself back. The rollback is a
script shipped by the package, not a host subcommand, because a binary that
will not start cannot be its own recovery. It rolls back once: a second failure
means the machine is the problem, not the binary.

Also refined the records checker, which produced a false positive: a proposed
record may extend another proposed one, because decisions are drafted in chains
and the alternative is marking things accepted to satisfy a check. An accepted
document resting on a proposed record still fails, and that was verified.

0057, 0058 and 0059 are all proposed.
2026-08-27 22:04:26 +02:00
jschoubben aeea2a9f9a Resolve the host lifecycle's open items, and say how the host is delivered
The upgrade question turned out to be a delivery question, so 0058 answers
both.

Today's third silo runs once per node and sends each one a command to install
and start. That is where the as-is records a package install that 404ed from
every mirror while the job went green, an image pull failure that did not fail
the deploy, and a verify stage that was built and never scheduled because it
was missing from a list.

The shape underneath all of those is that the thing reporting success was not
the thing doing the work. Meanwhile ADR 0037 has given every node a component
that applies state, reads back and reports -- so two mechanisms now change a
node and only one checks its work.

0058: a pipeline ends when the declaration is updated. Deploy stops sending
commands to nodes and becomes one write. The host applies it on its next
reconcile, and the host cannot report success it did not verify. The verify
stage disappears as a stage, which is the point -- verification stops being a
step that can be left off a list.

A pipeline result now means "the declaration is updated, and here is which
nodes have applied it". It does not wait for every node, because a node may be
legitimately switched off for a week. Outstanding is reported separately from
failed, since conflating them is how the old system produced a stall with no
error anywhere.

The host is delivered by exactly this path and needs no new resource type: a
`file` writes the package manager's config pointing at the mesh's repository, a
`package` names the version. Added a step I had missed -- before exiting for a
restart, the host runs the new binary once. A package can install something
that does not execute here, and that turns "the node never came back" into "the
apply failed and said why".

Six open items resolved: re-enrolment is decided when the token is issued and
revokes the previous identity; the mesh keeps a recovery copy of what each node
reports it owns, which un-strands the orphans; last-contact is reported with no
threshold, because a laptop off for three weeks is doing nothing wrong;
adoption always completes but a failed line makes a node ineligible for
assignment; a briefing is a structured document whose outcome is computed from
its lines; and the token is printed once and carried by hand, which is the
property that makes it worth anything.

Still open and named: automatic rollback of a host version that will not start.

0057 and 0058 are both proposed.
2026-08-27 21:53:38 +02:00
jschoubben 2204b01909 Design the node lifecycle end to end
The host was described as a component and never as something that runs for
years on a machine somebody else also uses. 09 covers every state a machine can
be in and every transition between them.

Four states: unmanaged, hosted, enrolled, disconnected. Only the last two are
nodes, and they are the same node in two situations. `hosted` -- the host
installed but never told which mesh it belongs to -- had no name before and is
where a machine sits between the two adoption commands.

Things that were unclear and now are not:

The first node walks the same path in an unusual order: reconcile from the
bundle, the control plane it just raised issues a token, enrol against it. Its
specialness lasts two commands. A side effect worth having -- enrolment is
exercised on node one, rather than being written and first used on node two.

Enrolment reports profile and inventory BEFORE the control plane decides
anything. The profile is the input to that decision, not a diagnostic; the
control plane cannot decide what a machine should run without knowing what it
can run.

Rebooting mid-apply is safe by construction. The store records each resource
after it worked, so a host that dies half way through comes back and applies
the rest. The rule that stops the host lying about what it did also makes it
crash-safe.

Retiring splits in two. Graceful is a final empty declaration. A node that is
gone will reconcile its last declaration forever -- the honest consequence of
making disconnection ordinary. The answer is not to make the host expire but
that the node holds nothing that outlives revocation: every grant is a per-node
credential revoked at the provider. A lost node keeps running and stops being
able to reach anything. Said plainly rather than implying the mesh can switch a
machine off, which it cannot and should not.

Losing the store is quiet and permanent, so it gets its own section. The host
re-enrols and re-applies fine; what does not come back is removal, because
resources it no longer has a record of become unowned and sit there
indefinitely.

Also corrects 0057, which said the mesh must not upgrade the host at all. That
conflated two acts. Replacing the binary is safe -- Unix keeps the running
inode. Stopping the unit is not. So the host may apply a package naming itself,
and restarts by finishing its apply and exiting cleanly, letting the supervisor
start it on the new binary. It never asks the service manager to restart it.
That makes a fleet-wide host upgrade an ordinary declaration, which the first
draft gave up on.

0057 remains proposed.
2026-08-27 21:16:44 +02:00
jschoubben 3ab11c96ef Say what the host process is: a root service, installed as a package
The design described what the host does and never what it is at runtime. The
words daemon, long-running, interval, poll and heartbeat appeared nowhere in it
or in the relevant decisions. What exists is a command that runs and exits;
what the design needs is a process holding a link. Nobody had written down that
those differ, so several questions had no answer.

0057 settles them. It runs on every node -- the host is what makes a machine
managed, so a machine without one is not a node. Root, because no useful subset
of the job is unprivileged. A systemd unit, because something must survive a
reboot to hold the link.

It never manages its own unit. The temptation is obvious and it ends with a
host stopping itself half way through an apply, leaving a machine with nothing
running to fix it. The installation owns the host; the host owns everything
else.

Installed as a package, with a tarball as the floor. The package carries the
unit file, the state directory and an upgrade path, which a bare binary does
not. But the mesh's package repository is hosted on the mesh, so any route that
needs the mesh to install the thing that joins the mesh is a circle -- the
tarball is the path that must never acquire a dependency.

Reconciles on start, on a declaration, on a timer and on reconnect. The timer
is the one easy to leave out, and without it `owned` reports what the host
applied rather than what is there -- ADR 0035 violated by omission.

The records checker caught this commit on its first attempt: 05 listed 0057 in
its frontmatter while 0057 is still proposed, and a to-be document may not rest
on an unaccepted record. The section now says so in the body instead.
2026-08-27 21:06:00 +02:00
jschoubben 2330d74c1b The host's vocabulary is complete; 05 and 07 said otherwise
All six shapes are built. 07 still said the last three did not exist, and 05
still described stage 2 as having built three of six.

Records what the lab still cannot do, because that is now the only thing
between here and an end-to-end substrate bootstrap: a sealed scenario cannot
fetch an image and its machines carry no container runtime, so package,
container and action were verified against a real machine instead.
2026-08-27 20:36:58 +02:00
jschoubben 03874f3fe2 Add a structural check over HQ's own records
Nothing in this repository was verified by anything but reading, which is how
a superseded decision stayed live in the constitution and in the to-be README
at the same time. Both were found by a person looking, and nothing stopped a
third.

Five checks: links resolve; `decisions:`/`extends:` name records that exist and
are accepted; a governing document citing a superseded record must name its
replacement in the same paragraph; supersession is symmetric; filename number
matches heading number.

Each was made to fail before it was made to pass. The live-citation check was
verified against a reconstruction of the actual incident -- the to-be README
citing ADR 0017 as live guidance -- and reports it with file and line.

It found one thing nobody had noticed: ADR 0018 never declared that it
superseded 0011, though 0011 has named 0018 as its superseder since August.
Fixed.

Deliberately not checked, and said so in the README: 02-DECISIONS and
01-RESEARCH may cite superseded records freely, because a decision record
discusses history and research records what was observed. 00-as-is may rest on
one, per 0056. Flagging those would put noise on correct documents, and a check
that cries wolf gets suppressed -- which costs more than not having it.

Two bugs found by running it: the frontmatter reader iterated an inline list as
characters, and the as-is exemption was missing entirely.
2026-08-27 20:18:35 +02:00
jschoubben e1f4c7d9e0 Approve 0054-0056, apply them, and fix the two smaller findings
0003 is now superseded by 0056. Nothing is left proposed.

Applied:
- 06 corrected from ten contexts to seven plus the api, each row now stating
  why it passes the more-than-one-node test. work, knowledge and stream are
  named as mesh-hosted rather than dropped; `ai` folds into config; `record`
  is deferred explicitly rather than listed. Its frontmatter now cites 0055.
- how-we-build §4 amended per 0054, and the derived page republished by
  playbook 05.

The sync found the drift the playbook exists to catch: the published §4 and
the source did not say the same thing. The source said "four accidents, not
four boundaries"; the published page said "one intent expressed four times",
and only the published page carried the scope caveat. Same rule, two texts,
already diverging. Verified the republish by reading back -- the new rule is
present and the old section's body returns nothing -- rather than trusting the
success message.

The two smaller findings:
- 0051 separated the transport identity from the declaring authority. It said
  the token carries "an address" and "the identity to expect" without saying
  what the node dials. It dials the broker, so pinning only that would make the
  control plane's authority transitive and let a compromised broker forge
  declarations -- which, since the host applies whatever the link delivers, is
  the whole machine. The token now carries four things, and declarations are
  signed and verified per declaration. Cost recorded: rotating the signing
  identity is fleet-wide.
- 0026 no longer restates 0022's rule about generated views. 0022's own words
  are "prose does not restate status; one place, and two is one too many",
  which is what 0026 was doing to it.
2026-08-27 02:21:34 +02:00
jschoubben f49d177a31 Draft three records for the contradictions the review found
0054 -- things that change together share an authority, not a package.
The constitution instructs agents to group "how a node is reachable" into one
module, citing superseded ADR 0017; ADR 0044 says there is no networking thing
to install. Since the constitution is injected where work is decided, the
superseded rule is the one actually steering work. The observation behind it
was right -- research 005 measured that reachability is the only place modules
genuinely change together -- but the conclusion was wrong: tight coupling means
a shared authority, not one artifact. wireguard and traefik deploy to different
node sets, so the merged module would be assigned where half is unwanted.
Requires amending how-we-build and republishing the derived page.

0055 -- the control plane is the node-coordinating contexts.
Three context lists were in circulation (0015 says nine, 06 says ten, the
README said eight) and none was decided. Research 006 said explicitly that the
change "belongs in a new record -- not written here", and the design used the
list anyway. Reconciling them shows `stream` and `ai` were dropped with no
reasoning at all. Applying 06's own test -- needs to know about more than one
node -- gives seven contexts plus the api, with work, knowledge and stream as
hosted applications and `ai` folded into config as an ordinary grant. The
record defers rather than lists.

The cost is stated rather than reassured away: a board composing across the
boundary reads more than one interface. That was raised before as "only moves
the problem up a layer", and the answer is that 0045 already requires surfaces
to read interfaces rather than stores -- what changes is the count, not the
kind of work.

0056 -- the authority is the control plane, not a database.
Every clause of 0003 has been decided against in four separate records and it
is still accepted and cited as live. The error underneath is the same category
error 0054 corrects: "source of truth" named a storage location when it meant
an authority, and once the store is the answer, shared schemas follow. The
half that was right -- the repository defines what exists, the mesh defines
what runs where -- survives untouched. Best consequence: the cache mode
disappears, so a node that has not heard from the mesh is no longer
indistinguishable from one that has.

0003 is left accepted until 0056 is.
2026-08-27 01:35:37 +02:00
jschoubben ef5dd0751b Approve 0049-0053; drop a to-be item superseded by ADR 0044
The 'domain grouping' item cited ADR 0017 as live guidance. 0044 superseded
it -- there is no domain module to group into, so there is no domain list to
settle.
2026-08-27 01:00:29 +02:00
jschoubben ccbbfa9c8a One node runs the control plane, and nothing takes over
Closes the two open questions in 06 and 08, which turned out to be one
question: how many control planes run, and what happens when the hub is down.
Both were drifting toward redundancy by default -- a standby plane, a second
hub, an election to pick between them. That is not one feature but a property
every layer must then honour, and each layer gets it wrong independently.

Not wanted, and not needed. A handful of machines with one node hosting the
registry is not a distributed system.

The argument for why this is sound rather than merely cheap is that the design
already tolerates it by construction. ADR 0036 makes reachability state rather
than class; the host reconciles from its own store (0043) and never needed to
ask anybody to hold the state it was last given. So the control plane being
down is not a new failure mode -- it is every node in the ordinary disconnected
situation at once. What is lost is change, not operation.

No node holds a contended role: the control plane is assigned like any other
module, and the overlay hub is declared (0050). No promotion, no quorum, no
fencing, no split brain, no replicated store, and no "which node is
authoritative" recurring at every layer.

Two consequences stated plainly rather than buried. The control-plane node is a
single point of failure -- deliberate, and said out loud so it stays
deliberate. And recovery is restore rather than failover, which makes backup
the availability story rather than hygiene.

The sharpest one is the clock: the control plane owns certificate issuance
(0049), so an outage outlasting a renewal window expires every public name.
That bounds how long recovery may take, and nothing measures it today.
2026-08-27 00:55:10 +02:00
jschoubben 4e80820e2f Design connectivity in full: overlay, resolution, exposure, filtering, certificates
Written as one document because the five are one design. They share inputs,
they must agree, and every one of them today is computed in a different place
by a different module from a different copy of the same facts.

The through-line is that none of the five can be answered by a machine alone,
so all five are decided centrally and delivered as `file` resources. That costs
no new host vocabulary and removes both remaining direct database connections
from nodes -- wireguard and traefik are the only two, and both are connectivity.

Three decisions fall out, all proposed:

0050 -- reachability is declared, not inferred from an address. The RFC1918
regex is wrong for carrier-grade NAT (100.64/10 tests as public, so an endpoint
is written to an address nothing can reach), wrong for IPv6, and wrong for a
routable address behind a closed firewall. The lab needing TEST-NET-3 to
satisfy the regex is the same bug from the other side. Also kills hub election
by address prefix, which fails silently and makes renumbering an outage.

0051 -- the enrolment token carries where the mesh is and how to recognise it.
Closes two circles with one mechanism: verifying the mesh needed the CA, and
obtaining the CA meant trusting whoever handed it over; and a node had to reach
the mesh before it could resolve any mesh name. An address plus a fingerprint,
carried out of band, resolves both -- and closes the CA question 0049 deferred.

0052 -- a filter rule names its source. `scope:` is declared in five manifests,
is part of no rule type, and is referenced by no code, so those manifests
appear to restrict ports and restrict nothing. Removed rather than implemented;
the general fix is refusing unknown keys, which the host already does and
manifests do not.

Also corrects two claims in 0049 asserting wireguard was already handled.
Research 006 says both modules still reach upward; neither is.
2026-08-27 00:36:49 +02:00
jschoubben 8d9282d86b Resolve the ingress gap: a route is a grant
ADR 0048 named ingress as an unclosed hole -- nothing said what terminates
TLS, how a public name reaches a container, or which tier owned it. Resolving
it needed no new concepts, which is why it survived: nobody had applied the
rules already written to it.

Ingress is not substrate. The control plane does not need a route to start,
and no node needs one to reach it -- the node dials out and has no listening
control surface. It grants itself a route afterwards, like a bucket.

A route is an instantiation edge under ADR 0044. The direction mirrors a
database -- the consumer supplies a target and receives a name rather than
credentials -- but it is the same edge.

The substantive finding is that exposure is three facts at two scopes: name
resolution and certificate issuance need to know which node is publicly
reachable, and only the proxy mapping is a single machine's business. That is
why it belongs to the connectivity context, and why Traefik doing all three on
the node is wrong.

Which matters beyond tidiness: research 006 counted traefik as one of two
modules opening a direct Postgres connection, reading nodes and mesh_ca. That
violates 0037, 0045 and 0039 at once, and is why every node permanently holds
a credential to the control plane's database. Deriving the config centrally and
delivering it as `file` resources removes it, costs zero new host vocabulary,
and closes the set 0039 identified -- wireguard was the other.

Left open deliberately: the mesh's internal CA is the other thing traefik
reads, and it belongs to the link's mutual authority, not to exposure.
Conflating the two is what made the gap hard to see.

Also fixes an inconsistency from the previous commit: 06 still claimed the
virtual host was raised from the bundle.

Proposed, not accepted -- for review.
2026-08-27 00:22:52 +02:00
jschoubben 4d19e93900 Name the substrate's actual products
The design layer described every service by role and never once by name:
Postgres appeared in zero design documents. That was over-application of the
research rule "never identify the mesh it observed", which is about node names
and domains, not software.

Two things were actually broken by it. substrate.lock pins images by digest and
a digest belongs to a named image, so the bundle could not be written from the
design. And a reader could not tell a settled choice from an unexamined one --
"a relational store" reads identically either way.

ADR 0048 names them: PostgreSQL, LavinMQ, MinIO, an OCI registry, Docker. The
argument for each is continuity, which is a real argument -- replacing a
substrate service migrates the mesh's own state. Role and product are now both
written, because the design depends on the protocol while the installer needs
the product.

Also separates two questions the substrate doc had merged: being substrate and
being in the bundle. Only Postgres must precede the control plane; the rest are
substrate by role and ordinary by delivery. Whether the bus joins it is left
open, because it turns on the control plane's internal shape.

Names the forge as Gitea, and records ingress/Traefik as an unclosed gap rather
than a naming one -- nothing says what terminates TLS or which tier owns it.

Fixes a miscount: the host's bootstrap vocabulary is six shapes, not five.
2026-08-27 00:11:38 +02:00
jschoubben c631cbd07c The bootstrap starts a step earlier than recorded
Asked whether postgres has to be installed, and the answer exposed a missing
step. The store is a container, so something must run containers before anything
else happens — and a container runtime is a PACKAGE, not a container.

Step 0 is where several threads meet. It is what the host's capability detection
already reports, and the first use of that report by something other than a
person. It is adopted rather than installed when the machine already has a
runtime with configuration somebody chose. And it is a package, needing the
machine's own package manager and a network, both of which ADR 0046 permits.

So the host's bootstrap vocabulary is six shapes: package, container, file,
directory, service, action. Stage 2 built three of them.

The node host design now names which three remain and why the lab cannot yet
exercise them — a sealed scenario fetches nothing and its machines carry no
container runtime, which is lab-installation work rather than a constraint on
the design, because production machines have a network.
2026-08-26 23:58:19 +02:00
jschoubben 93470f6162 ADR 0047 — the bundle may carry actions the link may not
The bootstrap's sharpest open question, and the framing was wrong. "State on
this machine" was being read as the filesystem and the service manager. A
service running on this machine IS part of this machine — writing a file and
creating a database in a local store differ in mechanism, not in scope.

The real question was underneath: must the host learn what a database is? It
must not. Giving it a `database` resource type means tier 0 knows Postgres, then
a bucket, then a virtual host — the host acquiring the substrate's vocabulary
one service at a time, which is what ADR 0037 exists to stop.

So the bundle declares an ACTION and the host runs it and verifies it. What a
database means stays with the module that provides one; the host knows only how
to run a declared action against something local and check the result. Its
vocabulary grows by one shape rather than by one resource type per service.

Actions are permitted in the bundle and forbidden over the link, and the
asymmetry is deliberate. A bundle arrives WITH the binary: anyone able to put a
hostile action in it could equally have put it in the host itself, so refusing
actions there buys nothing and costs the bootstrap. The link is a separate
party, reachable separately, and an action there is the unbounded blast radius
ADR 0039 refuses. That decision stands unchanged.

And ongoing provisioning is not the host's at all — the control plane does it
once a mesh exists — so the asymmetry costs nothing.

Which dissolves the earlier worry about one mechanism with a tier boundary
inside it: there are two mechanisms, with different actors, scopes and trust
models, and that is the answer rather than a compromise.

Named rather than hidden: this is the escape hatch research 011 warned about,
arbitrary code in the place hardest to remove later. It is bounded by being
bundle-only and by every action having to declare how it verifies itself, and
that boundary is the whole defence.
2026-08-26 23:56:39 +02:00
jschoubben 5b3d0ebd4f ADR 0046 — the installer fetches what it pins
The blocking question was where a container image comes from, and the version
that blocked assumed the machine might have no network. That assumption came
from the LAB: a scenario is a closed address space by design, which is what lets
two scenarios hold the same addresses without meeting. Production is not sealed
— a machine being adopted has a network, and one that does not is a machine
where very little works anyway.

So substrate.lock carries references, not payload: an image name and a digest,
fetched at apply time. A first node pulls from upstream because no mesh registry
exists yet; every node after that pulls from the mesh's own. The lab is the
exception and places images itself, the way it already places the host binary —
a property of a test environment, and letting it dictate the production design
would be the tail wagging the dog.

Pinned by DIGEST rather than tag. Reproducibility comes from pinning the
identity of a thing, not from carrying its bytes, which is what makes fetching
acceptable rather than a compromise.

ADR 0041 survives untouched, which was the point. "Copy it onto a machine and
run it" stays literally true — one binary, a few megabytes, which then fetches
what it was told to. Carrying images would have quietly redefined the property
that decision rests on.

Costs accepted and named: an apply can now fail because something is
unreachable, which a self-contained artifact could not, so it must fail legibly
— naming what it could not fetch and from where. And the lab needs a way to
place images into a machine that also has no container runtime, both of which
are lab-installation concerns and neither solved here.

Research 012's build-time-versus-apply-time reframing narrows accordingly: it
still holds for what a tailored installer contains, and no longer has to hold
for images.
2026-08-26 23:52:46 +02:00
jschoubben 0531d6fc38 ADRs 0044 and 0045 — the module design, closed; 011 graduates
011 opened asking what a graph deletes and found the graph already existed. The
work became design, worked through twenty cases and one provider in full. Two
decisions close it.

0044 — a module declares presence, instantiation and exclusion. Two kinds of
edge because a game wanting a database is not a game wanting postgres to exist:
one creates something per consumer, carries credentials back, can be revoked,
and leaves the provider holding state. Names are concrete unless providers are
genuinely substitutable — `terminal` passes, `database` fails, and the adapter
is what creates an interface. Where there is no contract there is a tag, which
describes and does not bind. Exclusion is a third relation and is not derivable.
A node provides names too, which makes capability checking stop being a separate
mechanism and makes the host's detection an input to resolution. Constraints,
never placement. Scope decides which provider and the binding is written down
and sticky, in a place that follows the scope. And there are three entities, not
two — the assignment carries what belongs to neither end, which is what
node-agnostic modules ran out of.

0045 — a context owns its store, exclusively. No shared writes and no read roles
on another context's store, because reading couples you to its layout just as
firmly and invisibly. The unit is the CONTEXT, not the process: a board showing
the mesh's own data is the mesh showing its own data. Asking or subscribing is
derived from ADR 0036 rather than chosen. And it is the first clear list of what
the design removes: grant kinds, table ownership, cross-context migration
ordering, and a class of permission modelling.

0017 is superseded rather than narrowed — its text unchanged, its status
changed. Folders assert relationships where edges record them, and the domain
module goes with it.

Left explicitly undecided in both: what a resolver delegates rather than
reimplements, how many instances a module should have, and what a provider hands
back.
2026-08-26 23:41:18 +02:00
jschoubben 60aea14935 Define the substrate, and answer 006's four-or-five conditionally
Same gap as the control plane: load-bearing and unpinned.

The substrate is what the control plane CONSUMES AND CANNOT GRANT ITSELF. Every
module needing a database asks provisioning for one; the control plane needs one
too and cannot ask itself, because it is not running yet. That circularity is
not an awkwardness to work around — it is the definition, and anything on the
wrong side of it must be raised by the bundle the host carries.

Which answers 006's open question in the honest form rather than with a number.
The identity provider is substrate only if the control plane DELEGATES
authentication — then it cannot serve anybody before the provider exists and
cannot grant itself a client. If it authenticates natively, the provider is an
ordinary hosted service. So the count follows from a decision not yet taken, and
asserting four was asserting that decision.

The test also rules out the tempting wrong answer: an identity provider, a mail
server and an analytics service are all infrastructure by any ordinary reading,
and none are substrate, because the control plane starts and runs without them.
Important is not the test.

Records why the bundle is pinned by hand — it is applied when no mesh exists, so
nothing can resolve a version or ask a registry — and why it must be
self-contained, which makes it an artifact built on a machine with a network for
a machine that may have none.
2026-08-26 23:39:11 +02:00
jschoubben 148395ca54 Define the control plane, which was used 79 times and defined nowhere
Nineteen files, seventy-nine mentions, no definition. That is how-we-build §5
failing on this repository's own vocabulary — ubiquitous language is checked,
not assumed.

The definition, and it is not arbitrary: the control plane is everything that
needs to know about MORE THAN ONE NODE. It follows from ADR 0037, which has the
host applying rather than deciding precisely because deciding needs knowledge
the machine does not have. So the line falls exactly there — writing a file is
the host's, choosing which nodes run the store is the control plane's, and
anything a single machine could answer alone does not belong here at all.

That last consequence is worth having: putting a single-machine concern in tier
2 is a mistake the tier rule will NOT catch, because the dependency direction
stays correct.

Also states what it is not — not the thing that changes machines, not a surface,
not the substrate, and not privileged on a node beyond what the declaration
vocabulary allows. And the property that makes tier 2 unlike the others: it is
itself a consumer, with the same requirements as any module, which is the
circularity the bundle exists to resolve rather than hide.

Scoped deliberately: this defines the term and does not design the contexts
inside it. Ten is the skeleton's claim rather than a settled list, and research
006 still asks whether the record belongs here or in the substrate.
2026-08-26 23:33:20 +02:00
jschoubben a4ab3e15c2 011: one interface, many contexts — and the constraint that hides in it
The objection is right: if every context runs its own service with its own
interface, the board is coupled to N interfaces instead of N schemas, something
has to compose them, and composition is logic — which tier 3 says a surface does
not hold. That moves the problem up a layer rather than solving it.

The skeleton already answers it, and the previous entry talked past it. `work`
and `knowledge` are not separate services; they are contexts INSIDE the control
plane, alongside the record, inventory and delivery — and `api` is listed there
as the one interface every surface speaks to.

So the board speaks to one interface. Behind it the contexts stay separate,
integrating through the record, but they are one tier, one repository, one
deployable — and coupling within a tier is not what the tier rule forbids. The
problem does move up a layer, and the layer it moves to already exists and has
this as its job.

The caveat is load-bearing and now recorded as an open question: this holds only
while the contexts are not separate deployables. The moment one becomes its own
service with its own interface, the board is back to N clients, something must
compose them, and the composition has nowhere to live that tier 3 permits. That
is a real constraint on how far the control plane may be split, and it is worth
knowing before splitting rather than after.
2026-08-26 23:29:56 +02:00
jschoubben a4a25ca7e3 011: one surface over several contexts is normal
The board visualises the mesh, the work engine, the knowledge base and more, and
the alternative — a web application per context — is worse for everyone using
it. Composing several sources into one view is what a surface IS, so this is not
a compromise with the ownership rule.

What changes is only where it reads from: each context's interface rather than
each context's store. Most of that already exists — 56 of 126 modules carry a
tool surface, more than carry a service.

And the unified board is what keeps those interfaces honest. A view that cannot
be built from a context's interface proves the interface inadequate, discovered
where it is cheap to notice rather than the first time something else needs the
same data and quietly reaches for the store instead.

If composing many calls proves too slow, the answer is a projection the board
owns and keeps current from events, not access to somebody else's tables.
2026-08-26 23:27:42 +02:00
jschoubben fa62c7f0e4 011: correct the rule — contexts, not processes
An earlier version argued a dashboard reading a dozen stores was caught by
exclusive ownership, because a dashboard is a surface and surfaces speak to an
interface. Wrong, and it drew the line in the wrong place.

The mesh's own board showing nodes, modules and deployments is not a separate
context reaching across a boundary — it is the mesh showing its own data.
Requiring it to go through an interface to reach facts its own context owns is
ceremony.

The rule is that a CONTEXT is granted what it exclusively owns. Everything
inside it — service, surface, tools — reads that store freely. What is forbidden
is a different context reading it.

Which is what the consumer count already showed: the problem was never surfaces,
it was three other contexts keeping their tables in the mesh's database.
2026-08-26 23:25:46 +02:00
jschoubben e71d532c2e 011: request or subscription is derived, not chosen
Asked what the distinction actually is, and the SQL half needed correcting
first: under exclusive ownership SQL runs against your own database and nothing
else, whatever transport a query might travel over. Both options are the mesh's
own channel and both ride the broker, so the transport is not the distinction.

The distinction is where the answer lives when you need it. A request asks at
the moment and waits — always current, costs a round trip, cannot answer when
the other side is down. A subscription keeps a local copy — instant, works
offline, as current as the last event received, and you must handle what you
missed.

What decides is not taste. ADR 0036 makes disconnection an ordinary situation
rather than an exception, so anything that must keep working while disconnected
CANNOT use a request: there is nobody to ask. And the converse — anything where
a stale answer is worse than no answer cannot use a subscription. A display can
lag; a decision about whether a grant is still valid cannot.

So an apparently open question turns out to be derived from a decision already
taken. What stays open is narrower: what a consumer does about the events it
missed while disconnected — replay from a point, ask once for a full picture and
resume, or rebuild. The question every projection has.

Also recorded: separate databases are required in the new design, and the shared
registry is a leftover rather than a pattern.
2026-08-26 23:23:39 +02:00
jschoubben 7e83723b7b 011: rewrite the question table, which had gone stale silently
Several edits to the overview matched nothing and returned success, so the
question table still carried answers superseded two or three exchanges ago —
"when two modules provide one name, who chooses" was still open in the table
while answered in the file it pointed at, and nothing recorded the instantiation
edge, instance counts, grants, bootstrap provisioning, the tool audience, or the
registry consumer check.

That is the fault this repository catalogues, committed by the thing cataloguing
it: a string replacement that found no match, reported nothing, and left the
document claiming a state it did not have. Rewritten from what the documents
actually say rather than patched again.

Nine questions settled, fourteen live, and the split is now visible instead of
implied.
2026-08-26 23:17:25 +02:00
jschoubben afcc355744 011: checked the registry's real consumers, and the question was the wrong shape
The exclusive-ownership rule turned on whether every reader of the mesh registry
could be served another way. Eighteen consumers open a direct connection. Four
groups, and only one is work.

The owner and its machinery keep reading, because they own it. The node appliers
are already resolved — ADR 0037 stops the host querying the mesh database,
decided for tier reasons with nothing to do with this.

The bulk are FOREIGN TENANTS. The work engine holds ten of its own tables in the
registry's database, the knowledge base two, pipeline logs one. Thirteen foreign
tables across three contexts, which is how-we-build §4's shared schema counted.

So the question was the wrong shape: the problem is not readers needing a new
route to data, it is tenants needing to move out. Tasks, agents and teams have
nothing to do with nodes and modules and are co-located by history. Give that
context its own database and its dependency on the registry shrinks to one
table.

A handful of genuine cross-context reads remain, small enough to enumerate
rather than estimate. The rule holds.

Left open: whether those reads want an interface or events. Asking which nodes
exist at the moment you need to know is a request; reacting when a node appears
is a subscription, and some consumers want both.
2026-08-26 23:16:13 +02:00
jschoubben 6b1aab6a1e 011: the dashboard case, and why exclusive ownership is the tier rule
Raised as the hardest test of the rule: a board showing nodes, modules,
pipelines, agents and tasks wants to read a dozen stores, and under exclusive
ownership it can read none of them.

It survives, and not by luck. The board is a SURFACE, and surfaces already may
not do this — the skeleton puts `api/` in the control plane as the one interface
every surface speaks to, and tier 3 as thin, no logic. A board reading stores
directly is a surface reaching past the context that owns the data, which the
tier rule forbids for reasons that have nothing to do with databases.

So it is not a counter-example; it is an instance the rule catches. And the two
rules turn out to be one rule seen from two sides: exclusive ownership is the
tier rule expressed in terms of storage.

The general shape for anything needing to see across many things: consume the
record and own your own view. A reporting context builds a projection from
events and reads its own store, never anybody else's.

The cost said plainly rather than buried: a projection is more work than a join,
and it lags. A board queries the mesh's own database directly today — ordinary,
working — and this rule makes that a migration rather than a preference. The
reason to pay it is §4's already-measured cost, not elegance.
2026-08-26 23:10:36 +02:00
jschoubben aa767d17a8 011: a module is granted only what it exclusively owns
Reconsidered by the operator — maybe shared databases should not be allowed at
all — and the stricter version is better and goes further than the schemas it
replaces.

No shared writes, and no read-only role on another module's database either.
Reading another context's tables couples you to its layout exactly as firmly as
writing them does, and the coupling is harder to see because nothing breaks
until the owner changes a column.

That is how-we-build §4 taken at its word rather than at its letter. The
permissive version — a per-consumer schema, revocable, with cross-context joins
possible but deliberate — kept the letter and left the temptation. A boundary
that is merely inconvenient to cross is a boundary that gets crossed.

The cost is cross-module reporting, and it is the point rather than a
regrettable side effect: anything wanting to know what several modules hold
consumes their events or calls their interface. That is §4's whole argument, and
the mesh already has both mechanisms. What gets harder is precisely the thing
that was making work belonging to one context keep having to be implemented in
another.

And it is the first clear instance of what this effort has been hunting — what
the design DELETES rather than adds. Grant kinds collapse to one: an exclusive
resource. With them go the question of who owns which table, the guessing at
revocation time, cross-module migration ordering, and a class of permission
modelling a shared store would otherwise need.

One thing it does not answer, recorded because it could make the rule
unworkable: the mesh's own registry is read directly by many things today, and
under this rule they consume events or call tools instead. Achievable in
principle. Whether EVERY current consumer can be served that way is unchecked,
and should be before this becomes a decision.
2026-08-26 23:09:44 +02:00
jschoubben 4c8515507a 011: where a binding lives follows the scope, and a grant is not always a whole resource
Two questions asked directly, and the second collides with a rule in force.

One module on two nodes sharing a database corrects something stated flatly: the
binding is not "recorded on the assignment". Where it is written down FOLLOWS
THE SCOPE. A shared grant belongs to the module and every assignment references
the same one — which is the answer for two nodes wanting one database between
them. A per-instance grant belongs to the assignment. Same relation, two homes,
and which home is what makes two instances share something or not.

Several modules adding their own tables to one database is three needs wearing
one sentence, and a provider offers KINDS of grant rather than one: a database
for a consumer whose tables are nobody else's business, a read-only role for one
that needs to see what another holds, and a SCHEMA within a shared database for
the case actually asked about.

Loose tables in a shared database is what how-we-build §4 warns against in as
many words — several domains sharing one forty-five-table schema, which is why
work belonging to one context keeps having to be implemented in another. Not a
style objection; the observed cost, already paid.

A per-consumer schema keeps what the request wants and drops what §4 objects to.
Same database, same connection, same backup, and a cross-schema read remains
physically possible when genuinely needed. What it adds is ownership: migrations
touch one namespace, two modules cannot collide over a table name, and revoking
drops the schema rather than guessing which tables belonged to whom.

So the fault §4 names is still possible and no longer accidental — a
cross-context join becomes something somebody deliberately writes rather than
the path of least resistance. And revocation becomes answerable, which the
whole-database version never was.
2026-08-26 23:08:32 +02:00
jschoubben fa9889536c 011: tools have a different audience, migrations cross the edge, provisioning is early
Three additions, and the third kills an assumption.

Tools are the most common content in the catalogue — 56 of 126 modules, more
than carry a service — and they survive the split without fitting either half. A
tool is not an artifact and not node state; it is a contract the mesh publishes
on a module's behalf, and what consumes it is an AGENT rather than another
module. That is a second audience the design has not described. Whether it is
one relation with two audiences or two relations is cheap to decide now and
expensive later.

A migration belongs to the CONSUMER and runs on the PROVIDER. A game's
migrations run against the database the store granted it: owned by the consumer,
hosted inside something it does not control, ordered after the provisioning edge
because there is nothing to migrate until the grant exists, and scoped to that
grant. Ownership crosses the edge, which nothing in provides and requires
expresses — and it gives a consumer's own install an internal order, provisioned
then migrated then started, that depends on an edge rather than on its contents.

And provisioning is EARLY, not late. The assumption worth killing is that it is
something the control plane does for consumers once a mesh is running. The
mesh's own registry database is provisioned before there is a mesh, and so is
its virtual host on the broker: the store runs from the carried bundle, a
database is created in it, the mesh's own schema is applied, and only then does
a control plane exist. Steps two and three happen before there is a mesh to do
them, so provisioning is part of the bootstrap and part of what the bundle has
to express.

Which strains ADR 0043. The host applies declared state ON THIS MACHINE, and a
database inside a running store is not a file or a unit. At bootstrap it is at
least local — the store is on the same machine. Afterwards a consumer on one
node provisioned from a store on another is the ordinary case and reaching it is
not the host's job. The same operation is local at bootstrap and remote later,
which is either two mechanisms or one with a tier boundary crossing inside it.
Currently the sharpest unresolved thing in the effort.
2026-08-26 23:02:52 +02:00
jschoubben a3c7e7e1f1 011: providing is a facet, and the assignment is a third thing
Any hosted service can be a factory — an identity provider grants clients, an
analytics service grants a tracking identity, a mail server grants mailboxes, an
application platform grants a project that is several of those at once.
Providing is a FACET a module may have, not a kind of module it is, which is the
same conclusion this effort reached about services and applications arriving
from the other direction. So `provider` stops being a category too.

Two relational stores from different vendors both grant "a database" and are the
sharpest possible test of the substitutability rule. They fail it completely —
different protocol, dialect, driver, client library compiled into the consumer —
so `database` stays a tag, now with two real providers rather than a thought
experiment.

The assignment is a third entity, recorded because the operator tried the
alternative: modules were once node-agnostic and it did not survive. Several of
a provider's properties belong to neither end — where its state lives, how it is
reached, tuning derived from the machine's hardware, which instance serves a
given consumer. Not the catalogue, because they differ per node; not the node,
because they are about this module. A design with only modules and nodes has
nowhere to put them, which is what node-agnostic ran out of. The current system
already stores environment values per module AND per node, arriving the same
way.

Which answers the question asked directly: two nodes both run a store, so which
serves a consumer? Neither obvious answer. Not the consumer naming a node — that
is placement in the consumer's manifest, a game edited because a database moved.
Not the consumer not caring — for presence it genuinely does not, for
instantiation it cares permanently.

What the consumer knows is the SCOPE of its own need: one instance shared across
every instance of itself, or one each. That decides, and needs no node named.
Then the mesh binds, and the binding is recorded on the assignment and is
sticky — a resolver that re-derives which store serves a consumer will one day
derive a different answer and relocate a database.
2026-08-26 22:56:36 +02:00
jschoubben 13c6068874 011: the provider shape generalises, and two things differ inside it
The broker has all nine properties the store has. So do the object store and the
image registry. A substrate service is a SERVICE PLUS A FACTORY, there are four
of them, and the pattern generalises past the substrate: anything granting
something per consumer has this shape.

Two differences matter more than the similarity.

The broker cannot be managed over the broker. ADR 0001 makes it the channel
every node takes work from and ADR 0039 makes it the security boundary, so the
module providing it is also the way modules are managed — a declaration cannot
be delivered to it over itself. Nothing else has that property; the store is
consumed by the control plane but is not how the control plane REACHES anything.
This is what the carried bundle exists for: the broker is raised from what the
host carries because there is no other way to raise it. A constraint on one
module, not a general rule, and a schema with no way to say so hides it.

And two modules of identical shape want opposite instance counts. The broker is
one per mesh by decision. The store cannot be, because a node that must keep
working while disconnected cannot depend on a database elsewhere. Which settles
what cases.md left open: how many instances is NOT derivable from what a module
is. It is a per-module decision, it has to be declared, and nothing in provides,
requires or excludes says it.

Revocation differs in consequence too. Dropping a database leaves data until
something removes it — a leak, recoverable. Dropping a virtual host loses
whatever was undelivered — silent, and not. Same relation, different blast
radius, which argues for the provider deciding what revocation means rather than
the mesh applying one rule.

File renamed: it was never really about postgres.
2026-08-26 22:54:49 +02:00
jschoubben f160b28a71 011: postgres worked through, and "one kind of edge" was wrong
The tidy version said a module provides names and requires names and that is the
only edge. Working postgres through completely disproves it.

A small game wanting to store data does not require postgres to EXIST. It
requires postgres to MAKE IT A DATABASE and hand back credentials. Those are
different relations in every way that matters: one creates something per
consumer, carries a payload back, can be revoked, and leaves the provider
holding state about who was granted what. The other creates nothing.

So: two kinds of edge, one graph. Instantiation implies presence; presence does
not imply instantiation. The current system already had exactly this split —
`dependencies` for presence, `requires: provision:` for instantiation, with the
resolver deriving one from the other. analysis.md called that derivation a
convenience. It is not: it is the correct relationship between two genuinely
different relations, and the design had collapsed them.

Postgres also turns out to be nine things, not one. A container. Persistent
state where moving nodes is a migration rather than a reschedule. Configuration
partly derived from the machine's hardware. A tool surface. A provisioner. Its
own bookkeeping about what it granted, which is not the data it stores. An
exposure decision per node it runs on. Credentials it generates, which means a
provisioning edge carries a secret. And health that is not "the container is up".

Four questions the worked example makes concrete rather than abstract. WHICH
postgres, when there are two — a consumer of `terminal` does not care and a
consumer of a database cares permanently. How many instances a module should
have, which cannot be a global rule because one-per-mesh is wrong for a store a
disconnected node needs and one-per-node is wrong for the mesh's own registry.
What happens to a grant when its consumer is removed, where dropping is data
loss and keeping is a leak. And whether a declaration is composed PER NODE from
what that node reported — because tuning follows hardware the control plane
cannot know, and the alternative is the host deciding, which ADR 0037 forbids.
2026-08-26 22:53:59 +02:00
jschoubben c9c2dfe686 011: what a feature is, and what it splits into
The operator wants features gone, and 006 left it open. Measured, and the answer
is that nothing replaces them because they were never one concept.

A feature is a kind of content a module carries, detected from its directory:
twenty-one of them, each with a handler owning six stages — build, publish,
install, configure, start, verify.

The structural finding: EVERY handler implements EVERY stage. `configs` writes
files onto a node, has nothing to build, and has a build stage. `npm` publishes
to a registry, has nothing to start, and has a start stage. One interface spans
build-time and apply-time, so every kind of content must implement both halves
and most do nothing in one — and a stage that does nothing looks exactly like a
stage that failed to do anything.

They split four ways, across three tiers. Artifacts built once per version and
published, where no node is involved — delivery. Resources that are desired
state on a machine, which is what ADR 0043 already describes and the host already
does — tier 0. Actions run once against something that is not this machine, like
a migration against a database on another node — delivery, and seeds go
entirely. And checks: the prerequisites are REQUIREMENTS IN DISGUISE, a module
saying what must be true before it can be installed, which is what an edge in
the graph says; the verifiers are the read-back the host already performs.

So `feature` is one word for four things spanning three tiers, which is why the
pipeline is hard to reason about.

One property must survive the split, and it is the thing the current design got
right: content is DETECTED, relationships are DECLARED. A module that says it
has migrations and has none is a fault nobody sees until it matters — but what
it requires and provides is not visible in a directory and has to be said.
2026-08-26 22:51:19 +02:00
jschoubben ae099482a9 011: twenty cases, and two axes nothing covers
Before settling a schema, what a module can actually be. Twenty kinds of thing,
with the hard ones at the end because they are the point.

The ordinary nine are unsurprising: a supervised service, a system package with
configuration, an application a person launches, a command-line tool, a library
that never runs, a one-shot task, a scheduled one, an adapter, and a standalone
application whose only difference is where its source lives.

The eleven that break a naive schema are where the work is. Something that is a
service AND an application — a git forge is consumed as a remote and operated
through a web interface, and neither reading is wrong. Something that provides
and consumes, because provider and consumer are ends of edges rather than kinds
of module. Something the mesh installs that then becomes a node CAPABILITY,
which means a node's provides-list is partly derived from what is installed on
it and not only detected. Something that must be adopted rather than installed.
Something that is a set rather than a thing. Something with exactly one instance
for the whole mesh, where assigning it twice is not redundancy but two meshes.
Something that is not software at all — a firewall policy, a DNS record, pure
desired state, which fits the host's declaration model exactly and an installable
package model not at all. An agent. The host itself, which is not a module and
needs a schema that can say so. And the things the mesh depends on and does not
control, which are why a node can be perfectly configured and still not work.

Nine axes come out of it. Two are covered by nothing anyone has proposed: HOW
MANY INSTANCES a thing may have, and WHETHER TWO CAN COEXIST — `excludes` covers
part of the second and nothing covers the first.

And one question the cases sharpen: is "runs" a property or a kind? The axes say
property — one schema with a field saying how it runs, `never` included. The
alternative is several kinds of module with different schemas, which is the
taxonomy this effort already rejected once for services and applications.
2026-08-26 22:48:20 +02:00
jschoubben f55ecc1a47 011: an abstract name needs providers that are actually substitutable
Two corrections from the operator, and the first improves the design rather than
narrowing it.

`database` is not an edge. The test it fails, and the test the proposal was
missing: can a consumer be switched from one provider to another WITHOUT
CHANGING? A module speaking Postgres does not speak MongoDB or SQL Server —
different wire protocol, dialect, driver — so a consumer declaring `requires:
database` and handed any of them breaks. The name promises what no provider can
deliver, and the resolver would report a requirement satisfied that is not.

`terminal` passes: anything that runs a command in a terminal works and the
consumer never learns which it got.

So the ADAPTER is what creates an interface. `ai-assistant` is legitimate exactly
because adapters normalise what is behind it. Without one there is no interface,
there is a category — and a category is a TAG. Tags describe, edges bind, and
keeping them apart is what stops the catalogue acquiring a second kind of
relationship that looks like a dependency and is not, which is what a folder
named after a domain already was.

And the domain module goes. A `networking` module gathering a firewall, a
resolver and a proxy under one name came from an older shape and does not fit —
there is no such thing to install. There is core infrastructure: concrete
modules named individually, not flavourable, with no grouping module standing in
front of them.

Fixed three places where the revision left the old rule standing, including an
example manifest still requiring `database` — the kind of contradiction that
would have been read as the design rather than as a leftover.
2026-08-26 22:44:37 +02:00
jschoubben e20a09ae80 011: one kind of edge
The design, rather than an account of what exists. A module provides names and
requires names, and that single relation absorbs three things this effort had
listed separately: requiring another module is requiring a concrete name,
requiring a resource is requiring an abstract one, and an interface is simply a
name with more than one provider. Nothing has to declare that it is an
interface — it either has one provider or several.

The move that does the most work: a NODE provides names too. Its profile is a
set of them — display-server, container-runtime, an architecture — so a module
requiring a display server is satisfied by the node exactly as one requiring a
database is satisfied by another module. One resolution instead of two, and a
graphical application cannot land on a node without a display server for the
same reason, through the same code, that it cannot land without its libraries.

Which makes the host's capability detection an input to resolution rather than
something a person reads. It was built to be read; it turns out to be a
provides-list.

`excludes` is the one genuinely new relation, because it is not derivable: two
modules that both provide message-bus look interchangeable when installing both
would break the machine.

Constraints are not placement. They say what must be true of a node, never which
node — which is the mistake the measurement found in the current catalogue,
where a module pins its database to a named node so a second node cannot provide
it without editing the consumer.

What it deletes, for the design: the module/resource distinction, the interface
as a kind of thing, capability checking as a separate mechanism, domain grouping
— folders assert relationships where edges record them, so a domain becomes a
query over the graph rather than a directory somebody keeps true — and possibly
tiers, if a tier is just a computed level.

What it does not delete, stated so it is not discovered later: a resolver still
has to exist, with version constraints and conflicts, and the design owes an
answer on what it delegates rather than reimplements.
2026-08-26 22:39:17 +02:00
jschoubben c3a2984b3e 011: measured, and the premise was wrong — the graph is not missing
The effort was opened to ask whether the catalogue's missing structure is a
graph. It is not missing. 126 manifests, 103 edges, no cycles, nothing dangling,
deepest chain of five — and a resolver in the SDK that topologically sorts them,
already called by the tool loader at startup, the installer when syncing modules
onto a node, and the delivery coordinator when expanding what a change affects.

It already does something this effort assumed would need designing: a
requirement on another module's provision is treated as an implicit edge to the
module that provides it. So "ordering by the graph", which ADR 0043 makes the
control plane's job, is a thing to call rather than a thing to build.

The one place the graph is wrong, it is wrong about the substrate. A module
needing a database declares `provider: postgres` inside `provisions:` — which is
what a module OFFERS — so the resolver, which reads `dependencies:` and
`requires:`, never sees it. Three edges are invisible this way, and they are the
mesh's own database, the mesh's own broker, and the work engine's database.

The consequence is measurable: computing what a working mesh needs from the
declared graph gives registry -> sdk -> mesh -> meshware. Four modules, four
levels, no database. Arithmetically correct and obviously wrong, for exactly one
reason — a field that means "depends on" is not read as one. That is
04-ISSUES/003 in a new form: not a key nothing reads, but a key read as
something other than what it means.

Two latent defects, both contrary to ADR 0008 and both in the component ADR 0043
makes responsible for ordering a host will apply without question: a cycle warns
and falls back to input order, and a dependency that does not exist warns and
continues. Neither has fired, because the catalogue currently has no cycles and
nothing dangling, which is why nobody has noticed.

And placement is decided in the catalogue: a provision pins itself to a named
node in the manifest. Which node runs what is an inventory decision — tier 2 by
the skeleton's own test — so a second node cannot provide the mesh's database
without editing the module that consumes it.

What the graph would DELETE is currently nothing. What it would add is three
declarations that no manifest uses today: excludes, a required node capability,
and an interface with adapters. Whether they would be used is not measured, and
zero usage is equally consistent with nobody needing them and nobody being able
to express them.
2026-08-26 22:35:39 +02:00
jschoubben 278f7427ed 012: the briefing carries an outcome, derived from its lines
Proposed by the operator: state plainly whether adoption succeeded, partly
succeeded or failed, with a severity per line.

Taken with one change — the overall is DERIVED as the worst mark present, never
written alongside. Two fields maintained independently drift, and a briefing
reading "full success" while carrying a failed line is exactly the fault this
record keeps cataloguing. An outcome computed from its lines cannot disagree
with them.

Four marks: ok, kept, unknown, failed. "unknown" is not a shade of success —
adoption will meet configuration it cannot parse and state it cannot read, and
folding those into "fine" is the same move as reporting an installed package as
a capability.

And adding severity reopens something the earlier rule did not cover. "Flags
inform, they do not block" was decided about CONFLICTS, where the mesh chose
deliberately and the machine still works. A failure is not "we chose" but "we
could not". Treating both the same makes a node where something the mesh needed
never happened indistinguishable from one where a log level differed.
2026-08-26 21:55:08 +02:00
jschoubben 0106318bcb 012: on conflict, keep the machine's configuration
Reversed by the operator, and both directions are recorded because the reasoning
for each is the useful part.

What is already on the machine stays, the conflict is flagged, adoption
completes. This buys non-destructiveness by construction: the class that made
the opposite rule dangerous — a storage driver against the filesystem it is
actually on, a data directory pointing at a mount that exists — cannot arise,
because nothing tied to the machine's physical reality is overwritten.

It exposes the mirror. The mesh's configuration is not only preference; some of
it is what a module needs to function. Keeping the machine's version there
produces a module that is installed and does not work, which is 04-ISSUES/007
arriving from a direction that issue did not anticipate. And a fleet where every
node kept its own settings is one where a module works on one node and fails on
another with nothing able to say why.

So neither direction is right as a blanket, and the question is not whose
configuration wins. It is whether the module REQUIRES the setting or merely
PREFERS it — required contradictions cannot be kept without breaking the module,
preferences should always yield to what is there.

That is a property of the module's declaration rather than of the adoption
algorithm, which makes it one more thing the graph would carry. Until modules
can say which of their settings are load-bearing, adoption is defaulting in the
dark, and the default chosen is the one that does not break the machine it is
adopting.
2026-08-26 21:40:25 +02:00
jschoubben bcb18c7329 012: on conflict, install the mesh's version
Decided by the operator. Where the existing configuration and the mesh's
disagree, the mesh's version is installed, the conflict is flagged, and it is
reconciled afterwards — the mesh's configuration is known to work, the machine's
is not, and a half-adopted machine is a state nobody understands.

So adoption always completes and flags inform rather than block, which also
settles what 'adopted with open questions' prevents: nothing. The node is a
node. The original is kept, so nothing is unrecoverable.

One class left open rather than folded in, because it is the one place the
oldest rule in this record argues the other way. 'Known to work' is true of the
mesh's configuration in isolation, not on this machine. Most disagreements are
preference and overwriting them is right. A few are tied to what is physically
present — a storage driver against the filesystem it is actually on, a data
directory pointing at a mount that exists — and installing ours there does not
discard a preference, it can make existing data unreadable. Restoring the
configuration file afterwards does not undo that.

The default is settled. The exception is not 'there is a conflict' but 'applying
ours would destroy something a configuration backup cannot restore', and
identifying that class is open.
2026-08-26 21:37:04 +02:00
jschoubben 60736199a4 012: keep the original, and flag what cannot be decided
Two additions from the operator, and the second answers a question this effort
had open with two bad answers.

Nothing is taken over without keeping what was there. Adoption happens on
machines somebody is already using, and the configuration being taken over is
configuration somebody chose. This is a never rule rather than a courtesy, and
it earns that by the same incident the mesh's strongest rule carries: the worst
loss in this record came from a tool acting on a path it did not own. Adoption
is that act made deliberate, which makes the safeguard obligatory.

And adoption produces a briefing, not just a result. It meets things a script
cannot decide — a runtime configured one way against a mesh wanting another, a
package pinned for a reason, local settings the mesh has no opinion about.
Silently winning is wrong in both directions and refusing outright makes a
machine in use unadoptable. So conflicts are FLAGGED: what it found, what it
took over, what it could not resolve, written to be read by a person or an agent
as the first thing a session on that node has to work with.

That is the declaration parser's principle at a larger scale — name every
problem at once, to somebody who can act on it.

The question it turns on is recorded rather than assumed away: are flags
advisory or blocking? A briefing nobody opens is worse than a failure, because
the machine is in service and the record says it went well — 04-ISSUES/003
again. Working position: the node is usable and the mesh KNOWS it has unresolved
adoption questions, as a state something can ask about rather than a document in
a log directory. What that state prevents is undecided.
2026-08-26 21:33:09 +02:00
jschoubben ddb8091f68 Research 012 — the minimum viable node, and adopting what is already there
Building tier 0 reached a wall that looked like a packaging problem and is not.
The host can be told to run a container or install a package; both need a file,
and asking where the host gets it produced a bad trilemma — carry everything,
download at apply time, or push the files in first. Downloading fails on the
first node, which cannot fetch the image registry from the image registry it is
trying to start.

The reframing came from the operator: the machine is not offline, and what
matters is WHEN the fetching happens. Move it from apply time to build time —
build the installer on a machine with a network, tailored to the target, apply
it on a target that then needs nothing. The same move the lab already made for
its router image.

Which makes the question not where artifacts come from but what is missing from
THIS machine, and that needs two things answered: the closure for a one-node
mesh, and how a machine already in use becomes one.

Adoption is the second half, and it is sharper than it sounds. Having a package
installed is not owning it: a container runtime found already present carries
settings somebody chose, and noticing the binary exists discovers none of them.

It was also the original path — 00-as-is/05 records adoption of a pre-existing
machine's configuration as the original mechanism, since made legacy and
explicitly out of scope for the lab. It returns for a different reason than it
was dropped for.

Two collisions recorded rather than discovered later. ADR 0004 has managed files
generated and never edited, and adoption needs a one-time import before that
rule starts applying — three states, and the middle one is new. And ADR 0043
says the host never touches what it did not create, which is exactly what
adoption does; that rule needs a companion rather than an exception.

Eight open questions, including whether 'tier' is just a coarse view of a graph
level, whether owning a package means owning its version, and what cannot be
precomputed at all — because tailoring moves the cost of building from source
rather than removing it.
2026-08-26 21:30:43 +02:00