Commit Graph
45 Commits
Author SHA1 Message Date
jschoubben 0bcfeb4e80 028 fixed — the mesh assigns the port, and knows what it cannot move
A module now says its port once, in `listens`, and the container's
mapping, the rule set and what a consumer is told are all derived from
one assignment. The three hand-written copies that agreed only because
one person wrote them are gone.

The half that made this an issue rather than an inconvenience was that
the substrate is not a module: nothing in the mesh had heard of its own
store, so it handed a database module the port the store already had.
The machine now says what it carries, and the mesh assigns around it.

Ports the protocol fixes became claims, which needed no new mechanism —
the mesh already had one for what is singular on a machine.

Left open, and unchanged by any of this: whether a module should publish
to the machine at all. Assignment makes publishing safe without making
it necessary.
2026-09-01 18:32:58 +02:00
jschoubben c2a37ab7f8 38 — the mesh assigns the port, and a module does not care
Jochen's call, and the right one: a module cannot choose a port well,
because it is written once and assigned anywhere. Any number it picks is
a guess about a machine it has never seen, and two modules guessing the
same number is not a mistake either of them made.

Writing it up turned up something the issue had missed. The same number
appears three times in every module — the rule set, what a consumer is
told, and what the runtime publishes — and nothing checks that they
agree. They agree today because one person wrote all three. A module
whose `serves` said one thing and whose container published another
would resolve, compose, apply, and hand every consumer a port that
answers nothing.

So the decision is one source with the other two derived, and an
assignment made once and kept, as a credential is.

The part that needed thought is ports that cannot move — mail on 25,
submission on 587. Those become claims, which is what the mesh already
has for what is singular on a machine. Two modules wanting 25 is the
same shape as two wanting the seat, and gets refused by name at
assignment rather than by a container runtime at apply. That makes this
mostly a matter of pointing an existing mechanism at ports.

Left open: whether a module should publish to the machine at all.
Assignment makes publishing safe without making it necessary.
2026-09-01 17:42:27 +02:00
jschoubben 6330abce5f 028 — two things want one port, and nothing says so until the machine
Found by fixing 027 and pushing again. The declaration is now accepted
and the database container still cannot start: the mesh's own store
holds 5432 on that machine, and the module publishes 5432.

Nothing catches it because the substrate is not a module. It arrives
from the bundle before there is a mesh to ask, so the control plane has
never heard of the store and does not know it holds a port. Resolution
can compare modules with each other and cannot compare one against what
the mesh is built on.

Nor does it compare modules with each other. A port is exclusive on a
machine in exactly the way a claim is, and the mesh has a mechanism for
that which ports do not use.

It has been met before: the end-to-end test that exercises a real
database publishes 5433 rather than 5432, inline, with nothing saying
why. That is how a constraint becomes folklore.

The open question is bigger than the bug. Whether a module should
publish to the machine at all decides how a consumer reaches it, and
changes what `serves` means.
2026-09-01 17:21:31 +02:00
jschoubben 03266fd4a2 027 — a container cannot follow a file, and a rotated credential is the case
Found by the forge failing to start. I had put `restart-on` on nine
containers so they would pick up a rotated credential; it belongs to a
service, and the host refused the whole declaration.

Removing it fixes the modules and leaves the reason I reached for it.
The mechanism is written against exactly this, in the host's own words:
a running service does not re-read its configuration, so replace the
file, find it already running, do nothing, and the machine keeps
behaving as before while every check passes. Every word of that applies
to a container, and nearly everything the mesh runs is one.

The cost is concrete. Rotation replaces the file and tells the provider
to accept the new credential. A provider reconciles, so it takes it. A
consumer is usually a container, so it does not — and the two ends hold
different passwords, which is the fault ADR 0001 records costing two
days. The test that proves rotation works uses a consumer that reads the
file on each attempt, so it does not meet this.

Two things a fix has to keep: it stays declared state rather than a
command, because the link may not carry an action; and where an env-file
changed, the honest verb is recreate rather than restart, because a
container's environment is fixed at creation.
2026-09-01 17:19:43 +02:00
jschoubben 3f00d3b413 026 filed; 025 corrected; the image store is a module
Three corrections, two of them to things I wrote today.

026 is the serious one. Four modules mounted fourteen host paths nothing
declared — the mail spool, the databases, the object store's data. The
runtime creates those as root, so owner and mode go unapplied, and the
rule that keeps a directory holding data the mesh did not put there is
written in terms of declared directories. It reached the configuration
and missed the data. The cause was carrying compose files across: a
container shape that can express one gets filled in like one.

025 claimed nothing turns a tag into a digest. That is false, and the
answer was designed and built before I wrote it. A module names an
artifact, not an image, and `kind: upstream` mirrors somebody else's
image into the mesh's own registry, pinned by the digest it lands with.
The two-document split the issue described as the shape of a fix is the
design. Pinning twelve images by hand was treating the symptom, and left
them pointing at a public registry rather than the mesh's.

And the image store was written up as something the mesh does. It is an
ordinary module — considered for the substrate and removed, because the
test is whether the control plane needs it before its first instruction,
not whether it can grant itself one. So somebody's own registry is the
same module as the mesh's.
2026-09-01 16:10:13 +02:00
jschoubben 2bbdc52440 025 — half done: nothing unrunnable reaches a machine now
The refusal landed, and the examples pin images that exist. What is
still missing is the part that makes it unnecessary: nothing in the mesh
turns a tag into a digest, so it was done by hand — which is precisely
what the issue says a person should not be asked to do.

Recorded because the resolution mechanism turned out to be trivial:
asking a registry what a tag points at takes about a second and pulls
nothing. That removes the main argument for leaving this open.

Also records the two faults that fell out of pinning for real. The mail
system named seven repositories that do not exist, because it publishes
to a different registry than the manifest assumed, and one of the seven
had been renamed upstream. Nothing checking only the shape of a
reference could have found either.
2026-09-01 15:13:52 +02:00
jschoubben fc7717f6b1 021 closed — the record was left open after the fix landed
The code and its tests went in hours before the record was touched, so
an issue that read `located` had been fixed all along. That is the exact
failure the frontmatter exists to prevent: status is meant to be
answerable from the record rather than by reading the code.

Closed with the commit that did it, and cross-referenced to 022 and 023,
which came out of the same mistaken instinct — treating the machine as
a boundary, then as an identity, then finding a consumer had a password
and no name to present with it.
2026-09-01 15:04:36 +02:00
jschoubben f6b1834ea9 024 fixed — the registry was addressed by hand and nothing else was
The cause was one line. Machines get a systemd-networkd unit with a
static address, so networkd finishes and reports the link configured.
The registry ran `ip addr add` inline, which leaves networkd waiting to
configure something it was never told about — and
systemd-networkd-wait-online has an infinite timeout.

So network-online.target was never reached and everything ordered after
it never started. On these machines that is Docker, so `docker load`
blocked on a socket whose daemon was queued behind a target that would
never come, and three bounded timeouts stacked to thirty-five minutes.
These machines have no DHCP by design, so that wait was never going to
end.

The hypothesis in this record was wrong and the record now says so.
Stocking had just been changed, so stocking looked guilty; stocking
takes 34 seconds and always did, timed directly before changing
anything.

Fixed with two things that made it cost hours instead of minutes: an
image placement now waits for the runtime and refuses after 120s naming
what systemd is waiting on, and the end-to-end test passes onProgress —
the raise reported every step and the test discarded it, which is why
thirty-five minutes and four minutes of silence looked the same.

The suite then ran to completion, 23 of 24, the one failure a check of
its own flagging a path as a credential because `/` is in the base64
alphabet.

Also recorded: a redirected log lags, because Node block-buffers stdout
to a file. Read as a stall twice, the second time right after the real
fix — where a buffering artifact argues the fix did not work.
2026-09-01 09:53:08 +02:00
jschoubben 407416e6d0 024 — a run stalls before the host is placed, and says nothing while it does
Seen twice today. Once mid-run: thirteen passes, then the process ended
with no summary, no failure and no receipt. Once from the start: the
first test ran 35 minutes against a measured 4.5 and was still running
when it was stopped.

Ruled out rather than assumed: not memory (84 GiB free, no OOM), not the
daemon (the stalled machine answered `incus exec` immediately), and not
the changes under test — the anchor VM had no host log and no
containers, so the run never reached placing the host.

What changed just before is that the rebuild went from two artifacts to
six, and every one of them is pushed into the scenario's registry, which
is the step the second stall sat in. Recorded as what changed, not as
the diagnosis.

The reason this is an issue and not a slow test: the suite prints
nothing between starting a scenario and finishing its first test, so
four minutes and thirty-five look identical from outside, and the only
recourse is to guess. That is how a workstation was left unbootable in
August. And a run that ends silently after thirteen passes is a run
somebody may believe.
2026-09-01 03:20:40 +02:00
jschoubben f583502fc9 023 fixed — the mesh says who a consumer is, and what it is bound to
Both halves had one cause: the mesh knew something and did not say it.

Who a consumer is now comes from one derivation, sent to the provider in
its grant and to the consumer in its binding, so the two agree by
construction. The provisioners use the name they are given and refuse to
invent one, because a name of their own would create a login the
consumer could never guess while everything reported success.

Bound values reach the file that needs them through the symmetric twin
of the sealed placeholder — simpler, because they are not secret, so the
control plane fills them in and the host gains nothing.

The lab run meant to prove this failed in a way that looked like the fix
being wrong: rotation could not authenticate against a real database.
The cause was the suite rebuilding the control plane's image and not the
provisioner's, so an image built that minute ran against a provisioner
built the day before. That is 005's family and is recorded with the
issue, because the misleading part is worth more than the fix.
2026-09-01 03:05:43 +02:00
jschoubben 80f18caf03 022 fixed; 023 filed — a password is not a connection
022 turned out to have a silent half worth recording: the provider
refuses loudly and names the modules, which reads as a decision, while
the consuming node does not refuse at all. Three modules wanting one
database produce one need, so two of them get no credential file and
each starts and fails to authenticate with nothing saying why.

023 is what remained after fixing it. A consumer now gets its own
password, in whatever shape its configuration wants, and still cannot
connect: the user name is invented by the provisioner and recorded
nowhere in the mesh, and the host and port sit in a JSON binding that an
application reading KEY=value cannot use.

The asymmetry is backwards and the coverage document now says so. The
secret is the hard case, because the mesh must not be able to read it,
and the secret is the part that arrives. The host and port are ordinary
facts the mesh holds in the clear, and they are the ones stuck.

Keycloak, Gitea, Mailu and MinIO all parse and resolve and none of them
can start. This is what stands between the module set and a running one.
2026-09-01 02:40:40 +02:00
jschoubben c86adbe3cc 022 — a credential belongs to a node, so a second consumer refuses
Found while checking whether the module vocabulary covers real use
cases. A node running three modules that all want a database cannot be
planned at all:

  anchor has 3 modules asking for "postgres-database" and they would
  share one credential: gitea, keycloak, umami

The refusal is right about what it says and wrong about what it implies.
They would share one credential, and sharing is worse than refusing —
but the arrangement being refused is the ordinary one, and the node this
mesh exists to take over runs eight modules against one database server.

The cause is the key: a credential is keyed by provision, consumer node
and provider node, so `consumer` is a machine. The provisioner inherits
it and names the role `mesh_<node>`. The refusal is not a check that
caught something; it is the only honest thing that function can do with
a key that cannot tell two consumers apart.

It is the same mistake as 021 with a different face. There the machine
was treated as a trust boundary; here it is treated as an identity, as
though "who is asking" is answered by naming a host. Two modules on one
node are as separate as two on different nodes.

Worth stating plainly: without the refusal, gitea's login would have
opened keycloak's database, and nothing would have said so — from the
provisioner's side it created exactly what it was asked to create.

Not a local fix. It crosses the control plane, the grant file naming and
every provisioner that names something after a consumer.
2026-09-01 02:28:39 +02:00
jschoubben 36d342f176 What a module must be able to say, measured against 127 that exist
Every manifest in the system being replaced was read and every key
counted, then set against what the new one can express. Three findings
worth more than the table.

**The most-used key was already covered and I expected a gap.**
Depending on another module — 65 manifests, the commonest thing any of
them says — is a requirement naming a module, which already means that
module rather than anything providing the name.

**The largest real gap is tool servers: 56 modules, over half.** A
module can already run one; what is missing is anything saying it offers
tools. That is plausibly a provision rather than new vocabulary, which
would need nothing added — not yet decided, and recorded as undecided.

**The gap most worth closing is health, at seven modules.** The mesh
knows a container is running, which is not whether it answers, and this
project has paid for that distinction twice. An action with a verify is
exactly the right shape and may not arrive over the link, so a module
cannot declare one.

Two things are missing deliberately and say so: stage hooks, because the
link may not carry an action and a module needing setup ships a program;
and flavours, retired in favour of claims.

Config merging is missing and should stay missing. A mechanism that
understands TOML gets asked for YAML, then INI, which is how the thing
being replaced became unholdable.

Also records what the survey found that is not about coverage: manifests
that had stopped matching what was actually brokered, one fact derived
in two places giving two answers, and a live listing returning
credentials in plaintext.
2026-09-01 02:20:34 +02:00
jschoubben 122405df8e 020: the server version is not it either
Pinned 2.5.0 rather than latest, on the suspicion that its draft
profiles extension was involved. Identical failure, so that is ruled out
and recorded — two of the three guesses in this issue have now been
tested and both were wrong, which is the useful half.

The scenario keeps the pin regardless; it should have had one from the
start.
2026-08-31 21:42:13 +02:00
jschoubben acd5a0d80b File 020 — a certificate is issued and never collected; close Phase 1
Against a real ACME server the proxy orders, the challenge is answered
at the name on port 80 through the proxy itself, the authorisation goes
valid, finalisation is accepted, and the authority issues a certificate.
The client then posts to an empty URL to collect it, and never does.

Read from the authority's own log rather than inferred. Across one run
it issued two certificates and accepted finalise three times: the client
reaches issuance every attempt and fails at the same step after it.

Ruled out and recorded, so nobody repeats it: the directory is complete;
the authority's API certificate covers the address; the challenge path
works. A hand-written server config was suspected and was wrong —
replacing it with the server's own default, changing only the challenge
port, gives the identical error.

Filed rather than pursued because what remains is interop between two
libraries against a server that exists to be a test server, and may say
nothing about a real authority. What the mesh needed to show, it showed:
a routed name gets a certificate ordered from a configured authority,
and an unrouted one gets nothing — that second assertion passes.

Phase 1 closes with this one item partly open. Two of its four tasks
needed no code at all, the network shape was built, and the next thing
to learn comes from moving a module rather than a fourth lab run.
2026-08-31 21:40:03 +02:00
jschoubben e823cc1cc5 The design record is read where it is written, never copied to be found
Decides the question 006 narrowed to. An agent reads this repository
directly and the search consults it, so these documents surface beside
ordinary results instead of only when somebody already suspects they
exist.

A scheduled sync into the mesh's memory was the option that works with
what exists today, and lost on the ground this repository can least
afford: it makes a second copy, and the copy that is searched quietly
stops matching the copy that is edited. A design record that has
silently diverged from the reasoning it claims to carry is worse than
one that cannot be found — the first misleads, the second merely fails.

Amends what 0019 promised rather than satisfying it: these documents
will not be indexed, they will be read. The commitment that survives is
the one that mattered — that a searcher finds them without already
suspecting they exist.

Gated on an agent that does not exist yet, so 006 stays open on the
build with a decided shape. What closes it is a check that fails today
by design: search the mesh's memory for a phrase that appears only in a
design document here, and require it back.
2026-08-31 15:45:00 +02:00
jschoubben c192810fba 004 and 008 resolved; 006 narrowed to the decision it actually needs
**004 — certificate issuance.** The resolver declared no authority at
all, so the client fell to its built-in production default: there was no
setting set wrongly, there was no setting. It is now a node property
defaulting to staging, which answers the first open question. Staging by
default rather than production-with-an-override, because the alternative
leaves the safe path depending on remembering to opt out of it — 005's
lesson, in a second place. The rollout is ordered and the order is the
dangerous part; recorded, not performed.

**008 — node rescue.** Read back from running nodes as the report asked,
and one of its own claims was wrong in a way that matters: the health
timer does exist and does fire. It simply never calls the rescue script.
A trigger that exists and does not do what the script claims survives a
halfway check, which makes it worse than the absence the report
described. Resolved by making the documentation true, not by
implementing rescue — the replacement host already supervises recovery,
and wiring unattended restart into the fleet being retired is a
deliberate decision rather than a tidy-up. Two "self-healing" claims
narrowed to what they actually do.

**006 — deliberately not closed.** Re-checked today: the indexing still
does not exist. What is gone is the reason it was an issue — the claim
is no longer load-bearing, because the README names the gap and the
decision's reasoning never invoked indexing. A signpost now points here
from the knowledge base, and was measured rather than assumed: it is
reachable, it is not surfacing. Closing it while the indexing does not
exist would be this repository's own named failure, one folder from
where it names it.
2026-08-31 15:20:11 +02:00
jschoubben 345bbe0552 005 resolved: a suite that cannot run on every push says when it last ran
Retired in favour of the lab rather than repaired — that answers the
first open question. The second finding is the one that generalises:
"nothing runs it, and nothing reports that nothing runs it" is not a
fact about that harness, it is a fact about any suite too expensive to
run on every push. The replacement inherited the fault it was replacing.

Records the three rules that now hold, and what the fix taught twice:
the remedy rebuilt the symptom inside itself, and the code that counts
results passed every test while reading nothing.
2026-08-31 15:02:26 +02:00
jschoubben 6ecd03694b Issue 019: a comment asserting a fact about a machine, which nothing checked
Twice in one file, a statement about a machine that read as reasoned and was
wrong — and the module's unit tests all passed while the daemon could not
start. That is what a unit test is: it confirms the assertion was made, never
that it is true of any machine.

003 in prose rather than in a manifest key.
2026-08-31 14:11:34 +02:00
jschoubben aafeb5c9df Three issues resolved: one closed by evidence, two answered by the replacement
012 named its own closing condition — a scenario with four images coming up —
and the scenario now stocks seven and has raised cleanly many times at the
memory the wrong diagnosis had raised.

001 is answered by the host reading the package database back after installing.
002 was NOT answered and was present here too, so it is a fix rather than a
note: a stale index is now named instead of reported as a failed install.
2026-08-31 13:00:03 +02:00
jschoubben 8448219de1 Issue 018: a provider on the same machine was never announced to its consumer 2026-08-31 01:35:54 +02:00
jschoubben 1ade18209d Issue 017: an action succeeded into a state its own verify rejects 2026-08-31 01:01:17 +02:00
jschoubben a9cd3de5be Issue 016: anything after the declaration in a file was ignored 2026-08-31 00:51:46 +02:00
jschoubben 6e7e77acbe Issue 015: the harness read a swallowed answer as success 2026-08-31 00:48:16 +02:00
jschoubben a6872ac099 A key that is present and unusable, and what the certificate work became
Issue 014: the node's serving key was stored in the host's own encoding, so
every check that reads the file passed and no server could start. Same shape as
013 — two halves of one mechanism designed separately, each correct about its
own half. Where a file exists so a third party can read it, the format is the
interface.
2026-08-31 00:42:13 +02:00
jschoubben 778efaba8b The mesh runs its own registry, certifies its own names, and computes its own filtering
Issue 003 is answered in both halves: manifests are parsed strictly, and a
module says what it listens on and from where rather than carrying a key
nothing reads. The design records what was built and how each part is checked.

Issue 013 is new, found by reading while writing the first module that has
both a computed file and a service that needs it. The file arrived second.
It failed, then the next reconcile fixed it, which is why nothing caught it.
2026-08-31 00:37:34 +02:00
jschoubben ba14e2b629 Issue 012 — the first diagnosis was wrong, and that is the useful half
Two things changed at once: a fourth image in the scenario, and scenario
machines raised from 1 GiB to 2 GiB. The bootstrap then failed every
time, and the image was blamed.

Removing the image did not fix it. Removing the memory increase did —
nine assertions pass again on three images with the machines back at
1 GiB. Three machines at 2 GiB on a host doing other work contend enough
that the store container does not come up at all.

The ordinary lesson, and it still caught me: two changes together, the
failure attributed to the plausible one, and an issue written recording
the wrong cause. What found it was reverting to the exact last-known-good
state rather than reverting the suspicious change.

What remains untested is whether a fourth image alone is fine. Probably.
Nothing has measured it, and the honest state of this issue is that what
it was opened about was never demonstrated.
2026-08-30 20:40:54 +02:00
jschoubben de2b6a2825 Issue 012 — a scenario machine cannot hold four images and raise a
substrate

Adding a fourth image to the two-machine scenario makes the bootstrap
fail every time, with the store's readiness check producing no output at
all — which says the container was not running rather than that the
database was slow. Three images pass nine assertions; four never get past
the store.

More memory did not change it, so memory is not the cause; the change is
kept because the reasoning holds on its own. Disk is the most likely
explanation and nothing has measured it.

It blocks proving the mesh runs its own artifact store, since the
registry module needs a registry image to mirror. The module is written
and accepted; what is unproven is a machine assigned it serving another.
2026-08-30 20:29:55 +02:00
jschoubben 1b49e684e1 Issue 011 — an action is a gate, which the first fix got wrong
The first fix continued past every failure, and the next lab run failed
at the bootstrap: the store did not answer in three minutes and then said
"the database system is shutting down". Carrying on past the readiness
gate had started the broker and the control plane against a machine that
was not ready, and on a small machine that is how a database still
initialising has its memory taken away.

An action is the only shape whose purpose is to make something true
BEFORE the next thing needs it, which is why it is the only one with a
verify. So a failed action stops what follows and nothing else does —
which fixes both this and the hostage problem the issue was opened for.
2026-08-30 20:07:41 +02:00
jschoubben 4a601ceb02 Issue 011 — one broken module stops every module after it
Found in the lab. A machine with one impossible module applied nothing at
all on every later push, and the mesh said "failed" without saying the
rest was never attempted.

Recorded with the evidence, including that the behaviour's test cited a
record which does not decide it: ADR 0010 argues about pipelines against
reconcilers and says nothing about whether one resource failing should
stop the next being attempted.

Fixed in mesh-host: everything is attempted, every failure reported.
2026-08-30 19:26:03 +02:00
jschoubben 6bcf0e4f9f Issue 010 fixed: origins keep the bundle and the mesh apart
The store records where each resource came from and each origin removes only
its own. Verified on the scenario that caused it -- eleven resources raised,
enrolled, sent the same two-resource declaration, and the store, broker and
control plane were all still running. A later declaration dropping a resource
still removed it, so removal by omission survived the fix.

Two more faults found while fixing it, both the same shape. A report published
to a routing key nobody bound vanishes: the broker accepts it, finds no queue,
drops it, and tells the publisher nothing -- so nodes announced what they had
applied into a void. And publishReport was discarding its error, so a node that
could not tell the mesh looked exactly like one that had.

Reports are mandatory now, so an unroutable one comes back and is said out
loud, and the binding covers every key a node may publish.
2026-08-29 16:43:28 +02:00
jschoubben 594ea10b07 Issue 010: the first declaration destroys the substrate
Found in the lab, doing the ordinary thing: raise a first node, enrol it, send
it a declaration. Both declared resources applied correctly and every container
on the machine was removed -- the store, the broker, and the control plane that
had sent the message. The link died mid-sentence because the broker carrying it
had just been torn down by what it carried.

Nothing is behaving incorrectly. Apply removes what the store holds and the
declaration does not name, which is what reconciliation means. The fault is
that the carried bundle and mesh declarations share one store, so the host
cannot tell what this machine raised for itself before there was a mesh from
what the mesh told it to have.

It is invisible until those two meet, which happens exactly once per mesh: on
the first node, after enrolment, the moment the control plane first speaks.

The report says what is not the answer, including the tempting one -- having
the control plane send the substrate back. It cannot: it was never told what
the bundle contained, and the bundle exists precisely because there was no
control plane to ask.
2026-08-29 16:23:01 +02:00
jschoubben 087a8f4144 Close 009: a sealed machine now pulls by digest
The resolution was the one the issue predicted -- a registry inside the
scenario -- and it is the real path rather than a stand-in, since that is what
every node after the first pulls from.

The digests are the lab registry's own, which satisfies the pinning rule: what
is required is a reference that is exact and cannot move, and one this registry
assigned is both. That was the insight that unblocked it; I had assumed the
upstream digest had to be preserved, which is what made it look impossible.

The fault worth keeping is recorded in the issue: the read-back checked that
the catalog endpoint answered by matching the substring 'repositories', which
an empty catalog also contains. It passed on a registry holding nothing. This
repository's own subject, arriving in the tooling built to catch it.
2026-08-29 00:05:20 +02:00
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00
jschoubben e1febe8e0f Renumber the records 1 to 23
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18,
19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only
the archaeology of what used to be there.

Renumbered contiguously. Renames run in ascending order, so every target number
is already free and no two files ever collide.

The reference rewrite is one simultaneous pass rather than a sequence of
replacements. Numbers moved into slots other numbers were vacating -- the node
host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time
would have cascaded and silently pointed things at the wrong record.

Seven plain-text references survived the merges as prose rather than links,
naming records that no longer existed: the enrolment token, the link boundary,
what a declaration is, reachability, the repository structure. Each mapped to
the consolidated record that now holds it.

Verified rather than assumed: every [ADR NNNN](path) link now has matching text
and target, checked across the whole repository, and the checker passes.

Frontmatter `consolidates:` lists dropped -- they named records that are gone,
and each consolidated record already says in prose what it absorbed.
2026-08-28 23:28:34 +02:00
jschoubben 77f3a4cea7 Consolidate: 65 decision records to 23
Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.

  the node host          8 -> 1    applies not decides, depends on nothing,
                                   per operating system, root service, the
                                   launcher, episodic, what a declaration is,
                                   actions from the bundle only
  a node and how it joins 4 -> 1   what a node is, joining, the link as
                                   security boundary, the enrolment token
  modules and the graph   7 -> 1   everything is a module, no domain modules,
                                   three edges, provisioning, the core library
  substrate and control   6 -> 1   the test, seven contexts, one control plane,
    plane                          the authority is not a database, the named
                                   products, the pinned bundle
  connectivity            3 -> 1   a route is a grant, reachability declared,
                                   filter rules
  delivery                5 -> 1   reconciliation not a pipeline, artifacts,
                                   the three silos, a failed step, the verdict
  the lab                 5 -> 1   (earlier)
  how this repository     10 -> 1  (earlier)
    works

Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.

The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.

The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
2026-08-28 20:03:24 +02:00
jschoubben 5e83ac2c22 Consolidate: 65 decision records to 52
Jochen: a normal application has 3-5 ADRs, maybe 10 for a large one, and we are
at 65. Fair, and the cause is mine -- I recorded every FINDING as a decision
rather than every fork in the road.

Two merges, both cases where one decision had been split across many records
because it was taken over several days rather than at once.

0019 absorbs ten records about how this repository works: what it is and that
it is public, the folder flow, the two design layers, the issue front door,
status in frontmatter, playbooks, the naming rule, the product name. Those were
never ten decisions -- they were one, seen from ten angles as the repository
took shape.

0016 absorbs the five about the lab: a node is a virtual machine, a router is
scenery, a scenario declares the underlay, a scenario is a closed address
space, and the two scenario classes. Same pattern -- one design, split by the
order it was worked out in.

The consolidated 0019 also raises the bar for what earns a record, since that
is what produced 65: a record is warranted when there is a genuine fork -- a
direction reversed, an alternative that will be proposed again, something
contested. A finding is not a decision, and a bug is certainly not. Everything
else belongs in the design document where the reasoning is actually read.

The checker earned its place here. Deleting nine records left 13 dangling links
across the repository and it named every one, including in AGENTS.md. Nothing
was found by reading.

Remaining clusters worth the same treatment: the host (8 records), delivery
(5), modules (6), connectivity (4), substrate and control plane (4). That would
be 52 down to roughly 30.
2026-08-28 18:53:19 +02:00
jschoubben f728c3fd98 File 009: a digest-pinned image cannot be placed in the lab
Two accepted decisions collide, and testing found it rather than review.

0046 pins images by digest and has the host refuse anything unpinned. The lab
places images by exporting them from the workstation, because a sealed scenario
cannot reach a registry -- and that loses the digest, since a repo digest only
exists for an image a registry served. Measured: the load says 'Loaded image
ID:' rather than 'Loaded image:', and the image lands dangling.

So a tag is refused by the host and a digest is unusable in the lab. There is
currently no declaration the lab can raise that exercises the container shape,
which matters because the container shape IS the substrate -- every bootstrap
step past the runtime is one.

The resolution is a registry inside the scenario, and that is not a workaround:
0048 already names an OCI registry as substrate and every node after the first
pulls from the mesh's own. It also removes the lab's export-and-push mechanism
rather than repairing it.

0046 now carries a pointer, since its own consequence is where the collision
was predicted -- half of it is closed and the other half turned out to be
harder than 'not solved here' suggested.
2026-08-28 01:56:49 +02:00
jschoubben 0a37d751e2 Graduate research 003; file the rescue that does not exist
A sweep of the nine active research efforts. 003 was answered five days ago and
nobody closed it -- the decision it asked for was taken without citing it, which
is how an effort stays `active` after being resolved.

Its recommendation is what the mesh adopted, and the match is exact rather than
approximate. "Run the daemons as containers, making Docker the supervisor for
everything" is ADR 0057. Its warning that a mesh-native supervisor inherits
fate-sharing "unless it sits outside the mesh's own process tree" is where ADR
0061 put the launcher. And its insistence that it cannot be all-or-nothing is
why the host itself is the one thing an init starts.

Its incidental finding does not graduate with it, so it is now issue 008: the
automatic node rescue the documentation describes does not exist. No unit
declares OnFailure=, nothing calls the rescue script on a timer.

That is worse than having no rescue. A rescue nobody wrote is a gap somebody
can see; a documented one that is absent is a gap nobody looks for, and the
documentation is read exactly when a node has failed and somebody is deciding
whether to intervene.

The issue names two honest resolutions -- implement it, or delete the
documentation and say a failed node needs a person -- and says the choice is
scheduling rather than technical, since the new host's recovery is built and
tested. It also says what would make the finding certain: it came from reading
the repository, and confirming it on a running node is the difference between
"no unit declares this" and "no unit in the source declares this".
2026-08-28 01:39:10 +02:00
jschoubben 4a628bf3fe Issue 007: the instance is fixed, the class is the issue
The incus hook landed and does the post-install work — group, subordinate
id ranges, both units, storage pool, bridge, default profile. Verified
independently here: group exists with the operator in it, both id files
carry the range, service active, pool reports CREATED. The only thing that
fails is a shell whose process tree predates the usermod, which is how
group membership works and not a defect.

So the mechanism was never missing. Hooks are the right place and they
work. The gap is narrower and worse: the hook did six things, six checks
were then performed by a human by hand, and nothing in the pipeline
asserted any of them. A pipeline that dispatched a hook which silently
never fired would have been green in the same 48 seconds — and a hook
named for a feature its module does not carry is skipped without
complaint, thirteen of which were found at once in the past.

The six manual checks are, almost word for word, the module's own
verification: outcomes rather than steps, which is exactly the shape the
lab design asks for. They currently live in a chat message. In the module
they would run on every delivery to every node.

Status moves to diagnosing rather than resolved, and fixed-by records the
instance explicitly as the instance only.
2026-08-23 22:21:59 +02:00
jschoubben b4904fec7e The lab comes first, and its first scenario has no pipeline
The lab was designed around a module under test, with a scenario being a
complete mesh — forge, coordinator, cascade, verify. That is unusable for
building the new mesh, because all four are tier 2 and do not exist yet.

And research 009 had the sequence backwards. It placed the lab at phase B
as verification of tiers already built, but tier 0 is the component that
takes over a machine's packages, services and network. It cannot be
developed against a machine anyone needs. The lab has to exist before the
thing it will test.

ADR 0029 splits scenarios into two classes. The bootstrap scenario is
virtual machines, the host binary and a pinned bundle, with the verdict
coming from what the host reports about the state it reconciled. The full
scenario is the designed one. The first is a strict subset of the second —
same virtualisation, same networking, same lifecycle, stopping before a
control plane exists — so the second is reached by addition rather than
rework.

The consequence worth having: raising a node from nothing stops being the
least-exercised path in the system and becomes the inner development loop.

It also settles the runner's two jobs. Scenario lifecycle is needed
immediately, because something must materialise and reset a mesh before
anything can be written against it. Assertion execution waits for the full
scenario.

Corrects a stale claim in the design while amending it: it argued
scenarios were affordable with system containers and would not be with
virtual machines. ADR 0016 superseded that reasoning and the text had not
followed.

Issue 007: the lab's first requirement is installed and unusable. The
virtualisation package is present and explicitly installed; both units are
disabled, the operator is in no group, and the client reports the server
unreachable. Not issue 001 again — that is an install failing while
reporting success. This is an install succeeding when success was not the
point. A package is files; a capability is a running service and an
identity permitted to reach it, and the module model has no vocabulary for
the second.
2026-08-23 21:57:14 +02:00
jschoubben c0ae8dec96 Remove two disclosures, and record Nox as the answer to 006
Found by a full scan before making the repository public, which is the
moment the public rule stops being aspirational.

A module name identified a specific laptop model — hardware inventory,
which is operational detail about one installation rather than a lesson
that travels. Generalised.

ADR 0028 named a forge username in a repository path, which the public
rule forbids, and the sentence had also gone stale: the repository it
described was subsequently verified empty of anything unique and removed.
Rewritten to state what happened without the username. Removing a
disclosure from a record is the same class as fixing a path — the rule
that permits it outranks the one that forbids editing.

Issue 006 gains its proposed direction: Nox works from within this
repository rather than these documents being synced into the knowledge
base. Better on three counts — no copy, so no drift; no fourth knowledge
system, which was the original objection; always current.

But it changes the promise, and the issue says so. ADR 0019 promised these
documents would surface BESIDE everything else in a symptom search. An
agent that must be asked is reachable, not surfacing, and the two differ
in precisely the case the operational memory exists for — someone
debugging an error with no reason to suspect HQ knows anything about it.
The question narrows to whether a symptom search finds this content
without the searcher already suspecting it.
2026-08-23 21:40:53 +02:00
jschoubben 87f4f29cc6 Novox Mesh, Nox, and HQ becomes company-scoped
ADR 0027 — the product is Novox Mesh, shortened to mesh internally. HAL was
never chosen: it arrived with the dotfiles repository this grew out of, it
is borrowed, and it is borrowed from the canonical untrustworthy machine
intelligence, which is an odd flag for infrastructure trusted with
credentials. Timing is the substance of the decision, not an aside — the
skeleton is not built, so renaming costs a search and replace now and a
migration later.

Nox is an identity of Novox, and specifically the agent of the MESH rather
than of a node. Nodes keep their own identities. Nox addresses them, and a
human mostly talks to Nox — which makes it the concrete form of the
mission's vision: state an intent, and the mesh works out which node holds
the thing. It holds no private channel. The gap this opens is recorded:
ADR 0012 binds every agent to a home node, and a mesh-scoped agent has
none, so the model needs extending.

ADR 0028 — HQ is company-scoped, novox/hq, with the mesh as its first
product. Checked rather than assumed: the company organisation already
holds live projects that the mesh builds and deploys, so they are tenants
rather than peers, and the mesh is the ground they stand on. There is also
company work outside the mesh already, which strengthens the case and means
the eventual split is closer than "some day" — so each document's scope is
fixed now, in a table, making that split mechanical instead of
archaeological. The folders are deliberately not restructured yet.

The skeleton takes the new vocabulary: mesh-host, mesh-substrate,
mesh-control, mesh-surfaces, mesh-catalog. Substrate drops to four services
now that identity is a hosted workload rather than a dependency.

Research 009 opens the migration, with the reframing that lowers its risk:
replace the control plane, do not move the workloads. Their data never
moves, so it is re-declared rather than adopted — which keeps adoption out
of scope, as the lab design requires. Self-hosting is the last phase, or a
failed cutover takes away the means to fix it.
2026-08-23 21:20:35 +02:00
jschoubben c0b35652d0 The numbering is the flow: decisions are 02, design is 03
papa-hq reads 01 research -> 03 decision -> 02 design. The order is a
scar, not a choice: 02-DESIGN existed from its initial commit, and when
adr/ was finally promoted on 2026-07-13 it took the next free number
rather than its place in the sequence. By then design was too settled to
renumber.

hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and
02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks
the process in the order it happens: research produces a decision, the
decision authorises a design.

00-GENESIS becomes 00-META, matching papa's rename from the same
restructure.

Every path reference rewritten across documents, frontmatter, playbooks
and skills. All links resolve; all 58 frontmatter blocks parse and their
path fields still point at files that exist.
2026-08-23 18:05:11 +02:00
jschoubben 702efca6bb Base layer: the mesh as it is, under the mesh as it should be
HQ held only the to-be. Every reader had to already know the system the
decisions were about, and an as-is claim had nowhere to live except inside
an intention.

Adds 02-DESIGN/00-as-is — eleven documents written from the implementation
and the operational record, not from intent, including the parts nobody
would choose again. The two existing designs move under 01-to-be. Layers
are declared in frontmatter and never mix: a design that ships does not
move, its as-is counterpart is written, and both stand.

Back-fills adr/0001-0014 for decisions taken in implementation and never
recorded — the broker, the module abstraction, the mesh database, managed
files, provisioning, migrations, the workspace removal, failing loudly,
the constitution, application placement, linking, the employee model, the
artifact, the three silos. Each marked reconstructed, dated from the
history, and citing the evidence it was recovered from. The two existing
records renumber to 0015 and 0016 so the ledger runs oldest first;
0017 extends 0015 to modules outside the core, principle only — the
domain list is deliberately not invented here.

how-we-build.md becomes the source of the mesh constitution, with a sync
playbook, so the enforced copy stops being the only one that is true.

Process becomes explicit: five playbooks, eight thin skills that defer to
them, a repository map, and AGENTS.md with CLAUDE.md as its include.

The five Observations become 04-ISSUES 001-005 where they can be owned and
closed. 006 is new and uncomfortable: HQ is not indexed into the knowledge
base. That claim is what decision 27 rests on, it was never checked, and
the README now says so instead of repeating it.

Also corrects the ADR index into something generated, the "02-DESIGN is
empty" claim, the VISION.md pointer that did not survive the repo split,
and a note asserting the symlink rule was contradicted — it was a
misreading; the rule forbids hand-made links, the installer links by design.
2026-08-23 03:08:26 +02:00