Commit Graph
64 Commits
Author SHA1 Message Date
jschoubben 37252f9c3e ADR 0100: the guard lets the machine itself through; in use is a non-loopback listener; openings say from where; 09 in step with the flip 2026-09-22 16:34:53 +02:00
jschoubben f3152d827f ADR 0100 after re-review: the bus and registry stay reachable for enrolment; the mesh guards the store in a table that only refuses; a machine in use defined; the flip refuses while a found container is held; held containers and returning to adopted spelled out 2026-09-22 16:32:37 +02:00
jschoubben 02c40bcab4 ADR 0100 after review: found means unrecorded; assigning prepares, taking cuts over; openings through the found firewall on both paths; the mesh guards its own ports; ports kept as node settings; a converged genesis refuses a machine in use; designs 05, 07, 08, 09 and 17 in step 2026-09-22 16:28:31 +02:00
jschoubben 111456abb5 ADR 0100 accepted; the node host, connectivity, the node lifecycle and raising a mesh amended for a node adopted before it is converged 2026-09-22 16:20:27 +02:00
jschoubben 5fc3cbde4c Research 012: migrating a node that is in use, measured on the control-node; ADR 0100 proposed — a node in use is adopted before it is converged 2026-09-22 16:16:28 +02:00
jschoubben becae7ba51 ADR 0080: the development cycle is checked, not trusted
The flow the process overview draws — idea/symptom -> decision -> to-be design -> code ->
as-is — was enforced by nothing. cycle.py now refuses a to-be design naming no decision, an
in-progress/implemented design naming no owning code, a located/fixed issue with no owner,
a fixed/resolved issue with no fix, and a graduated research overview that does not say what
it became. AGENTS.md carries the cycle and a where-to-look table so a fresh session (or a
cleared context) finds the chain in frontmatter instead of assuming it. Grounding the check
surfaced two real gaps, fixed here: the work-ahead design named no owning code, and research
003 listed one became target twice.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:28:03 +02:00
jschoubben 1111bd84d7 Establish the repo for the completed Phase 0-3 build
Settles the design repository now that the self-upgrade build is on main:
- Records the two decisions that shipped without a record — ADR 0077 (the
  controller/foundation/node vocabulary) and ADR 0078 (the store and broker are
  ordinary modules); accepts ADR 0075 and 0076, which shipped work rests on.
- Fills issue 051's amended-design and wires ADR 0078 into 07-the-foundation.
- Sweeps the repo rename (mesh-control -> mesh-controller) into the mutable docs
  now that the forge repo is renamed; updates the glossary note and repos.md.
- Fixes the six broken links from the design-doc renames, indexes the glossary,
  regenerates the decisions reading order.

Both checks (records.py, index.py) are green. Statuses stay honest: the build is
on main and lab-proven but not deployed as the production mesh, so the to-be docs
remain in-progress and the as-is layer (the hal mesh) is unchanged — graduation
to implemented + as-is belongs to deployment, not merge.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 00:04:58 +02:00
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00
jschoubben e1febe8e0f Renumber the records 1 to 23
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18,
19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only
the archaeology of what used to be there.

Renumbered contiguously. Renames run in ascending order, so every target number
is already free and no two files ever collide.

The reference rewrite is one simultaneous pass rather than a sequence of
replacements. Numbers moved into slots other numbers were vacating -- the node
host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time
would have cascaded and silently pointed things at the wrong record.

Seven plain-text references survived the merges as prose rather than links,
naming records that no longer existed: the enrolment token, the link boundary,
what a declaration is, reachability, the repository structure. Each mapped to
the consolidated record that now holds it.

Verified rather than assumed: every [ADR NNNN](path) link now has matching text
and target, checked across the whole repository, and the checker passes.

Frontmatter `consolidates:` lists dropped -- they named records that are gone,
and each consolidated record already says in prose what it absorbed.
2026-08-28 23:28:34 +02:00
jschoubben 77f3a4cea7 Consolidate: 65 decision records to 23
Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.

  the node host          8 -> 1    applies not decides, depends on nothing,
                                   per operating system, root service, the
                                   launcher, episodic, what a declaration is,
                                   actions from the bundle only
  a node and how it joins 4 -> 1   what a node is, joining, the link as
                                   security boundary, the enrolment token
  modules and the graph   7 -> 1   everything is a module, no domain modules,
                                   three edges, provisioning, the core library
  substrate and control   6 -> 1   the test, seven contexts, one control plane,
    plane                          the authority is not a database, the named
                                   products, the pinned bundle
  connectivity            3 -> 1   a route is a grant, reachability declared,
                                   filter rules
  delivery                5 -> 1   reconciliation not a pipeline, artifacts,
                                   the three silos, a failed step, the verdict
  the lab                 5 -> 1   (earlier)
  how this repository     10 -> 1  (earlier)
    works

Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.

The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.

The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
2026-08-28 20:03:24 +02:00
jschoubben 5e83ac2c22 Consolidate: 65 decision records to 52
Jochen: a normal application has 3-5 ADRs, maybe 10 for a large one, and we are
at 65. Fair, and the cause is mine -- I recorded every FINDING as a decision
rather than every fork in the road.

Two merges, both cases where one decision had been split across many records
because it was taken over several days rather than at once.

0019 absorbs ten records about how this repository works: what it is and that
it is public, the folder flow, the two design layers, the issue front door,
status in frontmatter, playbooks, the naming rule, the product name. Those were
never ten decisions -- they were one, seen from ten angles as the repository
took shape.

0016 absorbs the five about the lab: a node is a virtual machine, a router is
scenery, a scenario declares the underlay, a scenario is a closed address
space, and the two scenario classes. Same pattern -- one design, split by the
order it was worked out in.

The consolidated 0019 also raises the bar for what earns a record, since that
is what produced 65: a record is warranted when there is a genuine fork -- a
direction reversed, an alternative that will be proposed again, something
contested. A finding is not a decision, and a bug is certainly not. Everything
else belongs in the design document where the reasoning is actually read.

The checker earned its place here. Deleting nine records left 13 dangling links
across the repository and it named every one, including in AGENTS.md. Nothing
was found by reading.

Remaining clusters worth the same treatment: the host (8 records), delivery
(5), modules (6), connectivity (4), substrate and control plane (4). That would
be 52 down to roughly 30.
2026-08-28 18:53:19 +02:00
jschoubben 9d091c81e0 A build edge, a core library that is a domain, and 0063 corrected
Three things from walking a real dev cycle through 0063, all of which Jochen
caught by pushing on where I had glossed.

0064 -- a build edge is a third kind. Research 011 established presence and
instantiation, and both are RUNTIME edges: they answer what a module needs in
order to run. Delivery needs a different question -- what has to be rebuilt when
this changes -- and that relationship is fixed inside an artifact rather than
negotiated when it runs. So the graph as designed could not drive delivery,
which is the real reason 0063 was not approvable.

It is derived rather than declared, read from what a module actually imports,
because a declared list and the imports it describes drift and the imports are
the true ones. The runtime edges stay declared, and that asymmetry is not an
inconsistency: a runtime edge is an intention somebody has, a build edge is a
fact about code that exists.

It also makes design quality measurable. A module with many inbound build edges
is one whose every change is expensive, and the current shared library is
exactly that -- nobody could see it because nothing drew the edges.

0065 -- the core library is the mesh's domain. Jochen disagreed with 0030's
"types, not behaviour" and was right: that guard is aimed at the wrong thing. A
library everything depends on is a hub whether it holds types or code, and the
fan-in is what makes a change expensive. So types ship with the module that
owns them -- trading one wide edge for several narrow ones -- and the core
library holds what is true of the mesh regardless of context, which research
011 already found: a module, a node, an assignment.

The test is "would this still mean the same thing in a context that had never
heard of the one it came from". A node does; a pipeline stage does not.
Domain-driven is the point rather than the label: "who else might want this"
always answers yes, which is how the current one grew.

And it changes the check for the better. "The build output contains no runtime
code" would have enforced a rule now withdrawn. Inbound build edges is a
measurement rather than a prohibition, and it is visible while a hub is forming
rather than after.

0063 revised on both counts, plus a third: I had written "the lab judges it" as
though that were a step. A lab run takes tens of seconds, occupies a VM, and
fails for environmental reasons -- and a shared-library change produces dozens.
One expensive non-deterministic gate fails both ways, and neither failure looks
like itself. Verdicts are now tiered, and a run that failed environmentally is
explicitly not a verdict.

0063 also now carries what must exist before it can be implemented, rather than
leaving that to be discovered.
2026-08-28 18:25:28 +02:00
jschoubben 4ab8a0507f Delivery is reconciliation, not a pipeline; research 008 closes
Jochen: don't rebuild the current coordinator, use it as a pitfall list. That
reframed the last open question rather than answering it.

0058 stopped deploy being a stage that pushes to nodes, and said plainly what
it did not fix: detection. A merge that created no pipeline, and nothing said
so. That is not a defect in the detector -- it is what happens when correctness
depends on an event ARRIVING.

0063 applies 0058's move one level up. The control plane holds what source
exists and what has been built from it, and builds the difference. A change
becomes a build because source is ahead of artifacts, which is a comparison
answerable at any moment. An event makes it fast; nothing makes it necessary,
so a missed webhook costs latency and cannot cost correctness.

The mesh becomes one idea at two layers: the control plane reconciles artifacts
against source, the host reconciles machine state against declarations. The
pipeline as a state machine disappears, and with it the stage list that a
verify step was once omitted from.

That reframing answered the three questions still open in 008, so it graduates
with all six closed. A deployed state is two comparisons rather than an event.
A verdict is about an ARTIFACT and gates whether it may be declared -- sharper
than the question expected. And "before self-hosting" mostly dissolves, because
a reconciler needs source and artifacts as bindings where a pipeline's stages
name their targets.

Four costs recorded, and one is a real risk rather than a trade: a reconciler
that cannot reach its target retries forever, and without something noticing,
the failure is silence -- the exact fault this removes, reintroduced elsewhere.
Also named: the run identity people actually use is lost, and "did my change go
out?" needs a replacement or this will be worse to live with than what it
replaces, whatever its properties.
2026-08-28 02:59:42 +02:00
jschoubben 9dc57b4712 Graduate 005; record what 0058 answered in 008
Continuing the sweep. Both were answered by records that did not cite them,
which is the same pattern 003 showed -- an effort stays active because the
decision that resolved it was reached from another direction.

005 graduates. Three of its four questions are answered: provider modules do
not group (0044), the ~50 modules that co-change with nothing stay as they are,
and 'group or leave' was never the right pair -- 0054 reframes it as authority
versus package. Worth noting the debt runs the other way too: this effort's
measurement, that reachability is the ONLY place modules genuinely co-change,
is what 0054 rests on and why connectivity is a context while nothing else
needed one.

Its fourth question moves rather than closes. Whether applications leave the
monorepo before or after they group is a sequencing question, so it belongs to
009-migration.

008 stays active, with its central question marked answered: the coordinator
converges nodes on a declaration rather than dispatching stages (0058). The
three-silo split survives with the third redefined. What 0058 explicitly does
NOT answer is how a change becomes a pipeline reliably -- detection is upstream
of everything it changed and remains the fragile input.
2026-08-28 01:40:02 +02:00
jschoubben 0a37d751e2 Graduate research 003; file the rescue that does not exist
A sweep of the nine active research efforts. 003 was answered five days ago and
nobody closed it -- the decision it asked for was taken without citing it, which
is how an effort stays `active` after being resolved.

Its recommendation is what the mesh adopted, and the match is exact rather than
approximate. "Run the daemons as containers, making Docker the supervisor for
everything" is ADR 0057. Its warning that a mesh-native supervisor inherits
fate-sharing "unless it sits outside the mesh's own process tree" is where ADR
0061 put the launcher. And its insistence that it cannot be all-or-nothing is
why the host itself is the one thing an init starts.

Its incidental finding does not graduate with it, so it is now issue 008: the
automatic node rescue the documentation describes does not exist. No unit
declares OnFailure=, nothing calls the rescue script on a timer.

That is worse than having no rescue. A rescue nobody wrote is a gap somebody
can see; a documented one that is absent is a gap nobody looks for, and the
documentation is read exactly when a node has failed and somebody is deciding
whether to intervene.

The issue names two honest resolutions -- implement it, or delete the
documentation and say a failed node needs a person -- and says the choice is
scheduling rather than technical, since the new host's recovery is built and
tested. It also says what would make the finding certain: it came from reading
the repository, and confirming it on a running node is the difference between
"no unit declares this" and "no unit in the source declares this".
2026-08-28 01:39:10 +02:00
jschoubben 93470f6162 ADR 0047 — the bundle may carry actions the link may not
The bootstrap's sharpest open question, and the framing was wrong. "State on
this machine" was being read as the filesystem and the service manager. A
service running on this machine IS part of this machine — writing a file and
creating a database in a local store differ in mechanism, not in scope.

The real question was underneath: must the host learn what a database is? It
must not. Giving it a `database` resource type means tier 0 knows Postgres, then
a bucket, then a virtual host — the host acquiring the substrate's vocabulary
one service at a time, which is what ADR 0037 exists to stop.

So the bundle declares an ACTION and the host runs it and verifies it. What a
database means stays with the module that provides one; the host knows only how
to run a declared action against something local and check the result. Its
vocabulary grows by one shape rather than by one resource type per service.

Actions are permitted in the bundle and forbidden over the link, and the
asymmetry is deliberate. A bundle arrives WITH the binary: anyone able to put a
hostile action in it could equally have put it in the host itself, so refusing
actions there buys nothing and costs the bootstrap. The link is a separate
party, reachable separately, and an action there is the unbounded blast radius
ADR 0039 refuses. That decision stands unchanged.

And ongoing provisioning is not the host's at all — the control plane does it
once a mesh exists — so the asymmetry costs nothing.

Which dissolves the earlier worry about one mechanism with a tier boundary
inside it: there are two mechanisms, with different actors, scopes and trust
models, and that is the answer rather than a compromise.

Named rather than hidden: this is the escape hatch research 011 warned about,
arbitrary code in the place hardest to remove later. It is bounded by being
bundle-only and by every action having to declare how it verifies itself, and
that boundary is the whole defence.
2026-08-26 23:56:39 +02:00
jschoubben 5b3d0ebd4f ADR 0046 — the installer fetches what it pins
The blocking question was where a container image comes from, and the version
that blocked assumed the machine might have no network. That assumption came
from the LAB: a scenario is a closed address space by design, which is what lets
two scenarios hold the same addresses without meeting. Production is not sealed
— a machine being adopted has a network, and one that does not is a machine
where very little works anyway.

So substrate.lock carries references, not payload: an image name and a digest,
fetched at apply time. A first node pulls from upstream because no mesh registry
exists yet; every node after that pulls from the mesh's own. The lab is the
exception and places images itself, the way it already places the host binary —
a property of a test environment, and letting it dictate the production design
would be the tail wagging the dog.

Pinned by DIGEST rather than tag. Reproducibility comes from pinning the
identity of a thing, not from carrying its bytes, which is what makes fetching
acceptable rather than a compromise.

ADR 0041 survives untouched, which was the point. "Copy it onto a machine and
run it" stays literally true — one binary, a few megabytes, which then fetches
what it was told to. Carrying images would have quietly redefined the property
that decision rests on.

Costs accepted and named: an apply can now fail because something is
unreachable, which a self-contained artifact could not, so it must fail legibly
— naming what it could not fetch and from where. And the lab needs a way to
place images into a machine that also has no container runtime, both of which
are lab-installation concerns and neither solved here.

Research 012's build-time-versus-apply-time reframing narrows accordingly: it
still holds for what a tailored installer contains, and no longer has to hold
for images.
2026-08-26 23:52:46 +02:00
jschoubben 0531d6fc38 ADRs 0044 and 0045 — the module design, closed; 011 graduates
011 opened asking what a graph deletes and found the graph already existed. The
work became design, worked through twenty cases and one provider in full. Two
decisions close it.

0044 — a module declares presence, instantiation and exclusion. Two kinds of
edge because a game wanting a database is not a game wanting postgres to exist:
one creates something per consumer, carries credentials back, can be revoked,
and leaves the provider holding state. Names are concrete unless providers are
genuinely substitutable — `terminal` passes, `database` fails, and the adapter
is what creates an interface. Where there is no contract there is a tag, which
describes and does not bind. Exclusion is a third relation and is not derivable.
A node provides names too, which makes capability checking stop being a separate
mechanism and makes the host's detection an input to resolution. Constraints,
never placement. Scope decides which provider and the binding is written down
and sticky, in a place that follows the scope. And there are three entities, not
two — the assignment carries what belongs to neither end, which is what
node-agnostic modules ran out of.

0045 — a context owns its store, exclusively. No shared writes and no read roles
on another context's store, because reading couples you to its layout just as
firmly and invisibly. The unit is the CONTEXT, not the process: a board showing
the mesh's own data is the mesh showing its own data. Asking or subscribing is
derived from ADR 0036 rather than chosen. And it is the first clear list of what
the design removes: grant kinds, table ownership, cross-context migration
ordering, and a class of permission modelling.

0017 is superseded rather than narrowed — its text unchanged, its status
changed. Folders assert relationships where edges record them, and the domain
module goes with it.

Left explicitly undecided in both: what a resolver delegates rather than
reimplements, how many instances a module should have, and what a provider hands
back.
2026-08-26 23:41:18 +02:00
jschoubben 60aea14935 Define the substrate, and answer 006's four-or-five conditionally
Same gap as the control plane: load-bearing and unpinned.

The substrate is what the control plane CONSUMES AND CANNOT GRANT ITSELF. Every
module needing a database asks provisioning for one; the control plane needs one
too and cannot ask itself, because it is not running yet. That circularity is
not an awkwardness to work around — it is the definition, and anything on the
wrong side of it must be raised by the bundle the host carries.

Which answers 006's open question in the honest form rather than with a number.
The identity provider is substrate only if the control plane DELEGATES
authentication — then it cannot serve anybody before the provider exists and
cannot grant itself a client. If it authenticates natively, the provider is an
ordinary hosted service. So the count follows from a decision not yet taken, and
asserting four was asserting that decision.

The test also rules out the tempting wrong answer: an identity provider, a mail
server and an analytics service are all infrastructure by any ordinary reading,
and none are substrate, because the control plane starts and runs without them.
Important is not the test.

Records why the bundle is pinned by hand — it is applied when no mesh exists, so
nothing can resolve a version or ask a registry — and why it must be
self-contained, which makes it an artifact built on a machine with a network for
a machine that may have none.
2026-08-26 23:39:11 +02:00
jschoubben a4ab3e15c2 011: one interface, many contexts — and the constraint that hides in it
The objection is right: if every context runs its own service with its own
interface, the board is coupled to N interfaces instead of N schemas, something
has to compose them, and composition is logic — which tier 3 says a surface does
not hold. That moves the problem up a layer rather than solving it.

The skeleton already answers it, and the previous entry talked past it. `work`
and `knowledge` are not separate services; they are contexts INSIDE the control
plane, alongside the record, inventory and delivery — and `api` is listed there
as the one interface every surface speaks to.

So the board speaks to one interface. Behind it the contexts stay separate,
integrating through the record, but they are one tier, one repository, one
deployable — and coupling within a tier is not what the tier rule forbids. The
problem does move up a layer, and the layer it moves to already exists and has
this as its job.

The caveat is load-bearing and now recorded as an open question: this holds only
while the contexts are not separate deployables. The moment one becomes its own
service with its own interface, the board is back to N clients, something must
compose them, and the composition has nowhere to live that tier 3 permits. That
is a real constraint on how far the control plane may be split, and it is worth
knowing before splitting rather than after.
2026-08-26 23:29:56 +02:00
jschoubben a4a25ca7e3 011: one surface over several contexts is normal
The board visualises the mesh, the work engine, the knowledge base and more, and
the alternative — a web application per context — is worse for everyone using
it. Composing several sources into one view is what a surface IS, so this is not
a compromise with the ownership rule.

What changes is only where it reads from: each context's interface rather than
each context's store. Most of that already exists — 56 of 126 modules carry a
tool surface, more than carry a service.

And the unified board is what keeps those interfaces honest. A view that cannot
be built from a context's interface proves the interface inadequate, discovered
where it is cheap to notice rather than the first time something else needs the
same data and quietly reaches for the store instead.

If composing many calls proves too slow, the answer is a projection the board
owns and keeps current from events, not access to somebody else's tables.
2026-08-26 23:27:42 +02:00
jschoubben fa62c7f0e4 011: correct the rule — contexts, not processes
An earlier version argued a dashboard reading a dozen stores was caught by
exclusive ownership, because a dashboard is a surface and surfaces speak to an
interface. Wrong, and it drew the line in the wrong place.

The mesh's own board showing nodes, modules and deployments is not a separate
context reaching across a boundary — it is the mesh showing its own data.
Requiring it to go through an interface to reach facts its own context owns is
ceremony.

The rule is that a CONTEXT is granted what it exclusively owns. Everything
inside it — service, surface, tools — reads that store freely. What is forbidden
is a different context reading it.

Which is what the consumer count already showed: the problem was never surfaces,
it was three other contexts keeping their tables in the mesh's database.
2026-08-26 23:25:46 +02:00
jschoubben e71d532c2e 011: request or subscription is derived, not chosen
Asked what the distinction actually is, and the SQL half needed correcting
first: under exclusive ownership SQL runs against your own database and nothing
else, whatever transport a query might travel over. Both options are the mesh's
own channel and both ride the broker, so the transport is not the distinction.

The distinction is where the answer lives when you need it. A request asks at
the moment and waits — always current, costs a round trip, cannot answer when
the other side is down. A subscription keeps a local copy — instant, works
offline, as current as the last event received, and you must handle what you
missed.

What decides is not taste. ADR 0036 makes disconnection an ordinary situation
rather than an exception, so anything that must keep working while disconnected
CANNOT use a request: there is nobody to ask. And the converse — anything where
a stale answer is worse than no answer cannot use a subscription. A display can
lag; a decision about whether a grant is still valid cannot.

So an apparently open question turns out to be derived from a decision already
taken. What stays open is narrower: what a consumer does about the events it
missed while disconnected — replay from a point, ask once for a full picture and
resume, or rebuild. The question every projection has.

Also recorded: separate databases are required in the new design, and the shared
registry is a leftover rather than a pattern.
2026-08-26 23:23:39 +02:00
jschoubben 7e83723b7b 011: rewrite the question table, which had gone stale silently
Several edits to the overview matched nothing and returned success, so the
question table still carried answers superseded two or three exchanges ago —
"when two modules provide one name, who chooses" was still open in the table
while answered in the file it pointed at, and nothing recorded the instantiation
edge, instance counts, grants, bootstrap provisioning, the tool audience, or the
registry consumer check.

That is the fault this repository catalogues, committed by the thing cataloguing
it: a string replacement that found no match, reported nothing, and left the
document claiming a state it did not have. Rewritten from what the documents
actually say rather than patched again.

Nine questions settled, fourteen live, and the split is now visible instead of
implied.
2026-08-26 23:17:25 +02:00
jschoubben afcc355744 011: checked the registry's real consumers, and the question was the wrong shape
The exclusive-ownership rule turned on whether every reader of the mesh registry
could be served another way. Eighteen consumers open a direct connection. Four
groups, and only one is work.

The owner and its machinery keep reading, because they own it. The node appliers
are already resolved — ADR 0037 stops the host querying the mesh database,
decided for tier reasons with nothing to do with this.

The bulk are FOREIGN TENANTS. The work engine holds ten of its own tables in the
registry's database, the knowledge base two, pipeline logs one. Thirteen foreign
tables across three contexts, which is how-we-build §4's shared schema counted.

So the question was the wrong shape: the problem is not readers needing a new
route to data, it is tenants needing to move out. Tasks, agents and teams have
nothing to do with nodes and modules and are co-located by history. Give that
context its own database and its dependency on the registry shrinks to one
table.

A handful of genuine cross-context reads remain, small enough to enumerate
rather than estimate. The rule holds.

Left open: whether those reads want an interface or events. Asking which nodes
exist at the moment you need to know is a request; reacting when a node appears
is a subscription, and some consumers want both.
2026-08-26 23:16:13 +02:00
jschoubben 6b1aab6a1e 011: the dashboard case, and why exclusive ownership is the tier rule
Raised as the hardest test of the rule: a board showing nodes, modules,
pipelines, agents and tasks wants to read a dozen stores, and under exclusive
ownership it can read none of them.

It survives, and not by luck. The board is a SURFACE, and surfaces already may
not do this — the skeleton puts `api/` in the control plane as the one interface
every surface speaks to, and tier 3 as thin, no logic. A board reading stores
directly is a surface reaching past the context that owns the data, which the
tier rule forbids for reasons that have nothing to do with databases.

So it is not a counter-example; it is an instance the rule catches. And the two
rules turn out to be one rule seen from two sides: exclusive ownership is the
tier rule expressed in terms of storage.

The general shape for anything needing to see across many things: consume the
record and own your own view. A reporting context builds a projection from
events and reads its own store, never anybody else's.

The cost said plainly rather than buried: a projection is more work than a join,
and it lags. A board queries the mesh's own database directly today — ordinary,
working — and this rule makes that a migration rather than a preference. The
reason to pay it is §4's already-measured cost, not elegance.
2026-08-26 23:10:36 +02:00
jschoubben aa767d17a8 011: a module is granted only what it exclusively owns
Reconsidered by the operator — maybe shared databases should not be allowed at
all — and the stricter version is better and goes further than the schemas it
replaces.

No shared writes, and no read-only role on another module's database either.
Reading another context's tables couples you to its layout exactly as firmly as
writing them does, and the coupling is harder to see because nothing breaks
until the owner changes a column.

That is how-we-build §4 taken at its word rather than at its letter. The
permissive version — a per-consumer schema, revocable, with cross-context joins
possible but deliberate — kept the letter and left the temptation. A boundary
that is merely inconvenient to cross is a boundary that gets crossed.

The cost is cross-module reporting, and it is the point rather than a
regrettable side effect: anything wanting to know what several modules hold
consumes their events or calls their interface. That is §4's whole argument, and
the mesh already has both mechanisms. What gets harder is precisely the thing
that was making work belonging to one context keep having to be implemented in
another.

And it is the first clear instance of what this effort has been hunting — what
the design DELETES rather than adds. Grant kinds collapse to one: an exclusive
resource. With them go the question of who owns which table, the guessing at
revocation time, cross-module migration ordering, and a class of permission
modelling a shared store would otherwise need.

One thing it does not answer, recorded because it could make the rule
unworkable: the mesh's own registry is read directly by many things today, and
under this rule they consume events or call tools instead. Achievable in
principle. Whether EVERY current consumer can be served that way is unchecked,
and should be before this becomes a decision.
2026-08-26 23:09:44 +02:00
jschoubben 4c8515507a 011: where a binding lives follows the scope, and a grant is not always a whole resource
Two questions asked directly, and the second collides with a rule in force.

One module on two nodes sharing a database corrects something stated flatly: the
binding is not "recorded on the assignment". Where it is written down FOLLOWS
THE SCOPE. A shared grant belongs to the module and every assignment references
the same one — which is the answer for two nodes wanting one database between
them. A per-instance grant belongs to the assignment. Same relation, two homes,
and which home is what makes two instances share something or not.

Several modules adding their own tables to one database is three needs wearing
one sentence, and a provider offers KINDS of grant rather than one: a database
for a consumer whose tables are nobody else's business, a read-only role for one
that needs to see what another holds, and a SCHEMA within a shared database for
the case actually asked about.

Loose tables in a shared database is what how-we-build §4 warns against in as
many words — several domains sharing one forty-five-table schema, which is why
work belonging to one context keeps having to be implemented in another. Not a
style objection; the observed cost, already paid.

A per-consumer schema keeps what the request wants and drops what §4 objects to.
Same database, same connection, same backup, and a cross-schema read remains
physically possible when genuinely needed. What it adds is ownership: migrations
touch one namespace, two modules cannot collide over a table name, and revoking
drops the schema rather than guessing which tables belonged to whom.

So the fault §4 names is still possible and no longer accidental — a
cross-context join becomes something somebody deliberately writes rather than
the path of least resistance. And revocation becomes answerable, which the
whole-database version never was.
2026-08-26 23:08:32 +02:00
jschoubben fa9889536c 011: tools have a different audience, migrations cross the edge, provisioning is early
Three additions, and the third kills an assumption.

Tools are the most common content in the catalogue — 56 of 126 modules, more
than carry a service — and they survive the split without fitting either half. A
tool is not an artifact and not node state; it is a contract the mesh publishes
on a module's behalf, and what consumes it is an AGENT rather than another
module. That is a second audience the design has not described. Whether it is
one relation with two audiences or two relations is cheap to decide now and
expensive later.

A migration belongs to the CONSUMER and runs on the PROVIDER. A game's
migrations run against the database the store granted it: owned by the consumer,
hosted inside something it does not control, ordered after the provisioning edge
because there is nothing to migrate until the grant exists, and scoped to that
grant. Ownership crosses the edge, which nothing in provides and requires
expresses — and it gives a consumer's own install an internal order, provisioned
then migrated then started, that depends on an edge rather than on its contents.

And provisioning is EARLY, not late. The assumption worth killing is that it is
something the control plane does for consumers once a mesh is running. The
mesh's own registry database is provisioned before there is a mesh, and so is
its virtual host on the broker: the store runs from the carried bundle, a
database is created in it, the mesh's own schema is applied, and only then does
a control plane exist. Steps two and three happen before there is a mesh to do
them, so provisioning is part of the bootstrap and part of what the bundle has
to express.

Which strains ADR 0043. The host applies declared state ON THIS MACHINE, and a
database inside a running store is not a file or a unit. At bootstrap it is at
least local — the store is on the same machine. Afterwards a consumer on one
node provisioned from a store on another is the ordinary case and reaching it is
not the host's job. The same operation is local at bootstrap and remote later,
which is either two mechanisms or one with a tier boundary crossing inside it.
Currently the sharpest unresolved thing in the effort.
2026-08-26 23:02:52 +02:00
jschoubben a3c7e7e1f1 011: providing is a facet, and the assignment is a third thing
Any hosted service can be a factory — an identity provider grants clients, an
analytics service grants a tracking identity, a mail server grants mailboxes, an
application platform grants a project that is several of those at once.
Providing is a FACET a module may have, not a kind of module it is, which is the
same conclusion this effort reached about services and applications arriving
from the other direction. So `provider` stops being a category too.

Two relational stores from different vendors both grant "a database" and are the
sharpest possible test of the substitutability rule. They fail it completely —
different protocol, dialect, driver, client library compiled into the consumer —
so `database` stays a tag, now with two real providers rather than a thought
experiment.

The assignment is a third entity, recorded because the operator tried the
alternative: modules were once node-agnostic and it did not survive. Several of
a provider's properties belong to neither end — where its state lives, how it is
reached, tuning derived from the machine's hardware, which instance serves a
given consumer. Not the catalogue, because they differ per node; not the node,
because they are about this module. A design with only modules and nodes has
nowhere to put them, which is what node-agnostic ran out of. The current system
already stores environment values per module AND per node, arriving the same
way.

Which answers the question asked directly: two nodes both run a store, so which
serves a consumer? Neither obvious answer. Not the consumer naming a node — that
is placement in the consumer's manifest, a game edited because a database moved.
Not the consumer not caring — for presence it genuinely does not, for
instantiation it cares permanently.

What the consumer knows is the SCOPE of its own need: one instance shared across
every instance of itself, or one each. That decides, and needs no node named.
Then the mesh binds, and the binding is recorded on the assignment and is
sticky — a resolver that re-derives which store serves a consumer will one day
derive a different answer and relocate a database.
2026-08-26 22:56:36 +02:00
jschoubben 13c6068874 011: the provider shape generalises, and two things differ inside it
The broker has all nine properties the store has. So do the object store and the
image registry. A substrate service is a SERVICE PLUS A FACTORY, there are four
of them, and the pattern generalises past the substrate: anything granting
something per consumer has this shape.

Two differences matter more than the similarity.

The broker cannot be managed over the broker. ADR 0001 makes it the channel
every node takes work from and ADR 0039 makes it the security boundary, so the
module providing it is also the way modules are managed — a declaration cannot
be delivered to it over itself. Nothing else has that property; the store is
consumed by the control plane but is not how the control plane REACHES anything.
This is what the carried bundle exists for: the broker is raised from what the
host carries because there is no other way to raise it. A constraint on one
module, not a general rule, and a schema with no way to say so hides it.

And two modules of identical shape want opposite instance counts. The broker is
one per mesh by decision. The store cannot be, because a node that must keep
working while disconnected cannot depend on a database elsewhere. Which settles
what cases.md left open: how many instances is NOT derivable from what a module
is. It is a per-module decision, it has to be declared, and nothing in provides,
requires or excludes says it.

Revocation differs in consequence too. Dropping a database leaves data until
something removes it — a leak, recoverable. Dropping a virtual host loses
whatever was undelivered — silent, and not. Same relation, different blast
radius, which argues for the provider deciding what revocation means rather than
the mesh applying one rule.

File renamed: it was never really about postgres.
2026-08-26 22:54:49 +02:00
jschoubben f160b28a71 011: postgres worked through, and "one kind of edge" was wrong
The tidy version said a module provides names and requires names and that is the
only edge. Working postgres through completely disproves it.

A small game wanting to store data does not require postgres to EXIST. It
requires postgres to MAKE IT A DATABASE and hand back credentials. Those are
different relations in every way that matters: one creates something per
consumer, carries a payload back, can be revoked, and leaves the provider
holding state about who was granted what. The other creates nothing.

So: two kinds of edge, one graph. Instantiation implies presence; presence does
not imply instantiation. The current system already had exactly this split —
`dependencies` for presence, `requires: provision:` for instantiation, with the
resolver deriving one from the other. analysis.md called that derivation a
convenience. It is not: it is the correct relationship between two genuinely
different relations, and the design had collapsed them.

Postgres also turns out to be nine things, not one. A container. Persistent
state where moving nodes is a migration rather than a reschedule. Configuration
partly derived from the machine's hardware. A tool surface. A provisioner. Its
own bookkeeping about what it granted, which is not the data it stores. An
exposure decision per node it runs on. Credentials it generates, which means a
provisioning edge carries a secret. And health that is not "the container is up".

Four questions the worked example makes concrete rather than abstract. WHICH
postgres, when there are two — a consumer of `terminal` does not care and a
consumer of a database cares permanently. How many instances a module should
have, which cannot be a global rule because one-per-mesh is wrong for a store a
disconnected node needs and one-per-node is wrong for the mesh's own registry.
What happens to a grant when its consumer is removed, where dropping is data
loss and keeping is a leak. And whether a declaration is composed PER NODE from
what that node reported — because tuning follows hardware the control plane
cannot know, and the alternative is the host deciding, which ADR 0037 forbids.
2026-08-26 22:53:59 +02:00
jschoubben c9c2dfe686 011: what a feature is, and what it splits into
The operator wants features gone, and 006 left it open. Measured, and the answer
is that nothing replaces them because they were never one concept.

A feature is a kind of content a module carries, detected from its directory:
twenty-one of them, each with a handler owning six stages — build, publish,
install, configure, start, verify.

The structural finding: EVERY handler implements EVERY stage. `configs` writes
files onto a node, has nothing to build, and has a build stage. `npm` publishes
to a registry, has nothing to start, and has a start stage. One interface spans
build-time and apply-time, so every kind of content must implement both halves
and most do nothing in one — and a stage that does nothing looks exactly like a
stage that failed to do anything.

They split four ways, across three tiers. Artifacts built once per version and
published, where no node is involved — delivery. Resources that are desired
state on a machine, which is what ADR 0043 already describes and the host already
does — tier 0. Actions run once against something that is not this machine, like
a migration against a database on another node — delivery, and seeds go
entirely. And checks: the prerequisites are REQUIREMENTS IN DISGUISE, a module
saying what must be true before it can be installed, which is what an edge in
the graph says; the verifiers are the read-back the host already performs.

So `feature` is one word for four things spanning three tiers, which is why the
pipeline is hard to reason about.

One property must survive the split, and it is the thing the current design got
right: content is DETECTED, relationships are DECLARED. A module that says it
has migrations and has none is a fault nobody sees until it matters — but what
it requires and provides is not visible in a directory and has to be said.
2026-08-26 22:51:19 +02:00
jschoubben ae099482a9 011: twenty cases, and two axes nothing covers
Before settling a schema, what a module can actually be. Twenty kinds of thing,
with the hard ones at the end because they are the point.

The ordinary nine are unsurprising: a supervised service, a system package with
configuration, an application a person launches, a command-line tool, a library
that never runs, a one-shot task, a scheduled one, an adapter, and a standalone
application whose only difference is where its source lives.

The eleven that break a naive schema are where the work is. Something that is a
service AND an application — a git forge is consumed as a remote and operated
through a web interface, and neither reading is wrong. Something that provides
and consumes, because provider and consumer are ends of edges rather than kinds
of module. Something the mesh installs that then becomes a node CAPABILITY,
which means a node's provides-list is partly derived from what is installed on
it and not only detected. Something that must be adopted rather than installed.
Something that is a set rather than a thing. Something with exactly one instance
for the whole mesh, where assigning it twice is not redundancy but two meshes.
Something that is not software at all — a firewall policy, a DNS record, pure
desired state, which fits the host's declaration model exactly and an installable
package model not at all. An agent. The host itself, which is not a module and
needs a schema that can say so. And the things the mesh depends on and does not
control, which are why a node can be perfectly configured and still not work.

Nine axes come out of it. Two are covered by nothing anyone has proposed: HOW
MANY INSTANCES a thing may have, and WHETHER TWO CAN COEXIST — `excludes` covers
part of the second and nothing covers the first.

And one question the cases sharpen: is "runs" a property or a kind? The axes say
property — one schema with a field saying how it runs, `never` included. The
alternative is several kinds of module with different schemas, which is the
taxonomy this effort already rejected once for services and applications.
2026-08-26 22:48:20 +02:00
jschoubben f55ecc1a47 011: an abstract name needs providers that are actually substitutable
Two corrections from the operator, and the first improves the design rather than
narrowing it.

`database` is not an edge. The test it fails, and the test the proposal was
missing: can a consumer be switched from one provider to another WITHOUT
CHANGING? A module speaking Postgres does not speak MongoDB or SQL Server —
different wire protocol, dialect, driver — so a consumer declaring `requires:
database` and handed any of them breaks. The name promises what no provider can
deliver, and the resolver would report a requirement satisfied that is not.

`terminal` passes: anything that runs a command in a terminal works and the
consumer never learns which it got.

So the ADAPTER is what creates an interface. `ai-assistant` is legitimate exactly
because adapters normalise what is behind it. Without one there is no interface,
there is a category — and a category is a TAG. Tags describe, edges bind, and
keeping them apart is what stops the catalogue acquiring a second kind of
relationship that looks like a dependency and is not, which is what a folder
named after a domain already was.

And the domain module goes. A `networking` module gathering a firewall, a
resolver and a proxy under one name came from an older shape and does not fit —
there is no such thing to install. There is core infrastructure: concrete
modules named individually, not flavourable, with no grouping module standing in
front of them.

Fixed three places where the revision left the old rule standing, including an
example manifest still requiring `database` — the kind of contradiction that
would have been read as the design rather than as a leftover.
2026-08-26 22:44:37 +02:00
jschoubben e20a09ae80 011: one kind of edge
The design, rather than an account of what exists. A module provides names and
requires names, and that single relation absorbs three things this effort had
listed separately: requiring another module is requiring a concrete name,
requiring a resource is requiring an abstract one, and an interface is simply a
name with more than one provider. Nothing has to declare that it is an
interface — it either has one provider or several.

The move that does the most work: a NODE provides names too. Its profile is a
set of them — display-server, container-runtime, an architecture — so a module
requiring a display server is satisfied by the node exactly as one requiring a
database is satisfied by another module. One resolution instead of two, and a
graphical application cannot land on a node without a display server for the
same reason, through the same code, that it cannot land without its libraries.

Which makes the host's capability detection an input to resolution rather than
something a person reads. It was built to be read; it turns out to be a
provides-list.

`excludes` is the one genuinely new relation, because it is not derivable: two
modules that both provide message-bus look interchangeable when installing both
would break the machine.

Constraints are not placement. They say what must be true of a node, never which
node — which is the mistake the measurement found in the current catalogue,
where a module pins its database to a named node so a second node cannot provide
it without editing the consumer.

What it deletes, for the design: the module/resource distinction, the interface
as a kind of thing, capability checking as a separate mechanism, domain grouping
— folders assert relationships where edges record them, so a domain becomes a
query over the graph rather than a directory somebody keeps true — and possibly
tiers, if a tier is just a computed level.

What it does not delete, stated so it is not discovered later: a resolver still
has to exist, with version constraints and conflicts, and the design owes an
answer on what it delegates rather than reimplements.
2026-08-26 22:39:17 +02:00
jschoubben c3a2984b3e 011: measured, and the premise was wrong — the graph is not missing
The effort was opened to ask whether the catalogue's missing structure is a
graph. It is not missing. 126 manifests, 103 edges, no cycles, nothing dangling,
deepest chain of five — and a resolver in the SDK that topologically sorts them,
already called by the tool loader at startup, the installer when syncing modules
onto a node, and the delivery coordinator when expanding what a change affects.

It already does something this effort assumed would need designing: a
requirement on another module's provision is treated as an implicit edge to the
module that provides it. So "ordering by the graph", which ADR 0043 makes the
control plane's job, is a thing to call rather than a thing to build.

The one place the graph is wrong, it is wrong about the substrate. A module
needing a database declares `provider: postgres` inside `provisions:` — which is
what a module OFFERS — so the resolver, which reads `dependencies:` and
`requires:`, never sees it. Three edges are invisible this way, and they are the
mesh's own database, the mesh's own broker, and the work engine's database.

The consequence is measurable: computing what a working mesh needs from the
declared graph gives registry -> sdk -> mesh -> meshware. Four modules, four
levels, no database. Arithmetically correct and obviously wrong, for exactly one
reason — a field that means "depends on" is not read as one. That is
04-ISSUES/003 in a new form: not a key nothing reads, but a key read as
something other than what it means.

Two latent defects, both contrary to ADR 0008 and both in the component ADR 0043
makes responsible for ordering a host will apply without question: a cycle warns
and falls back to input order, and a dependency that does not exist warns and
continues. Neither has fired, because the catalogue currently has no cycles and
nothing dangling, which is why nobody has noticed.

And placement is decided in the catalogue: a provision pins itself to a named
node in the manifest. Which node runs what is an inventory decision — tier 2 by
the skeleton's own test — so a second node cannot provide the mesh's database
without editing the module that consumes it.

What the graph would DELETE is currently nothing. What it would add is three
declarations that no manifest uses today: excludes, a required node capability,
and an interface with adapters. Whether they would be used is not measured, and
zero usage is equally consistent with nobody needing them and nobody being able
to express them.
2026-08-26 22:35:39 +02:00
jschoubben 278f7427ed 012: the briefing carries an outcome, derived from its lines
Proposed by the operator: state plainly whether adoption succeeded, partly
succeeded or failed, with a severity per line.

Taken with one change — the overall is DERIVED as the worst mark present, never
written alongside. Two fields maintained independently drift, and a briefing
reading "full success" while carrying a failed line is exactly the fault this
record keeps cataloguing. An outcome computed from its lines cannot disagree
with them.

Four marks: ok, kept, unknown, failed. "unknown" is not a shade of success —
adoption will meet configuration it cannot parse and state it cannot read, and
folding those into "fine" is the same move as reporting an installed package as
a capability.

And adding severity reopens something the earlier rule did not cover. "Flags
inform, they do not block" was decided about CONFLICTS, where the mesh chose
deliberately and the machine still works. A failure is not "we chose" but "we
could not". Treating both the same makes a node where something the mesh needed
never happened indistinguishable from one where a log level differed.
2026-08-26 21:55:08 +02:00
jschoubben 0106318bcb 012: on conflict, keep the machine's configuration
Reversed by the operator, and both directions are recorded because the reasoning
for each is the useful part.

What is already on the machine stays, the conflict is flagged, adoption
completes. This buys non-destructiveness by construction: the class that made
the opposite rule dangerous — a storage driver against the filesystem it is
actually on, a data directory pointing at a mount that exists — cannot arise,
because nothing tied to the machine's physical reality is overwritten.

It exposes the mirror. The mesh's configuration is not only preference; some of
it is what a module needs to function. Keeping the machine's version there
produces a module that is installed and does not work, which is 04-ISSUES/007
arriving from a direction that issue did not anticipate. And a fleet where every
node kept its own settings is one where a module works on one node and fails on
another with nothing able to say why.

So neither direction is right as a blanket, and the question is not whose
configuration wins. It is whether the module REQUIRES the setting or merely
PREFERS it — required contradictions cannot be kept without breaking the module,
preferences should always yield to what is there.

That is a property of the module's declaration rather than of the adoption
algorithm, which makes it one more thing the graph would carry. Until modules
can say which of their settings are load-bearing, adoption is defaulting in the
dark, and the default chosen is the one that does not break the machine it is
adopting.
2026-08-26 21:40:25 +02:00
jschoubben bcb18c7329 012: on conflict, install the mesh's version
Decided by the operator. Where the existing configuration and the mesh's
disagree, the mesh's version is installed, the conflict is flagged, and it is
reconciled afterwards — the mesh's configuration is known to work, the machine's
is not, and a half-adopted machine is a state nobody understands.

So adoption always completes and flags inform rather than block, which also
settles what 'adopted with open questions' prevents: nothing. The node is a
node. The original is kept, so nothing is unrecoverable.

One class left open rather than folded in, because it is the one place the
oldest rule in this record argues the other way. 'Known to work' is true of the
mesh's configuration in isolation, not on this machine. Most disagreements are
preference and overwriting them is right. A few are tied to what is physically
present — a storage driver against the filesystem it is actually on, a data
directory pointing at a mount that exists — and installing ours there does not
discard a preference, it can make existing data unreadable. Restoring the
configuration file afterwards does not undo that.

The default is settled. The exception is not 'there is a conflict' but 'applying
ours would destroy something a configuration backup cannot restore', and
identifying that class is open.
2026-08-26 21:37:04 +02:00
jschoubben 60736199a4 012: keep the original, and flag what cannot be decided
Two additions from the operator, and the second answers a question this effort
had open with two bad answers.

Nothing is taken over without keeping what was there. Adoption happens on
machines somebody is already using, and the configuration being taken over is
configuration somebody chose. This is a never rule rather than a courtesy, and
it earns that by the same incident the mesh's strongest rule carries: the worst
loss in this record came from a tool acting on a path it did not own. Adoption
is that act made deliberate, which makes the safeguard obligatory.

And adoption produces a briefing, not just a result. It meets things a script
cannot decide — a runtime configured one way against a mesh wanting another, a
package pinned for a reason, local settings the mesh has no opinion about.
Silently winning is wrong in both directions and refusing outright makes a
machine in use unadoptable. So conflicts are FLAGGED: what it found, what it
took over, what it could not resolve, written to be read by a person or an agent
as the first thing a session on that node has to work with.

That is the declaration parser's principle at a larger scale — name every
problem at once, to somebody who can act on it.

The question it turns on is recorded rather than assumed away: are flags
advisory or blocking? A briefing nobody opens is worse than a failure, because
the machine is in service and the record says it went well — 04-ISSUES/003
again. Working position: the node is usable and the mesh KNOWS it has unresolved
adoption questions, as a state something can ask about rather than a document in
a log directory. What that state prevents is undecided.
2026-08-26 21:33:09 +02:00
jschoubben ddb8091f68 Research 012 — the minimum viable node, and adopting what is already there
Building tier 0 reached a wall that looked like a packaging problem and is not.
The host can be told to run a container or install a package; both need a file,
and asking where the host gets it produced a bad trilemma — carry everything,
download at apply time, or push the files in first. Downloading fails on the
first node, which cannot fetch the image registry from the image registry it is
trying to start.

The reframing came from the operator: the machine is not offline, and what
matters is WHEN the fetching happens. Move it from apply time to build time —
build the installer on a machine with a network, tailored to the target, apply
it on a target that then needs nothing. The same move the lab already made for
its router image.

Which makes the question not where artifacts come from but what is missing from
THIS machine, and that needs two things answered: the closure for a one-node
mesh, and how a machine already in use becomes one.

Adoption is the second half, and it is sharper than it sounds. Having a package
installed is not owning it: a container runtime found already present carries
settings somebody chose, and noticing the binary exists discovers none of them.

It was also the original path — 00-as-is/05 records adoption of a pre-existing
machine's configuration as the original mechanism, since made legacy and
explicitly out of scope for the lab. It returns for a different reason than it
was dropped for.

Two collisions recorded rather than discovered later. ADR 0004 has managed files
generated and never edited, and adoption needs a one-time import before that
rule starts applying — three states, and the middle one is new. And ADR 0043
says the host never touches what it did not create, which is exactly what
adoption does; that rule needs a companion rather than an exception.

Eight open questions, including whether 'tier' is just a coarse view of a graph
level, whether owning a package means owning its version, and what cannot be
precomputed at all — because tailoring moves the cost of building from source
rather than removing it.
2026-08-26 21:30:43 +02:00
jschoubben b9facf9375 Design the node host
Playbook 02 step 3, on four recorded decisions. Tier 0 has one job — apply
declared state on this machine — and the six absorbed concerns are instances of
it, not additions to it.

Specifies the six parts and what each owns, and the two properties that make
apply trustworthy rather than merely present: every applier reads back, because
setting a value is not evidence the value took; and what was applied is recorded
after it works, never before, because a failed apply leaves the machine wherever
it reached and nothing must claim otherwise.

Build order is staged so each stage is verifiable in the lab before the next
exists. Stage 1 is profile and inventory — no control plane, no declarations, no
network — and it is deliberately the smallest useful thing, because `place:` has
nothing to place and the lab therefore raises empty machines. Stage 1 ends that,
and every later stage is tested by a lab that already works.

Stage 2 is the one that could invalidate the tier boundary: whether one host can
raise the substrate alone is Move 1's assumption and has never been proved.

Every decision the design rests on is given the test that asserts it, per 0034 —
including the dependency-direction lint, which is what makes "the host never
queries the mesh database" a rule rather than an intention.

Six things left open and named, including the one that host-size.md could not
measure: zero dependencies, but still six vocabularies.
2026-08-26 00:10:24 +02:00
jschoubben 902739acb6 Research 011 — the module graph
The proposal to split modules into provisioning services and applications was
worked through and abandoned, for a reason worth keeping: it cannot be filed
consistently. A git forge is consumed as a service and operated through a web
interface; an analytics service grants tracking identity and is a dashboard.

The operator's correction is the sharper form — what runs on the machine is a
supervised container, not something a user started. That is a fact about HOW a
thing runs, not about what kind of thing it is. So it is a facet, and 0002
survives: everything is a module.

What the catalogue is missing is not a taxonomy but a graph. Grouping asserts
relationships; a graph records them. Five declarations, of which two exist:
requires/provides a resource (yes), requires/excludes another module (no),
requires a node capability (no). Plus interface modules that carry no
implementation, with adapters providing them.

Recorded because it matters: this is a package manager's model, and pacman
already has all of it — depends, conflicts, and provides as virtual packages,
which is exactly the interface/adapter idea. Arriving there independently is
evidence for the shape. It is also a warning about what not to reimplement.

Working position on capabilities, to be tested: intrinsic ones (hardware,
architecture, network position) are detected and never installed, and a module
requiring one it lacks is impossible rather than unresolved. Provided ones (a
display server, a container runtime) are not a separate kind of thing — they are
modules that provide a capability, so "may the mesh install a capability" is not
policy, it is dependency resolution. Issue 007 then bears directly: an installed
package is not a capability.

Also captured: the operator's assessment that the machinery around a module —
scheduled tasks, hooks, migrations, config and env — is worth keeping, seeds are
not, and the integration is wrong enough to need a major refactor. Research 005
found supporting evidence from another direction, that the densest apparent
coupling in the catalogue is manifest boilerplate churn.

The first open question is the one that decides whether this is progress: what
does the graph DELETE? If modules gain declarations and lose nothing, it is
motion.
2026-08-25 23:19:57 +02:00
jschoubben 72b22830f3 ADRs 0036, 0037, 0038 — what a node is, what the host does, how one joins
0036 (accepted): a node is a managed machine, and disconnection is a situation.
The open question posed a class distinction — full nodes and lesser presences.
There is none. Reachability is state, not kind, which promotes the host's local
store from a component to a requirement: it is what makes disconnection ordinary
rather than exceptional. The reduced contract the question reached for is real
but it is capability, and that belongs in the profile.

0037 (accepted): the host applies, it does not decide. Measured rather than
argued — the absorption is smaller than the machinery that already applies
state, and eight of ten adapters carry no dependency to move. The two that do
open a Postgres connection to the control plane, which inside tier 0 is the one
thing the tier rule exists to forbid. So each concern splits: deciding needs
every other node and stays in tier 2; applying needs root and locality and goes
to tier 0. The host carries ONE concern, of which the six are instances.

0038 (proposed): a node joins by linking first. The operator's two-modes
proposal, adopted as intent and corrected as structure. Two modes is two code
paths where the first runs once per mesh and rots — and the mesh already has
that fault in its worst form, as three hand-run shell scripts. Instead: one
behaviour, two sources of declaration. The first node is not a different kind of
node, it is a node whose mesh is not up yet, and its specialness is temporary
and self-erasing.

0038 also shrinks the migration 0037 called expensive: a joining node never
needs mesh-wide state, because the hard part of the overlay is only needed to
compute the WHOLE mesh. It needs one peer. The rest arrives.

Left open and said so: what may be pushed over the link and how a joining node
proves it is entitled to join, and whether one host can raise the substrate
alone.
2026-08-25 10:34:45 +02:00
jschoubben 42bce02bba 006: answer the host-size question by measuring it
The skeleton's biggest unproven claim was that absorbing six concerns makes a
binary whose whole argument is having no dependencies carry six of them.

Measured against origin/main, and the question turns out to ask about the wrong
axis. By size the absorption is SMALLER than the machinery that already applies
state on a node — 2755 lines of adapters against 3059 lines of meshware,
env-sync and config-sync. The host is not a new large thing; it already exists,
spread across three core modules.

The real risk is direction, and it is two modules wide rather than six concerns
wide. Eight of ten adapters already receive derived state and only apply it, so
absorbing them moves code that has no dependency to move. Two — wireguard and
traefik — open a Postgres connection to the control plane and compute their own
configuration, which inside tier 0 would be an upward dependency and is exactly
what the tier rule forbids.

And the split has already been happening without being named: dnsmasq-app needs
the same node data as wireguard and does not query for it, because hand-
duplicated state went wrong and someone derived it centrally instead. Eight of
ten adapters are on the far side of that migration.

So the absorption is not a move, it is a split: deciding stays in tier 2,
applying goes to tier 0. The claim survives with its scope corrected — the host
carries ONE concern, apply declared state on this machine, of which the six are
instances.

Stated open rather than glossed: the two unsplit modules are the two hardest,
six concerns is still six vocabularies even at zero dependencies, and what the
host must carry versus find is issue 007 and unresolved.

Question B also recorded as answered by the operator — a node is a managed
machine, and a disconnected node is still a node in a different situation. The
question posed a class distinction; there is none, and what varies is state.
2026-08-25 02:20:32 +02:00
jschoubben 4bf7a35568 Close the record on the lab
Playbook 02 and 04 were followed for the substance — decisions before design,
design before build — and skipped for the bookkeeping. This closes that.

004 graduates. Its one open item was "not yet stood up"; the lab is stood up,
and the substitution the effort turned on is now enforced by the validator
before anything is raised rather than left as a thing to remember. Its
certificate conclusion has a home in 01-end-to-end-testing and is designed but
not built — implementation is a third axis, and an effort graduates on its
conclusions.

One item leaves 004 without a home and is recorded rather than lost: the reverse
proxy does not set caServer, so it defaults to the production endpoint.

The two lab designs read `designed` while running in production of a sort, so
they become `in-progress`.

And the lab gets an as-is document, which it did not have. It records what runs
including the parts nobody would choose again: that `place:` is refused and the
lab therefore raises EMPTY MACHINES, that the drawing shipped with no design
document behind it, that a router is tagged as a machine for a reason found by a
bug, and that the integration suite raises two of five scenarios while both
faults found so far lived in the three it does not.

006 stays active, deliberately. Two of its open questions ARE the tier 0 design
— whether absorbing six concerns makes the host too large, and whether an
unprivileged node earns a place in the inventory. Playbook 04 is explicit that
an open question is a reason to research, not to build around.
2026-08-25 01:55:52 +02:00
jschoubben e88b448145 The fix is real: 76x, verified. And how the lab installs on a clean machine
Snapshot 9.9s -> 0.13s. Restore 10.4s -> 0.80s. Three snapshots sharing
1.36 GB instead of costing 4.8 GB. The projected four-machine reset cycle
falls from ~90s, unbounded at worst, to ~15s dominated by a boot that
cannot be avoided. ADR 0029's inner-loop argument holds with copy-on-write
and did not without it.

The consistency matters as much as the speed: three consecutive snapshots
took 0.13, 0.12 and 0.13 seconds, against a dir second snapshot that never
finished.

One honest counter-observation recorded: launching onto the fresh
copy-on-write pool was slower, 20.2s against 14.3s, because the image had
to be unpacked into a pool that had never seen it. Paid once per pool, and
dwarfed by what snapshotting saves, but it went the other way.

Doing the measurement produced the answer to how the lab installs on a
clean machine, because both failure modes appeared while doing it.

Installed is not available: the daemon was present with units disabled and
no group. Issue 007.

Available is not adequate, and this is worse: with the storage tooling
absent everything worked and snapshots were seventy-six times slower.
Nothing failed, nothing warned. That is a variant the mesh has not
catalogued — its usual failure is reported success and did nothing; this is
reported success and did it seventy-six times slower, which no error
surface catches because nothing is wrong.

So the lab verifies CAPABILITY, never installation, and refuses to run
degraded rather than warning — a warning about a slow inner loop is read
once and ignored forever. Prerequisites may arrive from a mesh module or
from the lab's own bootstrap, and the second path is required rather than
convenient: a lab installable only by a mesh cannot host the development
of the mesh that installs it.

The lab is the second thing installed by hand, after the node host, and for
the same reason: something has to be first, and pretending otherwise
produces a circularity papered over by a script nobody exercises.
2026-08-24 00:14:57 +02:00
jschoubben 98bcd5cc49 Measure the lab's inner loop — it is too slow, for a fixable reason
The lifecycle design closed on a question that was measurable rather than
arguable, so it was measured. One virtual machine on a workstation with
hardware virtualisation and NVMe.

Raising: the launch call returns in 3.4s, the machine is actually usable
after 14.3s. The gap is a design constraint — raise must wait for the
second number, because reporting the first would be transport reported as
effect, which is the mesh's own recurring failure.

Snapshot: 9.9s and 1.6 GB for a 1.5 GB instance. A dir snapshot is a full
copy; nothing is shared. Restore: 10.4s, usable again after 20.1s.

The second snapshot exceeded two minutes and never completed. That is the
more troubling number: snapshot cost here is not merely high, it is
unpredictable, and a loop with a variable multi-minute step is one nobody
trusts.

Projected to a four-machine scenario, a reset-and-rerun cycle is about a
minute and a half at best and unbounded at worst, before any of the mesh's
own work begins. That is too slow for an inner loop, and ADR 0029's whole
argument — that making the bootstrap path the inner loop turns the
least-exercised code into the most-exercised — holds only while resetting
is cheap.

The cause is not virtual machines. Hardware virtualisation is present and
machines boot in fourteen seconds. It is that the daemon offers exactly one
storage driver, dir, which has no copy-on-write and therefore no cheap
snapshot. The btrfs kernel module is available; btrfs-progs is simply not
installed, which is the entire reason the driver is absent.

The copy-on-write comparison was deliberately NOT run, because running it
would mean installing a package by hand — which the rules forbid and which
would have made the measurement unreproducible. So the honest statement is
that the current configuration is too slow and the likely fix is known but
unverified, rather than that btrfs fixes it.
2026-08-24 00:07:22 +02:00
jschoubben a72fea5342 ADR 0031 and the scenario declaration
The lab provides the underlay; the mesh builds the overlay. This is the
boundary that decides whether the lab is worth having: a scenario that
assigns overlay addresses, elects the hub and writes peer configuration
certifies its own work — if the mesh's peering is broken, that scenario
still comes up green. The most valuable thing the lab can test is exactly
the part pre-building would replace.

So a scenario declares what a hosting provider and a home router would
provide: segments, which machine sits where at which address, what NAT is
between them, which ports are forwarded, which machines are detached. It
declares nothing about overlay addresses, hubs, peering, names or
certificates, all of which become outcomes to observe.

The declaration has four parts — segments, machines, place, snapshot — and
the two scenario classes differ only in place. That is what makes one a
strict subset of the other rather than a fork.

Research 004's most important finding becomes a format constraint rather
than a footnote: the routable segment must use RFC 5737 documentation
space, because the mesh decides public versus private by matching the
address, and a private range there makes the hub test as unreachable while
the mesh silently never forms. A segment without behind: is routable, and a
non-documentation address in it should be refused before anything is
raised — ADR 0008 applied to a configuration file, since the failure it
prevents has no error at all.

Four things left open, including the one that matters most: a lab machine
is always privileged, so the user and edge profiles have no scenario that
exercises them.
2026-08-23 22:33:52 +02:00