15 Commits
Author SHA1 Message Date
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00
jschoubben e1febe8e0f Renumber the records 1 to 23
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18,
19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only
the archaeology of what used to be there.

Renumbered contiguously. Renames run in ascending order, so every target number
is already free and no two files ever collide.

The reference rewrite is one simultaneous pass rather than a sequence of
replacements. Numbers moved into slots other numbers were vacating -- the node
host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time
would have cascaded and silently pointed things at the wrong record.

Seven plain-text references survived the merges as prose rather than links,
naming records that no longer existed: the enrolment token, the link boundary,
what a declaration is, reachability, the repository structure. Each mapped to
the consolidated record that now holds it.

Verified rather than assumed: every [ADR NNNN](path) link now has matching text
and target, checked across the whole repository, and the checker passes.

Frontmatter `consolidates:` lists dropped -- they named records that are gone,
and each consolidated record already says in prose what it absorbed.
2026-08-28 23:28:34 +02:00
jschoubben 77f3a4cea7 Consolidate: 65 decision records to 23
Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.

  the node host          8 -> 1    applies not decides, depends on nothing,
                                   per operating system, root service, the
                                   launcher, episodic, what a declaration is,
                                   actions from the bundle only
  a node and how it joins 4 -> 1   what a node is, joining, the link as
                                   security boundary, the enrolment token
  modules and the graph   7 -> 1   everything is a module, no domain modules,
                                   three edges, provisioning, the core library
  substrate and control   6 -> 1   the test, seven contexts, one control plane,
    plane                          the authority is not a database, the named
                                   products, the pinned bundle
  connectivity            3 -> 1   a route is a grant, reachability declared,
                                   filter rules
  delivery                5 -> 1   reconciliation not a pipeline, artifacts,
                                   the three silos, a failed step, the verdict
  the lab                 5 -> 1   (earlier)
  how this repository     10 -> 1  (earlier)
    works

Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.

The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.

The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
2026-08-28 20:03:24 +02:00
jschoubben 93470f6162 ADR 0047 — the bundle may carry actions the link may not
The bootstrap's sharpest open question, and the framing was wrong. "State on
this machine" was being read as the filesystem and the service manager. A
service running on this machine IS part of this machine — writing a file and
creating a database in a local store differ in mechanism, not in scope.

The real question was underneath: must the host learn what a database is? It
must not. Giving it a `database` resource type means tier 0 knows Postgres, then
a bucket, then a virtual host — the host acquiring the substrate's vocabulary
one service at a time, which is what ADR 0037 exists to stop.

So the bundle declares an ACTION and the host runs it and verifies it. What a
database means stays with the module that provides one; the host knows only how
to run a declared action against something local and check the result. Its
vocabulary grows by one shape rather than by one resource type per service.

Actions are permitted in the bundle and forbidden over the link, and the
asymmetry is deliberate. A bundle arrives WITH the binary: anyone able to put a
hostile action in it could equally have put it in the host itself, so refusing
actions there buys nothing and costs the bootstrap. The link is a separate
party, reachable separately, and an action there is the unbounded blast radius
ADR 0039 refuses. That decision stands unchanged.

And ongoing provisioning is not the host's at all — the control plane does it
once a mesh exists — so the asymmetry costs nothing.

Which dissolves the earlier worry about one mechanism with a tier boundary
inside it: there are two mechanisms, with different actors, scopes and trust
models, and that is the answer rather than a compromise.

Named rather than hidden: this is the escape hatch research 011 warned about,
arbitrary code in the place hardest to remove later. It is bounded by being
bundle-only and by every action having to declare how it verifies itself, and
that boundary is the whole defence.
2026-08-26 23:56:39 +02:00
jschoubben a4ab3e15c2 011: one interface, many contexts — and the constraint that hides in it
The objection is right: if every context runs its own service with its own
interface, the board is coupled to N interfaces instead of N schemas, something
has to compose them, and composition is logic — which tier 3 says a surface does
not hold. That moves the problem up a layer rather than solving it.

The skeleton already answers it, and the previous entry talked past it. `work`
and `knowledge` are not separate services; they are contexts INSIDE the control
plane, alongside the record, inventory and delivery — and `api` is listed there
as the one interface every surface speaks to.

So the board speaks to one interface. Behind it the contexts stay separate,
integrating through the record, but they are one tier, one repository, one
deployable — and coupling within a tier is not what the tier rule forbids. The
problem does move up a layer, and the layer it moves to already exists and has
this as its job.

The caveat is load-bearing and now recorded as an open question: this holds only
while the contexts are not separate deployables. The moment one becomes its own
service with its own interface, the board is back to N clients, something must
compose them, and the composition has nowhere to live that tier 3 permits. That
is a real constraint on how far the control plane may be split, and it is worth
knowing before splitting rather than after.
2026-08-26 23:29:56 +02:00
jschoubben a4a25ca7e3 011: one surface over several contexts is normal
The board visualises the mesh, the work engine, the knowledge base and more, and
the alternative — a web application per context — is worse for everyone using
it. Composing several sources into one view is what a surface IS, so this is not
a compromise with the ownership rule.

What changes is only where it reads from: each context's interface rather than
each context's store. Most of that already exists — 56 of 126 modules carry a
tool surface, more than carry a service.

And the unified board is what keeps those interfaces honest. A view that cannot
be built from a context's interface proves the interface inadequate, discovered
where it is cheap to notice rather than the first time something else needs the
same data and quietly reaches for the store instead.

If composing many calls proves too slow, the answer is a projection the board
owns and keeps current from events, not access to somebody else's tables.
2026-08-26 23:27:42 +02:00
jschoubben fa62c7f0e4 011: correct the rule — contexts, not processes
An earlier version argued a dashboard reading a dozen stores was caught by
exclusive ownership, because a dashboard is a surface and surfaces speak to an
interface. Wrong, and it drew the line in the wrong place.

The mesh's own board showing nodes, modules and deployments is not a separate
context reaching across a boundary — it is the mesh showing its own data.
Requiring it to go through an interface to reach facts its own context owns is
ceremony.

The rule is that a CONTEXT is granted what it exclusively owns. Everything
inside it — service, surface, tools — reads that store freely. What is forbidden
is a different context reading it.

Which is what the consumer count already showed: the problem was never surfaces,
it was three other contexts keeping their tables in the mesh's database.
2026-08-26 23:25:46 +02:00
jschoubben e71d532c2e 011: request or subscription is derived, not chosen
Asked what the distinction actually is, and the SQL half needed correcting
first: under exclusive ownership SQL runs against your own database and nothing
else, whatever transport a query might travel over. Both options are the mesh's
own channel and both ride the broker, so the transport is not the distinction.

The distinction is where the answer lives when you need it. A request asks at
the moment and waits — always current, costs a round trip, cannot answer when
the other side is down. A subscription keeps a local copy — instant, works
offline, as current as the last event received, and you must handle what you
missed.

What decides is not taste. ADR 0036 makes disconnection an ordinary situation
rather than an exception, so anything that must keep working while disconnected
CANNOT use a request: there is nobody to ask. And the converse — anything where
a stale answer is worse than no answer cannot use a subscription. A display can
lag; a decision about whether a grant is still valid cannot.

So an apparently open question turns out to be derived from a decision already
taken. What stays open is narrower: what a consumer does about the events it
missed while disconnected — replay from a point, ask once for a full picture and
resume, or rebuild. The question every projection has.

Also recorded: separate databases are required in the new design, and the shared
registry is a leftover rather than a pattern.
2026-08-26 23:23:39 +02:00
jschoubben afcc355744 011: checked the registry's real consumers, and the question was the wrong shape
The exclusive-ownership rule turned on whether every reader of the mesh registry
could be served another way. Eighteen consumers open a direct connection. Four
groups, and only one is work.

The owner and its machinery keep reading, because they own it. The node appliers
are already resolved — ADR 0037 stops the host querying the mesh database,
decided for tier reasons with nothing to do with this.

The bulk are FOREIGN TENANTS. The work engine holds ten of its own tables in the
registry's database, the knowledge base two, pipeline logs one. Thirteen foreign
tables across three contexts, which is how-we-build §4's shared schema counted.

So the question was the wrong shape: the problem is not readers needing a new
route to data, it is tenants needing to move out. Tasks, agents and teams have
nothing to do with nodes and modules and are co-located by history. Give that
context its own database and its dependency on the registry shrinks to one
table.

A handful of genuine cross-context reads remain, small enough to enumerate
rather than estimate. The rule holds.

Left open: whether those reads want an interface or events. Asking which nodes
exist at the moment you need to know is a request; reacting when a node appears
is a subscription, and some consumers want both.
2026-08-26 23:16:13 +02:00
jschoubben 6b1aab6a1e 011: the dashboard case, and why exclusive ownership is the tier rule
Raised as the hardest test of the rule: a board showing nodes, modules,
pipelines, agents and tasks wants to read a dozen stores, and under exclusive
ownership it can read none of them.

It survives, and not by luck. The board is a SURFACE, and surfaces already may
not do this — the skeleton puts `api/` in the control plane as the one interface
every surface speaks to, and tier 3 as thin, no logic. A board reading stores
directly is a surface reaching past the context that owns the data, which the
tier rule forbids for reasons that have nothing to do with databases.

So it is not a counter-example; it is an instance the rule catches. And the two
rules turn out to be one rule seen from two sides: exclusive ownership is the
tier rule expressed in terms of storage.

The general shape for anything needing to see across many things: consume the
record and own your own view. A reporting context builds a projection from
events and reads its own store, never anybody else's.

The cost said plainly rather than buried: a projection is more work than a join,
and it lags. A board queries the mesh's own database directly today — ordinary,
working — and this rule makes that a migration rather than a preference. The
reason to pay it is §4's already-measured cost, not elegance.
2026-08-26 23:10:36 +02:00
jschoubben aa767d17a8 011: a module is granted only what it exclusively owns
Reconsidered by the operator — maybe shared databases should not be allowed at
all — and the stricter version is better and goes further than the schemas it
replaces.

No shared writes, and no read-only role on another module's database either.
Reading another context's tables couples you to its layout exactly as firmly as
writing them does, and the coupling is harder to see because nothing breaks
until the owner changes a column.

That is how-we-build §4 taken at its word rather than at its letter. The
permissive version — a per-consumer schema, revocable, with cross-context joins
possible but deliberate — kept the letter and left the temptation. A boundary
that is merely inconvenient to cross is a boundary that gets crossed.

The cost is cross-module reporting, and it is the point rather than a
regrettable side effect: anything wanting to know what several modules hold
consumes their events or calls their interface. That is §4's whole argument, and
the mesh already has both mechanisms. What gets harder is precisely the thing
that was making work belonging to one context keep having to be implemented in
another.

And it is the first clear instance of what this effort has been hunting — what
the design DELETES rather than adds. Grant kinds collapse to one: an exclusive
resource. With them go the question of who owns which table, the guessing at
revocation time, cross-module migration ordering, and a class of permission
modelling a shared store would otherwise need.

One thing it does not answer, recorded because it could make the rule
unworkable: the mesh's own registry is read directly by many things today, and
under this rule they consume events or call tools instead. Achievable in
principle. Whether EVERY current consumer can be served that way is unchecked,
and should be before this becomes a decision.
2026-08-26 23:09:44 +02:00
jschoubben 4c8515507a 011: where a binding lives follows the scope, and a grant is not always a whole resource
Two questions asked directly, and the second collides with a rule in force.

One module on two nodes sharing a database corrects something stated flatly: the
binding is not "recorded on the assignment". Where it is written down FOLLOWS
THE SCOPE. A shared grant belongs to the module and every assignment references
the same one — which is the answer for two nodes wanting one database between
them. A per-instance grant belongs to the assignment. Same relation, two homes,
and which home is what makes two instances share something or not.

Several modules adding their own tables to one database is three needs wearing
one sentence, and a provider offers KINDS of grant rather than one: a database
for a consumer whose tables are nobody else's business, a read-only role for one
that needs to see what another holds, and a SCHEMA within a shared database for
the case actually asked about.

Loose tables in a shared database is what how-we-build §4 warns against in as
many words — several domains sharing one forty-five-table schema, which is why
work belonging to one context keeps having to be implemented in another. Not a
style objection; the observed cost, already paid.

A per-consumer schema keeps what the request wants and drops what §4 objects to.
Same database, same connection, same backup, and a cross-schema read remains
physically possible when genuinely needed. What it adds is ownership: migrations
touch one namespace, two modules cannot collide over a table name, and revoking
drops the schema rather than guessing which tables belonged to whom.

So the fault §4 names is still possible and no longer accidental — a
cross-context join becomes something somebody deliberately writes rather than
the path of least resistance. And revocation becomes answerable, which the
whole-database version never was.
2026-08-26 23:08:32 +02:00
jschoubben fa9889536c 011: tools have a different audience, migrations cross the edge, provisioning is early
Three additions, and the third kills an assumption.

Tools are the most common content in the catalogue — 56 of 126 modules, more
than carry a service — and they survive the split without fitting either half. A
tool is not an artifact and not node state; it is a contract the mesh publishes
on a module's behalf, and what consumes it is an AGENT rather than another
module. That is a second audience the design has not described. Whether it is
one relation with two audiences or two relations is cheap to decide now and
expensive later.

A migration belongs to the CONSUMER and runs on the PROVIDER. A game's
migrations run against the database the store granted it: owned by the consumer,
hosted inside something it does not control, ordered after the provisioning edge
because there is nothing to migrate until the grant exists, and scoped to that
grant. Ownership crosses the edge, which nothing in provides and requires
expresses — and it gives a consumer's own install an internal order, provisioned
then migrated then started, that depends on an edge rather than on its contents.

And provisioning is EARLY, not late. The assumption worth killing is that it is
something the control plane does for consumers once a mesh is running. The
mesh's own registry database is provisioned before there is a mesh, and so is
its virtual host on the broker: the store runs from the carried bundle, a
database is created in it, the mesh's own schema is applied, and only then does
a control plane exist. Steps two and three happen before there is a mesh to do
them, so provisioning is part of the bootstrap and part of what the bundle has
to express.

Which strains ADR 0043. The host applies declared state ON THIS MACHINE, and a
database inside a running store is not a file or a unit. At bootstrap it is at
least local — the store is on the same machine. Afterwards a consumer on one
node provisioned from a store on another is the ordinary case and reaching it is
not the host's job. The same operation is local at bootstrap and remote later,
which is either two mechanisms or one with a tier boundary crossing inside it.
Currently the sharpest unresolved thing in the effort.
2026-08-26 23:02:52 +02:00
jschoubben a3c7e7e1f1 011: providing is a facet, and the assignment is a third thing
Any hosted service can be a factory — an identity provider grants clients, an
analytics service grants a tracking identity, a mail server grants mailboxes, an
application platform grants a project that is several of those at once.
Providing is a FACET a module may have, not a kind of module it is, which is the
same conclusion this effort reached about services and applications arriving
from the other direction. So `provider` stops being a category too.

Two relational stores from different vendors both grant "a database" and are the
sharpest possible test of the substitutability rule. They fail it completely —
different protocol, dialect, driver, client library compiled into the consumer —
so `database` stays a tag, now with two real providers rather than a thought
experiment.

The assignment is a third entity, recorded because the operator tried the
alternative: modules were once node-agnostic and it did not survive. Several of
a provider's properties belong to neither end — where its state lives, how it is
reached, tuning derived from the machine's hardware, which instance serves a
given consumer. Not the catalogue, because they differ per node; not the node,
because they are about this module. A design with only modules and nodes has
nowhere to put them, which is what node-agnostic ran out of. The current system
already stores environment values per module AND per node, arriving the same
way.

Which answers the question asked directly: two nodes both run a store, so which
serves a consumer? Neither obvious answer. Not the consumer naming a node — that
is placement in the consumer's manifest, a game edited because a database moved.
Not the consumer not caring — for presence it genuinely does not, for
instantiation it cares permanently.

What the consumer knows is the SCOPE of its own need: one instance shared across
every instance of itself, or one each. That decides, and needs no node named.
Then the mesh binds, and the binding is recorded on the assignment and is
sticky — a resolver that re-derives which store serves a consumer will one day
derive a different answer and relocate a database.
2026-08-26 22:56:36 +02:00
jschoubben 13c6068874 011: the provider shape generalises, and two things differ inside it
The broker has all nine properties the store has. So do the object store and the
image registry. A substrate service is a SERVICE PLUS A FACTORY, there are four
of them, and the pattern generalises past the substrate: anything granting
something per consumer has this shape.

Two differences matter more than the similarity.

The broker cannot be managed over the broker. ADR 0001 makes it the channel
every node takes work from and ADR 0039 makes it the security boundary, so the
module providing it is also the way modules are managed — a declaration cannot
be delivered to it over itself. Nothing else has that property; the store is
consumed by the control plane but is not how the control plane REACHES anything.
This is what the carried bundle exists for: the broker is raised from what the
host carries because there is no other way to raise it. A constraint on one
module, not a general rule, and a schema with no way to say so hides it.

And two modules of identical shape want opposite instance counts. The broker is
one per mesh by decision. The store cannot be, because a node that must keep
working while disconnected cannot depend on a database elsewhere. Which settles
what cases.md left open: how many instances is NOT derivable from what a module
is. It is a per-module decision, it has to be declared, and nothing in provides,
requires or excludes says it.

Revocation differs in consequence too. Dropping a database leaves data until
something removes it — a leak, recoverable. Dropping a virtual host loses
whatever was undelivered — silent, and not. Same relation, different blast
radius, which argues for the provider deciding what revocation means rather than
the mesh applying one rule.

File renamed: it was never really about postgres.
2026-08-26 22:54:49 +02:00