Author SHA1 Message Date
mesh-admin 017b1401e4 Merge pull request 'ADR 0135: a module version prepares its state before it runs' (#161) from decision/0133-0134-migrations-and-deploy-facts into main 2026-09-28 09:57:47 +00:00
jschoubben 476cda417d ADR 0135 supersedes 0133: a module version prepares its state before it runs
Two faults in 0133, both caught on review. It put the declaration on a container — one resource kind
the host applies — so every author would restate the machine's arrangement and a module's own
lifecycle would be tied to how its artifact happens to run. A module declares entrypoints for its
tools and its provisioner; preparing its state is the same vocabulary and nothing about a runtime.

And it derived the scope from the machine, which the facts already answer: a consumer is a module on
a machine (issue 022, migration 0015), so what the mesh provisions is per consumer. A module on three
machines has three databases, there is no shared state to race over, and the lock obligation 0133
invented was for a situation the mesh does not produce. The level question HAL answered with stages
dissolves — the scope of preparation is the scope of the state, and the mesh knows it.

0133 keeps its reasoning and gains a pointer; design 32 and issue 133 name the live record.
2026-09-28 11:57:16 +02:00
mesh-admin 74ba3ff1d4 Merge pull request 'ADR 0133 and 0134: who runs migrations, and the mesh saying what it applied' (#160) from decision/0133-0134-migrations-and-deploy-facts into main 2026-09-28 09:45:49 +00:00
jschoubben 891c7a945e ADR 0133 and 0134: who runs migrations, and the mesh saying what it applied
0133 — a module owns its migrations and the mesh owns when they run. A container declares what must
run before it; the mesh derives the gated step from the resource it precedes, so the image, the
environment and the credentials come from the one place they are described. The module owns the SQL,
the dialect and the lock; the mesh owns the moment and refuses to start a version whose step failed.
Per node, with no level: a step that ran once somewhere leaves every other machine ungated, and
'once, mesh-wide' is what holding a seat already means.

0134 — the pipeline is observable from a merge to an artifact and goes dark at the machine. What a
node now runs, and what it refused, become facts under the control plane's own seat, emitted when
what a machine runs changes rather than on every convergence pass.

Design 32's lifecycle carries both; issue 133 points at them as what ends the matter it opened.
2026-09-28 11:45:47 +02:00
mesh-admin 83cbeb8db8 Merge pull request 'Issue 133: the control plane's schema is migrated at birth and never again' (#159) from issue/133-the-control-plane-migrates-before-it-serves into main 2026-09-28 08:27:32 +00:00
jschoubben ab6db9369b Issue 133: the control plane's schema is migrated at birth and never again
The mesh replaced its own control plane with a build carrying a migration, applied none of it, and
then recorded no build for three quarters of an hour while saying everything was fine. ADR 0052
already prescribes the shape — a run-once step that gates the server — and the control plane was the
one module that did not use it.
2026-09-28 10:27:30 +02:00
mesh-admin 84571b4825 Merge pull request 'ADR 0132: a seat carries the tools its holder must serve' (#158) from decision/0132-a-seat-carries-the-tools-its-holder-must-serve into main 2026-09-28 08:17:01 +00:00
jschoubben d57196102d ADR 0132: a seat carries the tools its holder must serve
A role's tools belong to the role, not to whichever module holds it today: the seat declares them
with their schemas, serving them is a condition of occupying the seat, and what the mesh can do
becomes a read of its own records rather than a question nothing answers. A module keeps its own
tools — the same module may run without the seat, and then only its own name is true.

Design 33 follows: the three families, addressing a node-scoped seat, discovery, and what serves
this to an agent.
2026-09-28 10:16:59 +02:00
mesh-admin e6402cf777 Merge pull request 'Issue 132: a module can be recorded without the directory it lives in' (#157) from issue/132-a-module-can-be-recorded-without-its-directory into main 2026-09-28 07:20:07 +00:00
jschoubben 4f0d144833 Issue 132: a module can be recorded without the directory it lives in
Nine modules could not be rebuilt: their record named the repository and no directory, so every
build looked for a manifest at a repository root that has never had one. Resolved by mesh-controller
— `module add` takes the directory and the forge, and the rule is checked rather than described.
2026-09-28 09:20:05 +02:00
mesh-admin 98d94ef71e Merge pull request 'Design 28: 5.5 done, the mesh has one bus; issue 131 resolved' (#156) from design/28-one-bus-issue-131-resolved into main 2026-09-28 01:59:41 +00:00
jschoubben 4e13280604 Design 28: 5.5 done, the mesh has one bus; issue 131 resolved
The AMQP transport is gone from the control plane and the hosts (mesh-controller #112,
mesh-host #39). On the way: no build had ever recorded its bases, so every bases-first order
walked an empty graph; the builder now reports what it was handed and the graph is read from
builds (mesh-controller #113/#114). Issue 131 is resolved by the forge module's merge event,
the control plane following it, and those edges.
2026-09-28 03:59:39 +02:00
mesh-admin 8783a13448 Merge pull request 'Design 28: the mesh runs on the new bus' (#155) from design/28-the-mesh-runs-on-nats into main 2026-09-28 00:40:29 +00:00
jschoubben a31cfcf461 Design 28: 5.3 is built and was used for the hand-over 2026-09-28 02:28:06 +02:00
jschoubben 75b3861911 Design 28: the mesh runs on the new bus
Tasks 4.3, 5.2 and 5.4 are done as of 2026-09-28 02:25: every machine reports on
the new bus, the seat is held by the module that provides it, the old broker is
unassigned and forgotten, and every credential was minted afresh at the end.

5.2 records what it took, in the order it was found and each fixed on the trunk
before the next step, and how the bootstrap loop was broken once, by hand.
2026-09-28 02:27:49 +02:00
jschoubben 694555214a Merge pull request 'Design 26: which assignment holds a seat is on record, and changes as one act' (#153) from design/26-a-seat-is-held-on-record into main 2026-09-27 21:22:56 +00:00
jschoubben a7249541df Design 26: which assignment holds a seat is on record, and changes as one act
Until now the holder was derived — assigned and claiming — and a second eligible
assignment was refused, so a seat could not pass from one holder to the next
without a moment where nobody held it. The controller finds its own bus through
one of these seats, and that moment took the control plane down on 2026-09-27.

The holder is now a row the controller keeps, written by `seat <name> --to
<node>/<module>` in the same write that removes the previous one. No row means the
old rule, so nothing changes for a mesh that never hands a seat over; with a row,
another eligible assignment is silent rather than refused, which is what lets the
next holder run beside the current one until the switch. A holding is the
assignment's and goes when it does. Each rule names the test that checks it.

Under ADR 0131; design 28 task 5.3 is the work.
2026-09-27 23:20:36 +02:00
jschoubben 95d8253f71 Merge pull request 'ADR 0131: everything on the mesh speaks to the broker seat, and AMQP is not a provision' (#152) from decision/0131-everything-speaks-to-the-broker-seat into main 2026-09-27 21:06:02 +00:00
jschoubben 784b487bf9 ADR 0131: everything on the mesh speaks to the broker seat, and AMQP is not a provision
Taken during the outage of 2026-09-27, when the protocol leaked into the seat's
contract: to hold mesh-broker a module had to provide amqp, so the module that
will carry the bus could not hold the seat that names the bus, while the module
being retired could. Supersedes 0127. Modules depend on the seat and reach the
bus through the sdk; no manifest provides or requires amqp; the old broker's
module and the two modules that required it leave the catalogue; the AMQP
transport is deleted once every node reports on the new bus.

Design 28 step 5 rewritten under it: the seat handover becomes its own task and
is built first, because the seat the control plane dereferences cannot be empty
in between — that emptiness was the outage. The cost note now carries what was
measured rather than what was assumed.

0128 and 0130 extended 0127; each now rests on 0131 with a dated note and
changes nothing it decided. Every other citation of 0127 names its replacement.
records.py still fails on 0120/0112, which predates this branch.
2026-09-27 23:03:05 +02:00
jschoubben 7ae711ba0b Merge pull request 'Building the bus: the decisions the work needed, and what it taught back' (#150) from feat/nats-genesis into main 2026-09-27 17:06:40 +00:00
jschoubben 51739c4302 The order the repositories land in is part of the rollout
Derived while merging and visible from no single repository, so it belongs
written down rather than re-derived later: the client library before the
catalogue, because it is where the subject is derived and a converted module
against the old one publishes the local name itself; the catalogue before the
controller, because the controller refuses an old-style name outright and would
make every unconverted module unregisterable.

Two tested properties are what make it safe and not merely ordered. An old-style
name passes through the derivation untouched, so an unconverted module keeps
working at every step. And a converted name derives to exactly the key the old bus
published, so nothing moves on the wire until 5.2 sets the variable.

The failure avoided is 127's own, which is why this is worth a table: a publisher
and a subscriber disagreeing about a subject log nothing anywhere.
2026-09-27 19:04:45 +02:00
jschoubben 5cac268457 Two checks were wrong about where a seat is judged
Correcting design 26 to match what the merged code does, and fixing issue 112's
status, which used a word the vocabulary does not have.

A claim on a seat the manifest does not itself declare is refused at registration
rather than by the parser. A module may hold a seat another module declared —
that is why ADR 0126 has a caller name the seat and not its provider — so whether
the name exists is a fact about the whole catalogue.

And a declared seat may promise nothing. That is a marker seat, and most
node-scoped seats are markers: which module is this machine's packet filter. ADR
0126's "a declared seat carries a protocol" says what a holder must satisfy, not
that every seat offers something.

`records.py` still fails on ADR 0120 resting on a proposed ADR 0112, which is not
this branch's and not mine to decide.
2026-09-27 18:57:49 +02:00
jschoubben ce6ae943b7 Merge main: renumber this branch's records around the trunk's
Both lines of work numbered from the same point, so four decision records and one design
document existed twice with different content. The trunk keeps its numbers and this branch
yields — the only rule that scales, because the trunk's are already cited by what merged
before them.

  0117 the bus is the only broker        -> 0125
  0118 a module declares its own seats   -> 0126
  0119 amqp is a provision, not the bus  -> 0127
  0120 the mesh bus is required          -> 0128
  0123 a seat carries its role's protocol -> 0129
  0124 the predecessor is ending          -> 0130
  design 29, what a module declares       -> design 32

Applied to the code repositories too, because a stale reference is worse when numbers
collide than when they dangle: the reader lands on a real record that decided something
else.

Two reconciliations the merge forced, both real:

**0110 was marked wholly superseded and was not.** Its successor says in as many words that
everything 0110 decided about what a seat *is* stands untouched — and two records that
landed on the trunk rest on exactly that part. So it is accepted again, extended rather than
replaced, with a note saying which of its claims moved and where.

**A seat's protocol becomes columns, not fields.** The trunk moved the seat set out of
compiled code into a table the controller owns. This branch had added what a role accepts,
emits and serves to the Go slice. The decision is unaffected and the mechanism is better for
it: giving a role a protocol is now a write rather than a rebuild, which is the trunk's own
argument applied to what this branch added.

One check still fails and it fails on main too: a record resting on ADR 0112 while that is
still 'proposed'. Left alone — it is not this merge's to answer.
2026-09-27 18:23:41 +02:00
jschoubben 9ef9830dcf ADR 0122: the predecessor is ending, and its broker goes with it
ADR 0119 rejected giving the old broker a retirement condition and said why: "its
clients are not only the predecessor's, so the retirement condition describes a day
that will not come". The operator has said that day is coming — the predecessor is
deprecated, some of it still running, none of it being migrated, left to stop rather
than moved.

Recorded because three documents reason from the premise it overturns. Design 25 §5's
"no day anything is waiting for", §9's "the predecessor's clients never notice", and
design 28's closing note that the predecessor's world does not need to move.

**And it needs no new machinery, which is 0119 being paid off rather than revised.**
Because that record made the broker an ordinary provider rather than a compatibility
module, ending it is unassigning a provider whose provision nothing requires — something
the module system has done since it existed. So step 5.3 finishes instead of trailing
off, and the transitional double announcement of a build outcome has a date.

The consequence worth planning around: the predecessor's own mesh talks over that
broker, so shutting it down ends the tooling that reaches this installation's machines
from a workstation. The rollout is driven from the node, or driven before the broker
stops. That is a sequencing constraint on 5.2, not an afterthought.

What survives is `amqp` as a provision: a module that genuinely needs an AMQP broker can
still be given one. What retires is this broker's role as the predecessor's.
2026-09-27 18:11:11 +02:00
jschoubben dfac01a6dd 5.2: the readiness half is in, and what a failed move actually costs
`rollout check` answers from records whether this mesh could move its bus, and names the
next step for each thing missing. The move itself waits on that check having been run
against a real mesh — writing the irreversible half before its question has ever been
asked of something real breaks the plan's own rule about beds by another route.

And the cost of being wrong is written down rather than assumed: the old broker stays for
its other clients, nothing in a served request's path goes over the mesh's own bus, and
what a failed move costs is the mesh's ability to change things rather than the services
its modules serve.
2026-09-27 17:59:26 +02:00
jschoubben 51f5ce5c0f 4.5 done: the catch-up needed nothing built, which was the answer
Three ways to do the replay were weighed and the right answer was that the bus being
moved to already does it. A queue on the old bus receives only what is published after
it is bound, so everything built before the catalogue existed was announced to nobody. A
stream is a log and a consumer is a position in it: a consumer created later starts at
the beginning and the builds are simply there. Checked against a running server, because
the decision rested on it.

So "who replays" has no answer because nothing replays. The mechanism was never about
builds — it was about a queue that could not remember, and carrying it across would have
carried a workaround for a limitation that no longer exists, with nothing looking wrong.

Retiring it belongs to step 5, with the rest of what only the old bus needs.
2026-09-27 17:23:49 +02:00
jschoubben fc64a2c1a4 4.4 done: a person's account and their client
The account existed as a permission model and as nothing a person could be given; there
is a record and three commands now. The client is two surfaces over one thing, a command
line and an MCP server, both using the client a module's runtime uses — so what a person
may do is answered by the same permission list that answers it for a module.

Design 25 §7 says nothing of this is built before its bed passes, and this was built
before. Noted in the task rather than quietly ignored.
2026-09-27 17:07:28 +02:00
jschoubben 9fc4e74b0d Merge pull request 'to-be 31: a module declares its fail2ban jail, mesh composes them per node' (#149) from design/a-module-declares-its-jail into main 2026-09-27 15:00:16 +00:00
jschoubben 225dfa9451 to-be 31: a module declares its fail2ban jail, mesh composes them per node
A node's intrusion filter should be composed from its assigned modules, like
its firewall (the Filtering mechanism): a service module (postgres, mssql,
mailu) declares its jail in its manifest (filter + stanza, no node/path per ADR
0112), and the mesh writes the jails of a node's modules into the fail2ban
holder's jail.d. The base (sshd, recidive, ignoreip=mesh-range) stays the
fail2ban module's. Records the model after novox's HAL per-module jails were
lost as dangling symlinks; the ignoreip is now safe on disk, the service jails
need this to be restored.
2026-09-27 16:59:56 +02:00
jschoubben bfefb1dbdb 4.3: the installer can raise a mesh on the new bus
A foundation template that stands the server up, writes its settings and the mesh's
first user list beside them, and starts a controller on the new bus. The first user list
is the installer's because at genesis there is no mesh to compose one — a bootstrap
credential, rotated like the store's.

The carried list is checked against what the controller derives, since a mesh cannot be
raised twice to discover they disagreed. That check immediately found the composer
granting a role's whole event branch as well as the one event it follows.

What is left of 4.3 is running it, which is 4.1's bed.
2026-09-27 16:39:48 +02:00
jschoubben 85c0a3e567 4.2 done: a build is work submitted to a role, on both buses
Both sides behind a seam, one implementation per bus, and the outcome is the role's
own event so one publish reaches the asker, the controller and the catalogue. Checked
against a running server, including the part the decision rests on: a third party
hears the same outcome the asker does.
2026-09-27 16:02:17 +02:00
jschoubben a39765c924 Merge pull request 'ADR 0122: a seat is data the controller owns; a rename is a database update' (#148) from design/a-seat-is-data into main 2026-09-27 13:41:10 +00:00
jschoubben a8921fe737 ADR 0122: a seat is data the controller owns; a rename is a database update
Reviews 0110/0121 after a session where renaming seats cost three freezes, a
builder deadlock, and hand-resolved manifests. The seat rules were right; the
set being a compiled Go slice referenced by name-string everywhere was the
mistake. Seats become a table keyed by a stable id; claims/held/production code
reference the id; a rename is one UPDATE, no rebuild, no re-registration, no
freeze. The build machine reads the set from the mesh instead of embedding it,
removing the controller/builder seat coupling. Closed set and scope naming
unchanged; only storage and reference change. Outstanding renames (registry
seats, private-network scope) wait for this — as data each is a write.
2026-09-27 15:40:55 +02:00
jschoubben c8f430290e ADR 0121: a seat carries the protocol of its role
The mesh's own seats said who does a job and nothing about what may be said to
them or by them, and that gap showed up three times in one day looking like three
different problems: a build machine with three audiences for one outcome and no way
to derive a grant for any of them; an event genuinely about a role with nowhere to
live but the namespace of whichever module holds that role today; and a catalogue
catching up on builds, where every option needed a grant the design refuses.

One cause — the mesh has roles it cannot describe. So the `mesh-*` seats take the
same three fields a module's seat has, and the machinery that already derives
authority, queues and consumers from a declared seat does it for these too.

Builds become work submitted to a role, and `mesh.build.request`,
`mesh.control.built` and the BUILDS stream retire. A work queue shared by several
build machines is exactly what a seat's `accepts` is, so a second mechanism for it
was two places a permission could be wrong. The outcome is the seat's own event,
which means one publish still reaches whoever asked, the controller that records it
and the catalogue that places it — the fan-out a shared exchange gave for free,
written as a subject the mesh derived rather than a topology somebody configured.

That also avoids the grant that ruled out the alternatives: no holder needs
permission to publish into an asker's inbox.

The blocking gap is now named rather than incidental: the shared library has no way
for a module to publish on a seat. The build machine is Go and reaches the bus
directly, so it is unaffected; the artifact-store event waits.
2026-09-27 15:24:17 +02:00
jschoubben d98d6fca11 Merge pull request 'to-be 30: the mesh updates itself on a push' (#147) from design/the-mesh-updates-itself into main 2026-09-27 12:49:32 +00:00
jschoubben b8cdfce16d to-be 30: the mesh updates itself on a push
Records the manual update process (module moved -> build -> reconcile; and the
breaking-change freeze/re-register recovery), and the two things that make
self-update more than a webhook: the build-on-push trigger is currently HAL's
(hal-gitea-tools on :9877), a retirement gap the mesh must replace with its own
forge-webhook trigger wired to every repo including mesh-controller; and the
builder validates manifests too, so a breaking change couples controller +
builder + manifests + hosts, and renaming the builder's own seat deadlocks its
rebuild. Names the transition discipline (accept old+new for one release) that
self-update needs so a push does not auto-freeze.
2026-09-27 14:49:02 +02:00
jschoubben c4a8455e2e Issue 127 resolved; design 29 says what wildcards are and how the rule is checked
Every module named its events the way the old bus spelled a routing key, so on the
new bus every cross-module subscription pointed at a namespace nobody publishes to.
Nothing failed — the services started and none of them reacted. Converted, and the
rule now has checks at both scales: at registration for one manifest, and as a test
across the whole catalogue where a consumed event's emitter is present.

It was larger than the report said, in two directions nobody had looked. Forty-three
files of module code pass the event name at runtime, so the code mattered as much as
the manifests. And both clients had to learn the mapping — without that, converting
the modules would have broken the mesh that is actually running, which is the
opposite of what fixing this was for.

Design 29 gained three things it did not say: what a wildcard is (`*` for one name,
`**` for the rest, spelled the mesh's way and derived to each bus's own), that an
event about a role belongs on the seat and why that is not yet possible, and how the
rule is checked — because "a subscription that matches nothing is silence" is exactly
why nobody noticed thirty-seven manifests being wrong the same way.

4.2 and 4.3 are unblocked. The catch-up half of 4.5 is not: it is a decision, and it
narrowed rather than closed. It cannot be a reply to a module's inbox, because that
needs the blanket grant design 25 §4 refuses.
2026-09-27 14:45:08 +02:00
jschoubben bf7297b7ed Merge pull request 'ADR 0121: keep distribution, retire only verdaccio; node-* seats migrated' (#146) from design/0121-keep-distribution-fix into main 2026-09-27 12:36:54 +00:00
jschoubben 1bb0ef5658 ADR 0121: keep distribution, retire only verdaccio; node-* seats migrated
Records the reversal: distribution stays as the mesh's OCI registry (it serves
every artifact-store:// image); only verdaccio, a redundant second npm registry,
is removed. The 'consolidate onto gitea / retire distribution' direction was
dropped. Also records that the node-* rename was executed as one controlled
migration with a brief compose freeze, and why the delivering registry seats
are deferred rather than folded in.
2026-09-27 14:36:33 +02:00
jschoubben 5a9917f802 Merge pull request 'ADR 0121: a system seat is named for its scope; a module may define its own' (#145) from design/system-seats-are-named-by-scope into main 2026-09-27 12:15:22 +00:00
jschoubben 8a6ee9177c ADR 0121: a system seat is named for its scope; a module may define its own
The control plane's seats grew a second naming style (the-*) beside mesh-*,
and the closed set was the only place any seat could be defined. This settles
both: system seats are mesh-* (one, mesh-wide) or node-* (one per node), named
for scope; a module may define its own seat outside the closed set. Folds in
the seat review: mesh-build-machine (scope fix), mesh-private-network (one
server + client modules, dropping per-node VPN choice), showcase becomes the
first module-defined seat, node-uplink, and the node-* renames — plus the
registry consolidation onto gitea, which reshapes the registry seats and gates
retiring distribution/verdaccio. Records why the renames are a coordinated
migration and why distribution cannot be removed until gitea serves images.
2026-09-27 14:15:05 +02:00
jschoubben dd577e9ebe 1.7 is done; 4.1 waits on nothing but the bed
Minting on both halves, the file delivered per push, and the bus's objects
asserted on every start — verified against a real server that asserting twice
changes nothing, that a machine joining an already-raised bus is accepted, that
each node's consumer is bound to its own declaration subject, and that CONTROL
does not dead-letter before the controller gives up.

Which bus the mesh is on is one fact, and being told about both is refused at
start rather than warned about: a mesh half on each is one where a declaration
goes out on one bus and the report comes back on the other while every component
logs success — ADR 0074's failure arriving through configuration rather than code.

So 4.1 no longer waits on code. Both links speak NATS, the composition happens,
and every claim behind them has a unit test or a check against a running server.
What none of those can stand in for is a mesh raising itself, which is what the bed
is — this is where the code stops and the lab starts.
2026-09-27 03:20:42 +02:00
jschoubben a9b8f41570 Design 25 §4: the server verifies no client certificate, and writes only accounts
Two corrections of fact, both found by building the module's image and connecting
to it as a host would.

The first composed configuration said `verify: true`, which makes the server
demand a *client* certificate — and nothing in the mesh presents one. A host pins
this server's exact certificate and authenticates with the password the mesh
minted, and so does a module's runtime. Every connection in the mesh would have
died at the TLS handshake before any password was looked at, with an error that
reads as a fault in the client. TLS is still required; verify only decides whether
client certificates are checked. Mutual TLS is a later question and would need
machinery the mesh does not have — a certificate per module per node.

And §4 read as though the controller wrote the whole file. It writes the user list
and nothing else: ports, TLS paths and a store directory belong to the container
the module raises. The two files share one directory of necessity, because an
absolute include path is resolved relative to the including file's own directory.

The decision stands in both cases — accounts are composed, not called for, and
passwords are minted and sealed. What changed is what the file says and who writes
which half.
2026-09-27 02:51:40 +02:00
jschoubben 92a5e8fc05 Merge pull request 'ADR 0120: a roster fact carries its format as a template; rewrite to-be 29' (#144) from design/roster-fact-is-a-template into main 2026-09-26 23:51:21 +00:00
jschoubben 8c91ba1cfa 1.7 in progress: the list is derived and the keys are kept
What is in: a bus user's hash is recorded and its plaintext returned once, and
the user list is derived from the machines, what each runs, every manifest and
which machines hold a live token. Permissions stay derived rather than stored,
because a stored copy could disagree with the records it came from while both
looked internally consistent.

What is out, with what each needs, so the next person does not rediscover it:
delivery, which has one open question about what a module declares in order to
receive the file — design 29's ground, not this document's; minting, which is
transport-coupled because an enrolment reply carries one password and a node on
the old bus must not be handed a credential for the new one; and calling the
assertions from a start path.
2026-09-27 01:50:04 +02:00
jschoubben 0f7f628730 ADR 0120: note the shared/region interaction with hq 128
A roster fact may be shared — written into a marked region of the machine's
file (into: block, hq 128) rather than as the whole file. The template
renders the content; shared decides how the host lays it down. Composes with
hq 128: the region mechanism is the host's, the format is the module's.
2026-09-27 01:42:00 +02:00
jschoubben 970da74136 1.7: the composer exists and the composition does not
Correcting a tick and a claim I made one commit ago. 4.1 does not wait on an
enrolment user per live token; it waits on the whole composition, of which that
user is one input.

Tasks 1.3 and 1.4 are honest about what they built — the composer, the
derivation, the permission model, the stream and consumer definitions, the
asserter, all pure and held by unit tests and a golden composition. Nobody wrote
the caller. Measured: outside the package that defines them there is not one use
of the composer, the permission derivation, the stream set, the stream asserter or
the principal type. Step 1's "done when" claims every account and permission
composed from the manifests, and a mesh raised today would stand up a server with
no user list at all.

It also needs state the mesh does not keep. Design 25 §4 says the file holds
bcrypt hashes and that passwords are minted and sealed exactly as today — but
today the mesh mints one, hands it to the broker through a management call, seals
the plaintext to the holder and keeps nothing. With no management call the hash
has to survive every later recomposition, because the first thing a new module or
a person's access change touches is a file that must still hold every other
user's password. No bcrypt hash is stored anywhere in the controller.

Named as its own task rather than folded into 1.3, so the gap between "the parts
of step 1 exist" and "the mesh does any of it" is visible.
2026-09-27 01:39:20 +02:00
jschoubben 4d4012cdf6 ADR 0120: a roster fact carries its format as a template; rewrite to-be 29 around it
The facts mechanism formatted the roster in Go in the control plane — one
formatter per fact, in the consumer's own configuration language. ADR 0120
makes a fact a path and a template: the mesh owns the data, the module owns
the format, and the control plane holds no format at all.

to-be 29 (operator accounts + what lives under a home) is rewritten to ride
it: the ssh files become roster templates, the whole ~/.ssh is owned with a
found/owned boundary that cannot lock the operator out, keys are mesh-owned
through an SSH CA (existing keys adopted not regenerated, the operator's
personal key signed not minted), and the ssh-agent is a user-scoped service.
2026-09-27 01:33:12 +02:00
jschoubben 0a61531c42 WBS 3.5 done; 4.1 waits on one thing, an enrolment user per token
All three halves of the host's link are through seams, and the reply address in
the payload is now proved from both ends rather than one — the test asserts the
transport's own field held the consumer's ack address by the time the request
arrived, so a server that stopped claiming it fails a test instead of letting the
reason become folklore.

Two things had to be built for the host to hear anything at all: a node's
declaration consumer, which only the controller may create, and the enrolment
user's inbox, which design 25 §6 names and the composer granted none of. Both
were silent gaps — a node with either missing looks correct and hears nothing.

What remains is a single piece: something that composes an enrolment user per
live token. On the old bus that account is made imperatively through the broker's
management API; here there is no management API, so issuing a token has to
recompose the server's configuration. It is the only thing between the two links
and a mesh raised on NATS from nothing, so 4.1 now says so.
2026-09-27 01:32:26 +02:00
jschoubben 4b0b659084 Merge pull request 'ADR 0119: a taken tunnel's predecessor is retired once the take is proven' (#143) from decision/0119-a-taken-tunnels-predecessor-is-retired into main 2026-09-26 22:59:15 +00:00
jochen 0042ca9258 0119: link ADR 0118 now that it is on main 2026-09-27 00:58:55 +02:00
jochen 63c19456b4 0119 review: rollback needs the private network unassigned first; a configuration written back is retired again with the first original kept; only the interface's own file, never a link 2026-09-27 00:58:54 +02:00
jochen 2f195d501e to-be 08: the found tunnel's configuration is retired once the take is proven (ADR 0119) 2026-09-27 00:58:54 +02:00
jochen 8a78ff4efe ADR 0119: a taken tunnel's predecessor is retired once the take is proven 2026-09-27 00:58:54 +02:00
jschoubben 762300a380 Merge pull request 'issue 130: undeclaring a service stops it, even one the mesh only reloads or keeps running' (#142) from issue/130-undeclaring-a-service-stops-it into main 2026-09-26 22:58:43 +00:00
jochen 248c99ca6c 0118/130: link ADR 0117 now that it is on main 2026-09-27 00:58:17 +02:00
jochen 13208f0f48 0118: give a unit back the state it was found in — never-stop broke the converge rollback; process removal found and fixed 2026-09-27 00:58:17 +02:00
jochen ebd19c4c6c ADR 0118: undeclaring removes what the mesh made, gives back what it changed, leaves the machine's units as they are — resolves issue 130 2026-09-27 00:58:17 +02:00
jochen dd4cbabffb 130: ADR 0117 named, not linked, until it is on main 2026-09-27 00:58:17 +02:00
jochen 5b76a09da6 issue 130: undeclaring a service stops it, even one the mesh only reloads or keeps running 2026-09-27 00:58:17 +02:00
jschoubben a363a605cb Merge pull request 'issue 129: nothing makes a machine trust the mesh's own certificate authority' (#141) from issue/129-nothing-makes-a-machine-trust-the-meshs-authority into main 2026-09-26 22:57:30 +00:00
jschoubben 24835ab710 Merge pull request 'issue 128: the machine's hosts file is written whole, and on a workstation it is shared' (#140) from issue/128-the-hosts-file-is-written-whole into main 2026-09-26 22:57:02 +00:00
jschoubben ba397d4cbe Merge pull request 'ADR 0117: a machine's uplink is a seat — the mesh configures the manager, never the link' (#139) from decision/0117-the-uplink-is-a-seat into main 2026-09-26 22:56:27 +00:00
jschoubben f555d523c7 WBS 3.4 is done both halves; issue 127 holds 4.2, 4.3 and catch-up
The controller's inbound is through a seam with both transports behind it, and
the store window is now the server's rather than the controller's memory. Seven
claims about that were asked of a running server rather than reasoned.

Wiring the controller's own subscription is what found issue 127: every event
name in the catalogue is still written the way a routing key is, so design 29's
derivation turns a consumer's declaration into a subject no emitter publishes.
Thirty-seven manifests, one that cannot be composed at all. It fails on the first
mesh raised on the new bus and not before, which is why nothing had caught it —
the conformance fixtures pin one emitter against one subject, and both halves of
that pair are correct.

The node-facing flows are unaffected: those subjects are the mesh's own and
derive from nothing a module declares.
2026-09-27 00:55:07 +02:00
jschoubben 4f93d304d7 WBS: a person's account is done; the client is not blocked by step 3 2026-09-27 00:18:05 +02:00
jschoubben d898bd87e8 Step 5.4 was wrong from 0119 onward; removed
It waited on a retirement condition 0119 abolished when it made the
deprecated broker an ordinary provider. A step waiting for a condition
nobody set would sit open forever.
2026-09-27 00:16:45 +02:00
jschoubben 7e4da874a9 Design 25 §2: the eaten reply address is verified, not assumed
A claim the whole enrolment handshake rests on, now measured against a
running server rather than reasoned from documentation — and held by a test
so it cannot become folklore if a server version changes.
2026-09-27 00:15:00 +02:00
jschoubben 53092020eb WBS: asking a tool is through the seam; a build is a different shape 2026-09-27 00:11:44 +02:00
jochen df4a3538c3 0117 review: a holder's service declares no state — the manager's lifecycle is the machine's; a start-only setting's gap on a fresh machine 2026-09-27 00:07:42 +02:00
jschoubben 00817fb9e3 Design 25: the store window, and what moving it into the server changes
The guarantee is the same and the mechanism is simpler — a nak with a
delay, no parked list, nothing lost when the controller restarts. It costs
one thing: a naked message comes back whatever happened meanwhile, so an
older report is redelivered after a newer was applied. A report already
carries the digest of the declaration it answers, so supersession becomes a
check rather than memory — ordering settled by what a message says, not by
when it arrived.
2026-09-27 00:06:15 +02:00
jochen 90b8eeff6b 129 review: the example name is a routed name, not a doubled suffix 2026-09-27 00:03:36 +02:00
jochen a865fc7d79 128 review: located; the fix as built (block, at, never held, order) 2026-09-27 00:03:30 +02:00
jochen c3313f6e17 0117 review: seat table row + decision cited, networkd/dhcpcd lines match the modules, block placement, dispatcher scope, references 2026-09-27 00:03:15 +02:00
jschoubben 9946e852e1 WBS: 3.5's outbound half is in 2026-09-27 00:01:56 +02:00
jochen 60a53f9b18 issue 129: nothing makes a machine trust the mesh's own certificate authority 2026-09-26 23:58:48 +02:00
jochen ee31f9f761 issue 128: the machine's hosts file is written whole, and on a workstation it is shared 2026-09-26 23:58:48 +02:00
jochen 0e066473b3 ADR 0117 accepted; a manager that cannot reload takes the setting at its next start (dhcpcd, measured) 2026-09-26 23:58:48 +02:00
jochen 504adef221 ADR 0117: a machine's uplink is a seat — the mesh configures the manager, never the link 2026-09-26 23:58:48 +02:00
jschoubben 9f6aa7ea9c The bus is the mesh's centre, not a transport that replaced one
Two things. A paragraph from the superseded 0117 survived beside the 0119
correction that reversed it, so §5 said both that the amqp interface
retires and that it does not. The stale one is gone.

And the framing. §1 opened with "the bus carries five kinds of traffic
today, and this design keeps the five", with a column mapping each to the
queue it used to be — which describes the mesh's nervous system as a port
of something that did a fraction of this. It now says what the bus is: a
role addressable without knowing its holder, the mesh's own state, work
that queues until somebody can do it, and permissions derived from what a
module declared. Conditions, observation and a person's client land there
too as they are built.

Glossary gains `bus` and `the deprecated broker`, with a note on why not to
say "compatibility broker" or name it after a protocol — the second invites
exactly the backwards framing this commit removes.
2026-09-26 23:54:09 +02:00
jschoubben 8deecf620a Merge pull request 'to-be 29: a node has operator accounts, and the mesh owns what lives under a home' (#138) from design/29-a-node-has-operator-accounts into main 2026-09-26 21:53:39 +00:00
jschoubben 64bbdc30c1 to-be 29: a node has operator accounts, and the mesh owns what lives under a home
The mesh models machines but not the people on them — a node record
holds no username, and no module places anything under a home. So who
you are on each node (jochens/ace/jochen) is unknown to the mesh, and
nothing owns ~/.ssh, dotfiles or ~/.config. HAL knew it; the nox mesh
dropped it. Proposes the account as a node fact and a home-scoped
resource class (the ~/ mirror of ADR 0112's /var/lib placement), with
the login key staying the operator's (ADR 0051). Not urgent — HAL's
generators still run — load-bearing at node-by-node retirement. Found
generating ~/.ssh/config from HAL's registry, which nox has no
equivalent for.
2026-09-26 23:53:13 +02:00
jschoubben 672c994afa WBS: 3.4's seam is in, outbound half through it 2026-09-26 23:47:40 +02:00
jschoubben 2f9bb73685 WBS: the first fixtures are in, and what byte-for-byte means 2026-09-26 23:41:15 +02:00
jschoubben 5c193b3f54 WBS: 3.8's check is written 2026-09-26 23:34:42 +02:00
jschoubben fbf9440e8e WBS: 3.7 and 3.8 done, with the one check 3.8 still owes 2026-09-26 23:34:04 +02:00
jschoubben 9510bf5311 WBS: 3.2 done 2026-09-26 23:33:10 +02:00
jschoubben d940e14ec8 Design 19: the protocol on NATS
Task 3.2. ADR 0074's model is untouched — floor plus capabilities, partial
implementations legitimate, identity from the credential, dedup on
x-event-id, conformance as executable fixtures. The transport beneath it is
rewritten: exchanges and queues become subjects and streams.

Statements marked *verified* were checked against a running server while
the runtime's client was written, not reasoned from documentation. Three
of them are things the specification would otherwise have got wrong:

- the payload is the body alone, with metadata in NATS headers; an
  implementation that nested the whole envelope would agree with nobody
- a durable name may not contain a dot, while the ack subject joins two
  names with one — conflating them looks right in a permission list and is
  refused as a consumer name
- a certificate must carry a name the bus is dialled by, because the NATS
  client has no hook to replace hostname verification the way pinning did
  on AMQP

And one limitation lifts: a module may now call another's tool. Issue 049
recorded that a scoped account could not declare the reply queue a caller
needs, and ADR 0095 routed every ask through the control plane because of
it. Per-account inbox prefixes plus allow_responses replace that. ADR 0095
is not reversed — the control plane is still how a person asks — but
module-to-module calling stops being a question about capability and
becomes one about policy, which `uses` already answers.
2026-09-26 23:32:48 +02:00
jschoubben 80456981be WBS: 3.6 done, and the certificate constraint it surfaced
The NATS client has no checkServerIdentity hook, so pinning no longer makes
the name check redundant — the bus's certificate must carry a SAN matching
the address nodes dial.
2026-09-26 23:29:56 +02:00
jschoubben d2ed3152d3 Seat renames done; 0118 was wrong that it was a migration
A holding is derived at resolution from manifests, never stored, so there
are no recorded old names to rewrite. The work is an edit plus a kept rename
table — kept because a module lives in its own repository and may be
registered long after the catalogue stopped using an old name.
2026-09-26 23:06:55 +02:00
jschoubben 24d99ddd24 WBS: 3.9 done, and 1.4's client with it 2026-09-26 22:28:59 +02:00
jschoubben b5730525fe Merge pull request 'issue 127: a declaration that shrinks to empty is skipped, not sent' (#137) from issue/127-shrink-to-empty into main 2026-09-26 20:25:19 +00:00
jschoubben bb334e138b issue 127: a declaration that shrinks to empty is skipped, so the node keeps what it should drop 2026-09-26 22:24:58 +02:00
jschoubben b759e36bfd Design 29: tools, not serves; WBS 3.9 partly done
The manifest already uses serves for a provision's facts, so a module's
tools take their own key. Declaring them is itself new — until now a
module's tools existed only in a runtime environment variable.
2026-09-26 22:17:17 +02:00
jschoubben 0b8e84334f Why module events share one stream, checked against the server
Storage is not a property of a subject — a stream is a separate object that
covers one — so the question is always how many streams, not which topics
are durable.

Three facts decide it, two of them verified rather than assumed: NATS
refuses overlapping streams instead of merging them, so a shared stream
plus a per-module one is not available at all; a filter cannot express an
exception; and a stream per module turns one cross-module consumer into one
per module. So one stream, with per-subject caps for the fairness that
matters. Per-module age is genuinely unavailable, and a module that needs
it declares a seat.
2026-09-26 21:49:55 +02:00
jschoubben 39c802cbd4 Step 2: adoption recreates the bus once, on purpose
2.3 was already true and is now proved — the seat refusal is generic, and
three tests pin what matters: a second bus is refused by name, a different
bus implementation is refused too (which is what makes the bus replaceable),
and the AMQP broker no longer contends so both run on one mesh.

2.1/2.2 turned out not to be a no-op. The host keeps a container only when
its spec matches exactly; genesis raises the upstream image and the module
declares the mesh-built one carrying the entrypoint, so assigning it
recreates the container. That is ADR 0067's pivot and it is safe only
because the bus carries nothing yet — which is why step 2 comes before
anything speaks NATS. After it, never again: the config is a directory
mount, so accounts change without touching the container's spec.
2026-09-26 21:40:02 +02:00
jschoubben 86a084b7ff The mesh bus is required, not ambient
Design 29 said no module requires the bus. The catalogue disagrees: 49 of
72 modules take a broker credential and 23 do not, so an ambient connection
mints an account for a third of the catalogue that never speaks — and the
49 each hand-write the path it lands at, which is provisioning done badly
by hand.

The bootstrap argument that made it ambient was narrower than it looked.
"A provisioner needs an account before it can run" is true of a provisioner
process and says nothing about a provision the controller answers, and the
controller is not waiting on a bus account to compose one.

So: the mesh-broker seat delivers mesh-bus; a module requires it and gets an
address, a sealed credential and the trust to verify the server; a module
that requires nothing has no account at all. The requirement delivers the
connection, the declarations shape the authority, and declaring a subject
without requiring the bus is refused as incoherent.

mesh-bus and nats are deliberately two names: a module may run its own NATS
as a backing service exactly as one provides amqp, and a manifest saying
"nats" would otherwise mean either the mesh's nervous system or a private
queue.

The seat's Delivers was wrong twice today — amqp, then empty — and the
comment says so rather than reading as though it were always right.
2026-09-26 21:17:21 +02:00
jschoubben 85f972749a AMQP is a provision, not the bus
0117 went a step further than it had grounds for. It was right that the bus
is the only bus, and wrong that the amqp interface must therefore retire —
because it conflated two reasons to want a broker. Using one to reach
another module is a second bus and stays refused. Needing an AMQP broker as
a backing service, the way something needs a database, is ordinary, and
forbidding it would make the mesh unable to run normal software while
calling that architecture.

So the broker becomes a plain provider module: no seat, not foundation,
never raised at genesis, no retirement condition. lavinmq now claims nothing
and provides amqp; nats claims mesh-broker and provides nothing.

The rule that survives is about direction, not software: inter-module
communication goes over the bus. A module may hold a broker for itself; it
may not use one as a channel to another module. That is a review judgement
where 0117 could have used a parser, which is the honest cost.

0106's progressive insight was itself wrong and is corrected by a second one
there — nothing moves off the old broker, so its "one purpose" sentence does
not become true, it is just not what that server is.

The insight check needed two fixes it found itself: a date may carry
trailing words, and a bold run with a link is discussing an insight rather
than marking one. All four bad shapes still fire.
2026-09-26 21:08:40 +02:00
jschoubben 7abb268de6 Step 1 done but for its bed
1.1 to 1.4 built and tested. 1.5 turned out to need no controller change:
it already resolves the broker by seat and names no broker module in its
source, which is what ADR 0079 was for. The genesis module set naming is
scenario and installer config, carried with the bed.

Recorded what must NOT change yet: the amqps:// credential shape and the
5671 default are correct until the rollout, because steps 1-4 leave every
node on AMQP.
2026-09-26 21:03:35 +02:00
jschoubben 78a2274baf Designs 25 and 29 disagreed about the subject space; implementing found it
29 put a module's events and tools in one namespace, 25 kept mesh.events.*
and mesh.tools.*. One namespace is right — a module's authority over its own
name becomes a single pattern the server enforces — but it needs a kind
token, because a stream is a subject filter and mesh.mod.*.> would persist
every tool call in the mesh. Tools stay on core NATS for the reason 25
already gives.

So: mesh.mod.<module>.event.<name>, .tool.<name>, and seats the same shape.
2026-09-26 21:02:18 +02:00
jschoubben 3f9b316015 Design 25: a scoped inbox needs allow_responses, or nothing can answer
Found composing the first real configuration. Scoping every inbox to its
owner is right and leaves a responder unable to reply, because the answer
goes to the caller's inbox. The fix is not a wider grant but the server's
own allow_responses: one reply to the subject of a message the user actually
received. Without it every tool call times out while the permission list
looks correct.
2026-09-26 20:58:34 +02:00
jschoubben e05825a881 Design 29: versioning, provisioning and secrets on the bus
Versioning: additive is free; a breaking change is refused while callers are
bound, and the refusal names them, because the mesh already holds the uses
graph; a real break versions the subject, not the seat name, so the role
does not fork; binding is a recorded pin, not a drift to whatever is newest.
Semantic change stays open — no fingerprint sees it, and saying so beats
implying the check is complete.

Provisioning: a provisioner's create/remove/holds IS a serves protocol, so a
provision interface is a seat that also delivers a credential — which is why
design 26 already allowed that. The per-consumer resource is what stops the
two collapsing into one.

Secrets: sealed, so the bus is never trusted with plaintext — but sealed is
not enough, because a stream persists and a durable ciphertext is an archive
the day a key leaks. So a secret never enters a stream: core request/reply
only, and a declaration names a secret rather than carrying one, which is
0098's fetch-don't-store applied where carrying is worst. The vault's own
credential and the bus's own accounts are the two bootstrap exceptions,
resolved the way 0067 resolves the control plane.

Also rewrote the addresses paragraph, which was too compressed to follow:
on-bus addresses disappear because nothing stores them, off-bus ones are
untouched and still 0098's problem, and the bus's own address is the one
that cannot be a subject.
2026-09-26 20:44:16 +02:00
jschoubben 7b4916e9ec Modules declare their own seats; the mesh reserves mesh-*
The architecture 0117 opened needs a module to offer a service as a role on
the bus — one holder, addressed by what it does. A closed table in the
controller cannot express that: a capability a module contributes would
require changing the mesh itself.

But 0110 closed the set for a good reason — nothing could say what seats a
mesh had, and the hand count came out at eleven of thirteen. That argues for
enumerable, not hardcoded, and 0110 weighed free-form against a fixed table
without considering a third option: closed at any moment and derived from
the catalogue. A derived list cannot drift, which is how the count broke.

So: the mesh's seats stay the mesh's, reserved by the mesh- prefix so the
prefix is the rule and there is no list to maintain; ten seats are renamed
to restore 0079's convention; everything 0110 decided about what a seat IS
survives untouched.

Design 29 carries the declaration model: three namespaces, subjects derived
from local names so a manifest survives the wire changing, queues never
declared, five relationships (the job and state shapes 0041 had no room
for), and the build-publish-deploy lifecycle with hard, soft and build-time
dependencies distinguished.

0041 gets a progressive insight: "no per-consumer setup, only a
subscription" was a fact about a topic exchange, and a JetStream durable
consumer is a real object someone creates.

WBS 1.3/1.4 were wrong and say so: streams come at registration and
consumers at assignment, so only the foundation set belongs at genesis.
2026-09-26 20:34:32 +02:00
jschoubben 3f14264c7b Merge pull request '112 is fixed: a carried peer is nameable, and the resolver answers for all four machines' (#136) from issue/112-fixed into main 2026-09-26 18:13:53 +00:00
jschoubben 25d599094f 112 is fixed: a carried peer is nameable, and the resolver answers for all four machines 2026-09-26 20:13:37 +02:00
jschoubben b5b68e8852 Design 25: a host directory bind, not a named volume
Issue 115 is resolved and converted four modules away from named volumes;
the bus's own data is not the place to bring one back. Also: NATS carries
TLS on the client port rather than beside a plaintext one, so there is no
5671/5672 pair to mirror.
2026-09-26 19:34:54 +02:00
jschoubben 814c9e563f The bus is the only broker; step 1 starts
A module does declare requirements the provisioner fulfils — but the broker
it gets that way is a private vhost, the analog of a database, not the
mesh's bus. Two modules of the new mesh depend on it, so the compatibility
broker was never single-purpose and its retirement would have stranded them.

NATS is the heart: one bus, a module's messaging is subjects on it scoped by
what it declares, and no module is handed a server of its own. The seat
delivers nothing; the interface retires with the broker. Also closes the
EVENTS question — one stream, on the bootstrap argument, not preference.

Designs 25 and 28 go in-progress: step 1 is starting.

The insight check caught a false positive on its own first real use — its
bold-run pattern crossed newlines and joined an unrelated `**` to the
marker. Constrained to one line, still catching all four bad shapes.
2026-09-26 19:26:07 +02:00
jschoubben db4ca9b043 Merge pull request 'The bus in five steps: a decomposition, a breakdown, and progressive insight' (#134) from feat/the-bus-in-five-steps into main 2026-09-26 17:05:43 +00:00
jschoubben 5a6d0e111d Merge remote-tracking branch 'origin/main' into feat/the-bus-in-five-steps
# Conflicts:
#	02-DECISIONS/README.md
2026-09-26 19:05:36 +02:00
jschoubben c94e2ece53 Merge pull request 'Give 0115 a home, and match its batch's status — main is failing both checks' (#135) from fix/0115-checks-on-main into main 2026-09-26 17:04:50 +00:00
jschoubben 0af6479b2c Give 0115 a home, and match its batch's status
Both repository checks fail on main. 0115 is cited by no design, and it is
marked accepted while resting on 0112, which is proposed.

Design 27 already states the rule the record decides — "a module is assigned
at most once to a node, and that pair is the assignment's identity" — so it
is the home, and now says so. And 0112, 0113, 0114 and design 27 are all
proposed: the batch is under review, so the record is too. Promoting 0112
instead would be marking a record accepted to satisfy a check, which
check_rests_on names as a failure this repository already made once.
2026-09-26 19:03:14 +02:00
jschoubben 77a1493df4 Renumber to 0116: another record took 0115 on main
PR #133 landed a different 0115 while this branch was open. The bus record
is now 0116, with every citation in designs 19, 25, 28 and the index
following it.

Note: cycle.py and records.py both fail on main as merged, on that record —
nothing cites it, and it rests on 0112, which is still proposed. Both
pre-date this branch and are left for their own change.
2026-09-26 18:56:34 +02:00
jschoubben 21b54ee52f Merge main: 0115 was taken by another record 2026-09-26 18:54:59 +02:00
jschoubben 3950c2b75d Merge pull request '0115: one assignment of a module per node — the requirement is dropped' (#133) from decision/0115-one-assignment-per-module-per-node into main 2026-09-26 16:53:12 +00:00
jschoubben 9516d31a62 0115: one assignment of a module per node — the requirement is dropped, the module's name is the identity 2026-09-26 18:52:59 +02:00
jschoubben 1c808898a5 Allow progressive insight, and apply two to the bus record
A record can assert a fact that goes stale while the decision it supports
stays right. Superseding for that buries a sound record under a second one
and makes every reader work out which is live. So a correction of fact is
now made in place, marked and dated, with the old wording quoted — bounded
by three conditions and checked by records.py, which fires on an unmarked,
undated or back-dated note. Judgements still supersede.

Applied to 0115: no conformance suite exists to recapture, and the full
genesis bed cannot run until the links exist. Designs 25 and 28 follow.
2026-09-26 18:39:38 +02:00
jschoubben 6ab113e6c6 Break the bus work down, measured, in dependency order
Counting the surface first changed the plan twice: the genesis bed cannot run
until the links exist, so it belongs to step 4, and there is no conformance
suite to recapture — step 3 builds one against the current bus before moving
it. Both corrections are recorded in the breakdown rather than edited into
ADR 0115. Also indexes design 25, which was never listed.
2026-09-26 18:24:14 +02:00
jschoubben fe0c1e9da2 Divide the bus work into five steps, each proved on its own
The NATS change was recorded as one undivided item, which hid three gaps:
a mesh already running had no adoption path, the protocol specification did
not know its transport was being replaced, and nothing was runnable until
everything was. Dividing it is what surfaced them.
2026-09-26 18:19:56 +02:00
jschoubben dc80116b09 Merge pull request 'to-be 27: the three gaps answered' (#132) from design/to-be-27-gaps-answered into main 2026-09-26 15:44:53 +00:00
jschoubben db142ffa26 to-be 27: the three gaps answered — root is a node setting, the mesh's writes need no module-visible reservation, resolution is the controller's at composition 2026-09-26 17:44:38 +02:00
jschoubben e38814962d Merge pull request 'Issues 125 and 126: two apply-layer gaps the novox session hit live' (#131) from issue/125-126-from-the-novox-session into main 2026-09-26 15:25:41 +00:00
jschoubben 7842457d4b Issues 125 and 126: two apply-layer gaps the novox session hit live
125: a hold is not a line in the apply report — sixteen resources held
for an untaken module while four surfaces reported success, and the
operator stopped the edge's predecessor on their word (the route-proxy
flip outage). 126: a changed volume path neither recreates a running
container nor warns, and a roll-out upgrade policy makes a build a
deployment — together they turned a data-path migration into a forge
outage (the /var/lib move). Filed as 119/121 in the novox session
before syncing; renumbered past the other session's 119-124.
2026-09-26 17:25:27 +02:00
jschoubben 782d5ace04 Merge pull request 'Issues 119 and 122: what a host path is, and what a node's layout would replace' (#130) from issue/119-122-what-a-host-path-is into main 2026-09-26 15:21:34 +00:00
jochen f709e8e0fb Issues 119 and 122: what a host path is, and what a node's layout would replace
Not every host path names this machine. A system file the mesh owns is at that path on every machine
of the kind — the path is the fact. The operator's shared data is already answered as an access. A
path inside a container is the software's contract. What is left, and what a node's default layout
would replace, is 514: a module's own data, and what the mesh writes for that module.

Documents the reservation model as the records already have it — a root per node, one directory per
assignment, a placement for adopted data — and the three things nothing states: where the root comes
from, what sits beneath it, and the order the 514 are retired in.

Stops there deliberately. Changing where a definition looks without moving the data does not fail: the
mesh creates the directory, the container starts, the service comes up empty. Retiring these is a data
migration with a verification step, module by module, and belongs with whoever can see the machine.
2026-09-26 17:21:14 +02:00
jschoubben c5e1a9a8d5 Merge pull request 'Issue 122: count host paths, not paths' (#129) from issue/122-host-paths-not-container-paths into main 2026-09-26 15:11:33 +00:00
jochen c4d9b515ea Issue 122: count host paths, not paths
A path inside a container is not a fact about the machine — /run/secrets and the directory a server
keeps its data in are the software's own contract, true in any mesh that runs it. Only the host side
of a mount names where it landed.

The first sweep matched path-shaped strings, so it counted both halves of every mount and every
in-container location a value mentioned: 798. Counted by role — directory and file resources, the
host side of mounts, accesses, and the targets of binds, grants, receives and secrets — it is 698
across 70 definitions.
2026-09-26 17:11:18 +02:00
jschoubben 1012fff607 Merge pull request 'Issue 122: count it properly — 30 of 71 definitions name this installation' (#128) from issue/122-the-census into main 2026-09-26 15:09:17 +00:00
jochen d5a4cb0d4f Issue 122: count it properly — 30 of 71 definitions name this installation
The five modules this report first named were what a first look found. A sweep of all 71: 49 public
domains across 26, the node's own name 58 times across 19, a routable IP 12 times in one, 798
absolute paths across 70 (issue 119's number, grown), and no email addresses at all.

The sharpest case is not a domain: a mail module states the node's public IPv4 as the address it
trusts a real-IP header from, so a node that moves or gains a second address stops attributing mail
correctly, silently. Two upstream resolvers are excluded deliberately — naming a public DNS service
is a policy default, true of any mesh, not a fact about this one.
2026-09-26 17:09:01 +02:00
jschoubben 5db79cd168 Merge pull request 'Issue 124: a consumer cannot be told a value its provider derived for it' (#127) from issue/124-a-consumer-cannot-be-told-what-its-provider-derived into main 2026-09-26 14:47:25 +00:00
jochen 146fd6b3a8 Issue 124: a consumer cannot be told a value its provider derived for it
The object store derives each consumer's bucket from the login the mesh minted, and never reads the
one a definition named. The consumer still has to tell its own software which bucket to use, and has
no way to be told: bound values come from the provider's serves, which is a literal block identical
for every consumer, and a provisioner returns nothing. So all three consumers wrote the answer down
by hand and one of them wrote the predecessor's bucket — a key scoped to one bucket and software
asking for another, which reads like a credential fault and is not one.

Records the general shape: any interface where the provider names the resource forces the consumer to
reproduce the provider's rule, kept in agreement by hand and checked by nothing.
2026-09-26 16:47:06 +02:00
jschoubben 007e3f4b03 Merge pull request 'Issue 123: the image registry is named after a role, and *artifact* is defined as one format' (#126) from issue/123-the-image-registry-is-named-after-a-role into main 2026-09-26 13:38:00 +00:00
jochen 3abb3a3c07 Issue 123: the image registry is named after a role, and artifact means one format
Three wordings disagree, and the confusion is the damage: the glossary defines artifact as an OCI
image while the build vocabulary already names four kinds in use, two of which are not images; the
image registry's seat is named after its job while ADR 0079 names foundation seats after their servers
and ADR 0109 names package seats after their ecosystem; and prose that says 'the module's image' reads
as though a module were an image.

Records the question the naming hides: ADR 0075 keeps two provisions because packages and images are
two protocols, and already allows the forge to provide the artifact store. The second implementation
rests on a bootstrap argument, and the forge has the same upstream-server shape the store and broker
have, which ADR 0078 raises as plumbing and adopts in place.
2026-09-26 15:37:41 +02:00
jschoubben 112963524f Merge pull request 'Issue 121: retract step 4 — the forge's service is not built' (#125) from issue/121-step-4-retracted into main 2026-09-26 13:36:36 +00:00
jochen a1d8b478ec Issue 121: retract step 4 — the forge's service is not built
Step 4 claimed the forge cannot exist as a container before the builder has built its image. The
forge's server is an upstream public image pinned by digest; only its runtime sidecar is built. The
store and the broker have the same shape, and ADR 0078 raises both at genesis as plumbing and adopts
them in place — so the forge can be raised the same way and serve git, packages and OCI before
anything is built.

What survives: a grant is minted by the provider's runtime sidecar, which is built, so the question
is whether raise-service, grant, build-sidecar simply works. Sequencing inside the mesh, not images.

The wrong version came from taking a record's bootstrap argument at face value instead of comparing it
to how the store and broker are raised — one command away in the manifests.
2026-09-26 15:36:19 +02:00
jschoubben 1489d17238 Merge pull request 'Issue 108: the second door was attached to the wrong thing' (#124) from issue/108-one-door-after-all into main 2026-09-26 13:22:55 +00:00
jochen c343fc68c5 Issue 108: the second door was attached to the wrong thing
Two things were treated as one. The mesh's own artifact store holds the store seat and is internal by
design — reached by name over the overlay, no accounts, ADR 0082. Serving a registry publicly is a
service the mesh can host: a module with its own name, accounts and storage, like anything else it
runs for somebody. The conversion this report was written beside gave the seat holder a second public
door over the same filesystem, which is neither.

Keeps the original issue whole — no garbage collection, and the settings a routine needs are not
enabled — drops the two-door complication, and sharpens one thing: deletion on the only door is
deletion on a door with no accounts, which the predecessor kept behind its authenticated one.
2026-09-26 15:22:09 +02:00
jschoubben 859ff49ff2 Merge pull request 'Issue 122: the pattern is already on main, in five modules' (#122) from issue/122-the-instances-already-on-main into main 2026-09-26 13:07:42 +00:00
jochen e399a2c148 Issue 122: the pattern is already on main, in five modules
Asked whether merging the three reviewed changes would set a precedent. It would not: five modules
already carry a name belonging to this one mesh — a workflow module stating its host, protocol and
absolute webhook URL, and two carrying a full clone URL for a repository on the mesh's own forge.

That changes what the issue is for. There is no version of this catalogue today that does not name
the mesh it was written in, so refusing three changes buys nothing and a mechanism is the only thing
that removes any of them. The three were merged on that reading, each PR saying so.
2026-09-26 15:07:22 +02:00
jschoubben bfcb660a8e Merge pull request 'Issue 122: a module cannot ask for its own public name' (#121) from issue/122-a-module-cannot-ask-for-its-own-public-name into main 2026-09-26 12:58:41 +00:00
jochen d7078061ea Issue 122: a module cannot ask for its own public name
Three open module changes independently wrote one mesh's names into the catalogue — two literal
public URLs, because the software generates absolute URLs behind a proxy, and one bucket renamed to
match what exists here. None was careless: the mesh composes <label>.<public-domain> for the proxy
and never hands it back to the module that asked for the route, and no interpolation yields a public
name, so writing the answer down is the only expressible option.

Files it rather than blocking the three, because the fix is a mechanism and the instances are live
needs. The cost is stated: a second mesh installing the identity provider gets the first mesh's
hostname, and nothing distinguishes a literal domain from a version number.
2026-09-26 14:58:18 +02:00
jschoubben bba9371832 Merge pull request 'The seats as they run, and to-be 26 implemented' (#120) from design/26-the-seats-implemented into main 2026-09-26 12:50:02 +00:00
jochen 3ff38d2ee1 The seats as they run, and to-be 26 implemented
Both halves are on their main branches, so the seats stop being an intention. Writes the as-is
document from the controller's code and the catalogue's manifests: the closed set of fourteen, the
three refusals a claim meets, the holder being an assignment and nothing else, and the one place a
seat changes resolution — which of several providers answers, never whether a requirement may go
unanswered.

Two things the as-is layer exists for are stated rather than smoothed over: a seat cannot answer
before it is held, which is the standing condition issue 121 records; and capacity is not
implemented at all, so the design's bench has no counterpart in the code.
2026-09-26 14:49:44 +02:00
jschoubben a89b0b5586 Merge pull request 'Issue 121 diagnosed: the seats work renamed the requirement, not the order' (#119) from issue/121-diagnosis into main 2026-09-26 12:48:12 +00:00
jochen 9766a3afce Issue 121 diagnosed: the seats work renamed the requirement, not the order
Asked first whether the seats change fixed this in passing, since it landed the same day and
touches both manifests the report names. It did not: the seat's holder is consulted only where
several nodes provide the thing, and with none providing it resolution refuses outright. Genesis
has no exemption — the unchecked first pass exists to learn what each node offers, and a
declaration is never built from it.

Records the part that did change: the three tests left failing on purpose were deleted by the
controller's seats PR and replaced with passing seat-based ones, so the gap is invisible again.
Adds the resolver to located-in, since that is where the refusal is.
2026-09-26 14:47:45 +02:00
jschoubben 8211dfd72c Merge pull request 'Accept ADR 0110 and ADR 0111; the seats design is in progress' (#118) from decide/0110-0111-accepted into main 2026-09-26 12:28:42 +00:00
jochen a7b2db9efc Accept ADR 0110 and ADR 0111; the seats design is in progress
The seats half of to-be 27's review is settled, so the two records it rests on are accepted and
the vocabulary catches up: the glossary's *seat* becomes a named role from a closed set, held by
an assignment and possibly delivering a provision, and 23 — Choosing a provider gains the seat
step in resolution, with ambiguity still refused rather than guessed. Both were held back when
0110 was proposed, because a document may not rest on a record that is not accepted.

26 — The seats moves to in-progress rather than designed: it names the files that implement it,
and naming a file claims implementation, which is only defensible once those files are on the
owning repositories' main branches. It becomes implemented when mesh-controller #63 and
mesh-catalog #69 land.

0112, 0113 and 0114 stay proposed; to-be 27 stays proposed with them.
2026-09-26 14:28:23 +02:00
jschoubben 99ffa6474d Merge pull request 'Issue 121: builder's real package-registry grant deadlocks a genesis bootstrap' (#117) from issue/117-builder-package-registry-deadlocks-genesis into main 2026-09-26 12:20:53 +00:00
jschoubben cb770f95c4 Merge pull request 'Issue 118: the analytics store answers the dial and times out the query' (#116) from issue/118-umami-store-query-timeout into main 2026-09-26 12:20:46 +00:00
jochen 09e502058b Issue 121: builder's real package-registry grant deadlocks a genesis bootstrap
Renumbered from 117, which is taken on main by 'a module's own code is a container in one record
and a process in another' — two reports claimed the same number and git would not have said so.

Scrubbed the node's name and a real registry path; this repository is public.
2026-09-26 14:20:05 +02:00
jschoubben 74b88d8efc Merge main 2026-09-26 14:19:58 +02:00
jochen b083790b21 Issue 118: the analytics store answers the dial and times out the query
Keeps 118: the other claimant to this number is on main as issue 119, where ADR 0112 points.

Scrubbed the service's public name — this repository is public — and completed the report's
frontmatter with the fixed-by and amended-design keys every other report carries.
2026-09-26 14:19:40 +02:00
jschoubben 0740de9d14 Merge main 2026-09-26 14:19:33 +02:00
jschoubben 79c692959e Merge pull request 'To-be 27 (proposed): a module requires, the mesh resolves — with ADRs 0109–0114, research 016 and issue 119' (#113) from design/27-a-module-requires-the-mesh-resolves into main 2026-09-26 12:17:56 +00:00
jochen d2044fb7b4 Merge main: the bus design, issues 113/114/117/120 and research 017 landed
# Conflicts:
#	04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md
2026-09-26 14:16:05 +02:00
jschoubben 01c6b89cc5 Merge pull request 'Design 25: the bus on NATS — proposed architecture for review; issue 103 resolved' (#93) from design/25-the-bus-on-nats into main 2026-09-26 12:15:19 +00:00
jschoubben 93502a05dc Merge pull request 'Issue 117: a module's own code is a container in one record and a process in another' (#110) from issue/117-a-modules-own-code-is-a-container-and-a-process into main 2026-09-26 12:14:38 +00:00
jschoubben 1272a97fd6 Merge pull request 'Research 017: a mesh that heals itself' (#115) from research/017-a-mesh-that-heals-itself into main 2026-09-26 12:14:24 +00:00
jschoubben de0c9b6cfc Merge pull request 'Issue 120: a provisioner remembers what it did, not what is there' (#114) from issue/120-a-provisioner-remembers-what-it-did-not-what-is into main 2026-09-26 12:14:17 +00:00
jschoubben 77934991c4 Merge pull request 'Issue 113 resolved by the repin, and what issue 064 did not cover' (#108) from issue/113-record-the-repin-and-fold-114 into main 2026-09-26 12:13:52 +00:00
jochen b5512bed32 Research 017: a mesh that heals itself
The operator's wish written as intended behaviour for the NATS bus:
every loop compares against what is, repairs by the ordinary path, never
destroys, and raises a condition for what it cannot fix. What is done
before NATS is limited to what survives the move.
2026-09-26 00:54:27 +02:00
jochen 0c5eac0218 0110 amends 0109: its seats are provisions, and moving npm takes a holdable seat 2026-09-26 00:44:35 +02:00
jochen b3bd50c588 Issue 120: a provisioner remembers what it did, not what is there
The harness compares against its own memory, so a backend that loses
what was provisioned (the cache's ACL users on a server restart) is
never provisioned again, silently.
2026-09-26 00:44:10 +02:00
jochen e387c4bd0e Apply review: two credentials, staged admin rotation, a ninth provider
The fact-check found mailu, whose user is its mailbox, so 0114 rotates
over two credentials rather than two logins, the adapter choosing what a
credential is. Also: minio keeps non-empty buckets; five backends take
their admin credential only at first init, so single-party rotation is
staged; postgres ownership moves to a non-login role; the harness keys by
consumer; rotation state lives with the vault. Consistency fixes across
0110-0113, 26 and 27; issue 103 resolved by mesh-host PR #22.
2026-09-26 00:38:06 +02:00
jochen 43f63ed41c ADR 0114: a two-party credential rotates over two logins
Graduates research 016. Retiring a login is separated from removing a
consumer, which closes a data-loss path in five providers; single-party
secrets rotate in place; the number of parties decides, not the provider.
2026-09-26 00:22:21 +02:00
jochen 942ebe350f Research 016: survey how each provider can rotate a credential
Overlap as drafted in 0113 would have deleted consumer data: seven of
eight providers name the resource after the login and five drop it on
remove. Rotation is now undecided in 0113 and to-be 27, pending the
survey. Also: a requirement naming a seat resolves to its holder, a
person chooses among remaining candidates at assignment, the controller's
secrets are requirements of its definition, genesis seals to the
control-node key, and moving the vault or broker is break-glass.
2026-09-26 00:14:23 +02:00
jochen 805df3f81e ADR 0113 and to-be 27: rotation overlaps old and new credentials
Decided with the author. A credential is never changed in place: each consumer has two logins, both
derived by the mesh, and uses one at a time. An applier adds the new login beside the old through the
adapter's existing create, and confirms both work; only then are readers released to the new one and
restarted by derivation; only when every reader has confirmed is the old login retired through the
existing remove.

It closes the three cases review found in applier-first rotation: an offline reader keeps working on
the old login until it returns; a bus account's owner keeps its bus until it has moved; a provisioner
restarted mid-rotation is still delivered both values. Nobody is ever without a credential that works,
which replaces to-be 13's all-or-nothing rule with a stronger one.

No consumer module changes. The alternation is the provider loop's. A provider's adapter gains one duty,
giving both logins the same rights over the consumer's data — in postgres, membership of one role that
owns it. The mesh derives two logins per consumer, both within ADR 0049's limit, which 0113 now names
among what it amends. Every rule has a check: overlap, offline reader, bus account, restarted
provisioner, equal rights, login length, and confirmation only once the old login is gone.
2026-09-25 23:50:20 +02:00
jochen 4a1b218706 Seats held by assignments, one assignment per module per node, and 0113's bottom of the stack
Decided with the author:
- A seat is held by one assignment, not claimed by a definition. A definition says which seats a module
  can hold; an assignment says which it does. The store module can run on every node and one assignment
  holds mesh-store; moving a role changes an assignment, never a definition. The foundation's seats name
  what the mesh itself uses and route no consumer — database and amqp consumers use co-location, the
  holder included. This replaces the wrong rationale that the foundation's store is "provider to nobody",
  which contradicted ADR 0078 and to-be 21. 0079's one-postgres rule becomes one mesh-store holder.
- A module is assigned at most once to a node. The instance identity in 0112 and 27 is withdrawn, and the
  login-length problem with it.

Review fixes to 0113:
- The bottom of the stack: the vault is installed as soon as the shared runtime base exists, and genesis
  generates everything needed until then — including the permanent controller's, the control-node
  agent's, the builder's and the broker provisioner's bus accounts, and the controller's store login.
  Genesis creates those accounts until the broker's provisioner runs and adopts them.
- Genesis's values are delivered recorded as the mesh's own, so 0092's never-replace rule for operator
  values does not make them unrotatable.
- Backend-issued secrets (a forge's once-only API token) enter through the vault. Non-module parties
  (the controller's logins, node agents' accounts) are answered the same way, the controller asking on
  their behalf; an enrolment token reaches the controller only as what verifies it.
- A secret with no provisioner to apply it is marked not rotatable by the mesh and refused, instead of
  a restart reported as done. Unused password generators in six provider clients are removed, and a
  catalogue scan checks no module mints.
- Rotation's lock-out cases (offline reader, bus account owner, restarted provisioner) are recorded as
  open, with overlap and re-confirm-with-safeguards as the two answers, to be chosen before acceptance.

0110, 0111 and 26 are marked proposed: they changed in meaning and are under review, and an accepted
record must not rest on proposed ones. To-be 23 and the glossary are restored to main; they change when
these records are accepted.
2026-09-25 23:47:26 +02:00
jochen 1b5f2c2c1a ADR 0113 and to-be 27: address the review of the vault rework
Two decisions taken with the author:
- Genesis delivers and the vault adopts. The vault cannot run first — it is built on the runtime base
  the installation makes after the store, broker and controller, and it learns its work over the bus.
  Genesis generates the foundation's first shared secrets, seals them to the operator key, and
  delivers them to the vault through the path an operator's value takes; from then on the vault holds
  and rotates them. This answers ADR 0085's own reason for rejecting vault-only minting, which 0113
  now names instead of stepping around.
- Rotation re-confirms on every pass. An applier repeats its confirmation until acknowledged, so a lost
  message costs one pass; an applier that stops after applying locks readers out until its supervised
  restart, and that window is stated and shown, not claimed away.

Fixes:
- Scope: a shared secret is made by the vault; a private key (node sealing keys, the operator's key,
  the certificate authority) is made where it is used. The inventory adds the makers the first version
  missed: node and builder broker passwords, and enrolment tokens.
- Broker accounts are created by the broker's provisioner, not the controller, so the controller never
  holds their plaintext; mesh-broker delivers amqp again — one broker per mesh — and only mesh-store
  delivers nothing.
- secret is a reserved provision: only the mesh-vault holder may provide it, and no pin routes around it.
- A secret's contract says whether a recipient applies it or reads it at start; appliers are never
  restarted for it, init-only secrets are applied, and confirmation is to-be 13's standard.
- Operator secrets are one rule everywhere: a secret requirement answered by the vault (0112 no longer
  says otherwise). A data provider's adapter may return fields; the data-return check names a lab consumer.
- 'Holder' now means a seat's holder only; a secret has recipients.
2026-09-25 23:27:58 +02:00
jochen fa2c09a2c5 ADR 0113 and to-be 27: the vault makes every secret — provisioning all the way down
A secret comes into being seven ways today: provider credentials, own secrets (54 modules), broker
accounts through a command that is easy to forget, a vault that only records what the controller
mints (6 modules), operator values, licences, and root secrets. The vault was built to end own secrets
and did not; the old path was never retired.

0113 is rewritten as a waterfall. The vault makes every secret and nothing else does. A provider that
needs a secret for a consumer requires it from the vault, declared once in its provision's contract
and expanded per consumer by resolution; the vault delivers it to both holders, each sealed to its own
node, so a provider's code is unchanged. Own secrets, broker passwords, operator values and licence
credentials take the same path. Genesis is not an exception: it raises the vault first and asks it,
so there is one way a secret is made from the first one on. The vault can sit at the bottom because it
requires nothing but a broker account.

One shared mint function in the SDK was considered and rejected: generation becomes uniform but custody
stays spread over every provider's machine, and each SDK language needs its own implementation.

Rotation is asked of the vault and is provider-first: the value goes to the holder that accepts it,
which confirms, before the holder that presents it gets it, so the lockout window shrinks to the
consumer's own restart, and an unconfirmed provider holds the rotation rather than half-doing it. The
host derives which processes to restart or recreate from the requirement a definition reads, so no
definition declares restart-on for a secret. A rotation shows unconfirmed until each consumer restarted
and passed its health check. Issue 103 becomes a prerequisite.

The file is renamed to match what it now decides. 0112 follows.
2026-09-25 23:10:52 +02:00
jochen 6e3373c879 Design pass: address the review
0113 — the plaintext claim was false under its own mechanism: handing a provider's answer to the
controller puts every secret on the broker and in the controller in the clear. The provider now seals
each secret field itself, to the consumer node's public key the mesh hands it, and the controller
carries sealed fields it cannot open. That is stricter than today, where the controller holds every
minted credential in the clear. Option 3 (plaintext to the controller) is recorded and rejected. The
foundation exception now covers root-secret rotation (0085) and forms like the broker admin's hash, so
no phase claims to remove the broker's bootstrap step. To-be 24 and 13 are named among what it amends.

27 — resolution is consistent with 0110: co-location and the only provider apply only where no seat
delivers the provision, so an unheld seat is refused even with one provider. The secret-field rule now
matches 0086 exactly (a declared env-file, never a container environment value). The seat placeholder
is the controller's, and the one module reading it moves to a host port. Contracts are held by the
controller and written down in phase 1, so they can be checked; every rule has a check. An operator's
secret is still the operator's, with the vault as custodian. Which seats a module holds is listed as
not settled.

0110 — the unheld-seat-with-one-provider case and the one-answer-for-everyone rule have checks; the
claim about moved manifests is corrected. 26 — the table governs and the code catches up, not the
reverse; scope and capacity agree with the glossary; moving a seat is described as it really is today.
0112 — aligned with 27, and lists 0049 and 26 among what it changes.

Issue 118 is renumbered 119: another branch took 118 first. 'Control-plane' is gone from 0110 and 0111.
2026-09-25 22:46:10 +02:00
jochen aad92ea8fe To-be 27 and ADR 0113 (proposed): a module requires, the mesh resolves
The design pass. Everything a module needs is a requirement: a name, a contract, and one of four kinds
of provider — a module, the node's host, the mesh, the operator. Installing a module resolves every
requirement or refuses, naming everything missing at once. It retires six mechanisms that grew
separately: provisions through bindings, settings, assigned ports, machine facts, minted secrets and
literals in the definition.

ADR 0113, proposed: a provider makes what it provides, and the mesh carries it back sealed to the
consumer's node. It is the return path ADR 0048 left "to a separate decision", now needed three ways:
data provisions with nothing to answer with, contracts needing a value the controller cannot make, and
a vault that generates nothing. Who a consumer is stays the mesh's (ADR 0049). Genesis is the one
exception. On acceptance it supersedes 0048 and amends 0085.

ADR 0112 is revised from three sources to that single concept.

ADR 0110 is amended for two points raised in review. The vault gets the mesh-vault seat (issue 106).
A seat's holder outranks co-location for a provision it delivers. Writing that down exposed an
inconsistency: mesh-store delivering postgres-database would have sent every database consumer to the
control-node, against to-be 23's node-local stores. So a seat delivers a provision only where the mesh
has one answer for everyone — artifact store, npm registry, git, vault — and mesh-store and mesh-broker
deliver nothing. 23 and 26 follow.

'Control plane' becomes 'controller' in the records written today.
2026-09-25 22:34:46 +02:00
jochen 7668190154 ADR 0112 and issue 118: address the review
- Secrets follow ADR 0085 as amended: a module's own secret is a provision the controller mints and
  the vault records. The previous commit had that backwards. Whether the vault should generate
  instead is recorded as an open question, not decided.
- A directory's contract is owner and mode only. The persistence flag was the keep flag ADR 0030
  refused; a directory is kept while it holds anything, and disposable data is a named volume (0107).
- An operator's shared data stays an access (ADR 0051), which rejected an operator-owned directory.
  Only where its path is written moves to the assignment.
- The records it changes on acceptance are named: 0051, 0091, 0046 (settings keyed by instance),
  0084 (a provider is a node and an instance), and the glossary, which gains its new words only
  when the record is accepted.
- How it is checked covers every stated rule. Container-side paths are no longer flagged by the
  host-path rule, and code fallbacks are covered.
- Provisions are what other modules provide. A seat's occupant is not listed as one, and the vault
  is not described as selectable per assignment.
- 'Control plane' becomes 'controller'. The provider count is ten of eleven, not eleven of twelve.
2026-09-25 22:29:38 +02:00
jochen 90fb7ae4ad ADR 0112: a module's own secrets are a provision from the vault, not something the mesh generates
The first draft listed minted secrets under what the mesh generates. ADR 0085 made a module's own
secret — a password, an internal token, an external key it was handed — a secret provision answered
by the vault, like a database by the store. What the mesh still mints is the delivery credential for
each provision a module takes (ADR 0048), the vault's own included.
2026-09-25 22:29:38 +02:00
jochen ca235e775f Issue 118 and ADR 0112 (proposed): a module definition names no node, no mesh and no path
Issue 118 records what a review of where module code reads its files found: 789 host-path strings
in 70 of the catalogue's 71 definitions, every one a decision the definition makes about a machine.
Mounts are checked (ADR 0091); the same paths retyped as values are not. It records what that has
already allowed — a DNS provider that would provision nobody silently, a contributions file that
names credentials by host path and so forces every provider to mount at the identical path, an SDK
loop that treats an unwritten contributions file as empty without a word, defaults in code that
disagree with their own manifests — and that no module can be assigned to one node twice, because
every identity is keyed by the module's name.

ADR 0112, proposed for review, answers it the way ADR 0038 answered ports: a definition names
variables, and installing it resolves every one or refuses, from three sources — the assignment's
own configuration, provisions the mesh resolves against a contract, and what the mesh generates or
knows. A directory becomes a provision: the module requires one by name with its owner, mode and
persistence, and where it lands is the assignment's. The mesh's own files stop carrying host paths.
An assignment gets an identity of its own, so a module may run twice on one node.

Checking copies for agreement was rejected as checking something that should not exist; rewriting
paths per assignment was rejected as inferring which strings are paths by their shape. Syntax, a
node's default layout, and when a second instance becomes possible are left to the design.
2026-09-25 22:29:38 +02:00
jschoubben 74ae0609cb issue 118: umami's store answers the dial and times out the query 2026-09-25 22:01:28 +02:00
jochen d94fe8f638 To-be 26: name the files that implement the seats and the build source 2026-09-25 20:47:48 +02:00
jochen c4cac767f8 ADR 0110: admit the-private-network, claimed by a manifest the control plane composes in code
The enumeration behind the first set read manifests in two repositories and missed a claim made
in the control plane's own code: the private-network module it ships claims the-private-network at
node scope. A closed set without it would refuse the control plane's own module. Thirteen claims in
use, naming twelve seats.
2026-09-25 20:35:31 +02:00
jochen dbe100ca96 ADR 0110 and 0111: a seat is a module assignment from a closed set, and a build source may live on the git seat
Seats have been doing two jobs and neither is written down. The mechanism ADR 0009 introduced is
enforced — a second holder is refused — but any well-formed name becomes a seat by being claimed,
and nothing can say which seats a mesh has or who holds them: holdings are assembled while planning
and discarded. The enumeration done while preparing this missed the control plane's own manifest,
because core modules' manifests live in its repository rather than the catalogue.

0110 closes the set. Each seat has a name, a scope, what occupying it delivers, and the record that
made it one; a claim outside the set is refused. A seat is held by a module assignment, and what the
mesh knows about the holder is what it knows about that assignment — nothing is stored beside it. A
seat may deliver a provision, and then its holder answers for it among several providers: pin, then
the holder, then the only provider, then refused. That keeps 0009's "refused, never guessed": the
seat is the choice made once, mesh-wide, instead of a pin per consumer node. The first set is the
eleven seats already claimed plus 0109's npm-package-registry, so nothing in use is refused.

Two concepts — seats for exclusion, a new word for consumable singulars — was rejected: both mean
"this mesh's one X", and the overview a person wants is one list.

0111 gives the mesh a git seat and makes a build source one of two explicit forms: a repository on
the seat's holder, recorded by its path and cloned from wherever the holder runs at build time; or
an external URL, recorded and cloned exactly as given. Recognising self-hosted sources by matching
URLs against the forge's address was rejected — it fails in the one case it exists for, after the
forge moves. Credentials for private repositories are left undecided and said so.

Design: new to-be 26 (the seats); 23 gains the seat step in resolution; 18's source entry names
the two forms; the glossary's seat and provision entries say where they meet. 0109 is carried from
its own branch so every link here resolves.
2026-09-25 20:33:14 +02:00
jschoubben 59c93dcfe4 109: a package registry seat is one per ecosystem, not one for all of them
Extends ADR 0075. Surfaced fixing builder's hand-faked package-registry
binding tonight: gitea's manifest declares the provision once with a single
npm-path, conflating what should be independently assignable per ecosystem
(npm/cargo/docker/...) the same way artifact-store and package-registry
were themselves split. Cited in 22-the-work-ahead.md's Phase 2, where the
target state this decision points at was already described a week ago.

Numbered 109, not 108: route-proxy's policy feature (mesh-controller PR
still-unwritten decision record — reserved but never committed. Renumbered
around it rather than colliding.
2026-09-25 20:29:22 +02:00
jschoubben 2ca63ae54e 117: builder's real package-registry grant deadlocks a genesis bootstrap
Fixing builder's hand-faked package-registry binding tonight (requires:
package-registry, a real mesh grant instead of a hardcoded JSON fragment)
broke three tests describing a deliberate carried-binding fallback for
exactly this: gitea's own image is built by builder, so builder cannot
yet hold a real grant from gitea the first time either has to exist.
Invisible on novox (already bootstrapped, gitea already live) — real on
any genesis from scratch. Fix left in place, tests left failing rather
than reverted or hacked, so the gap stays visible.
2026-09-25 16:59:04 +02:00
jochen 82a6badc7c Issue 114: land the controller's container-or-process question, renumbered
Filed 2026-09-24 on a branch of its own and never merged, numbered 113, which is taken. 114 is
free because a sibling branch folded it, so it takes that number and keeps its commit.

Kept separate from issue 117 rather than folded into it. 117 asks the same question of every
module and locates the missing decision; this asks it of the controller, where `network: host`
means container network isolation — the property that resource type usually buys — is not in use.
That observation is this report's own and is nowhere in 117, and folding would lose it.

Its first open question is answered by 117's diagnosis and now says so: the host's `process` shape
is built, applied and tested, restart and run-to-completion semantics included, so deciding this
does not wait on host-side work.
2026-09-25 16:16:21 +02:00
jschoubben 10a2b706c6 Issue 113: should the controller be a container or a process the host supervises
Filed after a session where every mesh-controller interaction went through
docker exec — its manifest runs it as a container with network: host, using
none of the isolation that resource type usually buys, while ADR 0006 makes
it the mesh's single point of coordination. Open question, not a claimed
defect: does type: container get the controller anything type: process
(supervised the way the host supervises its own unit, per ADR 0005) would not.
2026-09-25 16:15:27 +02:00
jochen 35db2aaa41 Issue 117: a module's own code is a container in one record and a process in another
Asked what the "sidecar" is and whether a supervised process would do instead. The repository
answers both ways. ADR 0047 (accepted, unsuperseded) says a module with tools or events runs a
container carrying its compiled code. To-be 18 and 20 (both proposed) define a `process` resource
type — the module's own code, a unit the machine's supervisor keeps up — and the worked guide says
plainly "it is why these are `process` rather than four containers." Neither design doc names 0047,
and no decision record mentions a `process` shape at all.

Diagnosed rather than left open, because the ground truth settles what the report could not.
The shape is real: mesh-host defines TypeProcess, applies it, and tests it, and the host's
vocabulary is twelve shapes rather than the nine ADR 0029 counted. So the alternative the report
offered — that two proposed documents describe a type that does not exist — is disproven.

ADR 0029's mechanism is intact and was not enough. The vocabulary-count test names the decision
behind each addition: network 0029, access 0051, opening 0100. The eleventh names a *proposed
design document*, and TypeProcess is the only shape in the vocabulary whose doc comment cites no
ADR. Requiring every addition to name something does not require it to name a decision.

The argument this issue asked for already exists — as a Go test comment. "It is a full-host shape
rather than a portable one: it needs a process supervisor to install into. It does NOT need a
container runtime, which is the point — only software that genuinely needs isolation asks for a
container." That is a decision's context and consequences, in another repository.

What the catalogue does is a third thing: 115 container declarations against 3 process, all three
in showcase — the module to-be 20 documents. There the tools resource is a container running
`sleep infinity` on a bare upstream base with the broker credential mounted, and the tools and
provisioner entrypoints are run by nothing. That is the condition 0047 was written to end, back
in a new shape.

Where the isolation argument leaks is narrower than expected and worth having precisely: the
serving key and the credential shape both conform. But serveTools serves every registered module
over one broker connection, the runtime takes its modules from a comma-separated list, and
x-source is stamped from the single credential — so two modules in one runtime means the second's
events are attributed to the first. Nothing refuses it and no test asserts against it.

Located on hq rather than on a code repository: the implementation and the design layer agree,
and the missing thing is the record. Which shape is right is left open, deliberately — this
establishes that the question was answered in practice and never written down, not which answer
is correct.

One correction kept in the trail: the first search here was for len(Vocabulary()), found nothing,
and was two steps from being written up as "the mechanism ADR 0029 relied on is gone." The test
binds the slice to a local first. A negative search result read as a fact about the world is the
same error issue 113 recorded.
2026-09-25 16:15:00 +02:00
jschoubben 56669ee23b Merge pull request 'Issue 116: route-proxy has no authentication or IP-restriction mechanism' (#109) from issue/116-route-proxy-has-no-auth-or-ip-restriction into main 2026-09-25 12:28:50 +00:00
jochen 367df38e6d Issue 116: resolved by mesh-controller PR #58
The gap is closed in the proxy: policy applies, the four capabilities exist, the table is keyed
by host and path with a total ordering, and the two failure modes that rot quietly are held by
tests — a declaration carrying a credential refused rather than served, an unreadable secret
failing closed.

Resolved rather than left open because the issue reports a gap in the proxy and that gap is
gone. But the record says plainly what it does not yet allow: an operator still cannot move the
affected routes, because that needs the mesh side — a manifest able to declare these values and
the controller minting the secret auth names. Until both exist the capability is reachable only
by writing the routes file by hand. That is the ordinary build-out of a contract this issue's
decision created, and it belongs to to-be 08 rather than here.

The open questions are marked answered and kept rather than deleted, pointing at ADR 0108 —
what was rejected and why is the half worth having, and a section still saying "the fix should
not be written before these are answered" after the fix was written reads as though nobody
looked.

One finding kept in the record: priority was read with the reader for ports, which caps at
65535, and the one real rule this reproduces is declared at 100000. It parsed to zero, so
refusal and path scoping would have shipped looking complete and doing nothing on the only case
that motivated them. A validator borrowed from a neighbouring field is a silent default.
2026-09-25 14:25:24 +02:00
jochen a11da86591 ADR 0108: a route carries the policy applied to a request
Issue 116 found the mesh's proxy applies nothing to a request — host lookup, forward. Against
what the replaced ingress actually relies on, four capabilities are missing: authentication
(three dependents, each gating an admin surface with no login of its own), refusal scoped to a
path (one, a live incident mitigation), path-scoped routing with priority, and redirect.

Policy goes on the route rather than beside it. A proxy-side settings layer keyed by route name
would keep the grant literally clean, but then "what protects this route" is answered from two
files nothing keeps in step — and a route's protection is part of what a route is.

The set is closed at those four, so a fifth is an amendment and each addition is earned by a
dependent that exists. An open middleware surface was rejected: it recreates what is being
replaced, and narrowing one later is far harder than widening a closed one.

Where policy needs a credential the declaration names a secret and never carries the value,
which keeps the existing secret machinery the only thing holding credentials. Inlining a hash
was rejected as the first credential in a declaration — a precedent easier to set than withdraw.

This re-keys the routing table by host and path with priority, which follows from the decision
rather than being a separate one: two of the four need one host routed more than one way. Equal
priorities must resolve identically every time or the proxy stops being reproducible.

The record says how it is checked, including the negative case that rots quietly — a
declaration carrying a credential value rather than a reference must be refused, so the
rejected option cannot return by accident.

08-connectivity §3 names the record and gains the subsection; issue 116 gains amended-design.
2026-09-25 13:48:18 +02:00
jochen c839d9ac26 Issue 116: scrub the disclosure, and correct the count and the shape of the gap
Two things the report got wrong, and one it could not have found the way it looked.

Disclosure first: it carried a real hostname and an absolute node path, in a public
repository. Both are gone; the ingress, the modules and the routes are named by role, as the
rest of 04-ISSUES does.

The count was low. Basic authentication has three dependents in the catalogue, not one — the
key-value store's browser UI, a database web UI, and the ingress's own dashboard. All three
are credential-less admin surfaces whose only gate is a middleware the mesh's proxy lacks.
The earlier version read only the node's dynamic configuration directory, which cannot see
what modules declare as container labels; counting needs both sources, and the report now
says so.

Two gaps were missing entirely. Redirect rules: two live routes canonicalise a www name onto
its apex, they exist only on the node and not in the catalogue, and they fail silently rather
than erroring. And path-scoped routing with priority, which is the one that reorders the
issue: the table maps host to exactly one target, so a host cannot be routed two ways, and
the refusal rule matches a path on a host already routed elsewhere. Authentication and a
source filter would not make it expressible. Path scoping is a prerequisite, not a sibling.

Also corrected: the refusal rule was described as an address-scoped deny. It is an allow-list
holding a single documentation-range address — deny-everyone — so reading it as address-scoped
points at the wrong fix. And its severity was understated: its own header records it as
incident response closing an abused write primitive, which is not "a real exposure" but a live
mitigation.

The open questions now say plainly that they are design questions and the fix should not be
written before they are answered, and one is added: whether a declaration may carry a
credential at all.
2026-09-25 13:24:34 +02:00
jschoubben 226d556743 Issue 116: route-proxy has no authentication or IP-restriction mechanism
Comparing route-proxy against what HAL's actual Traefik config does today,
not Traefik's general feature set, per the standing rule that the nox mesh
must do at minimum what the HAL mesh it replaces already does. Everything
else checked out even or better; these two are real, confirmed gaps —
RedisInsight has no login of its own and depends entirely on Traefik's
basicauth middleware, and the gitea-internal route depends on an IP-scoped
deny rule. Neither has any equivalent in route-proxy's single-lookup
request path.
2026-09-25 11:51:56 +02:00
jochen ab7d216fc6 Issue 113 resolved by the repin, and what issue 064 did not cover
The object-store module was repinned to a maintained fork of the withdrawn server image, its
runtime sidecar built rather than pulled, and its data moved off the predecessor's live
directory. The instance is closed; the three general points the report makes are not, and What
was done says so rather than letting a resolved status imply otherwise.

Folds in the one thing a duplicate report of this symptom had that this one did not: issue 064
asked whether the build environment can reach a declared vendor image and assumed that, once
declared, it stays fetchable. Withdrawal is the case that assumption does not cover. The
duplicate is not merged — it carried the reading this report's diagnosis retracts.
2026-09-25 00:55:13 +02:00
jschoubben cdcd4da27e Merge pull request 'Research 015: reopen the comparison — the premise for narrowing to one candidate was false' (#105) from storage/015-reopen-the-candidate-comparison into main 2026-09-24 16:47:18 +00:00
jochen 1d524a1fa8 Research 015: rewrite the comparison — wrong axis, and a missing candidate
The previous version ranked candidates on whether they preserved single sign-on to the
object store's console. That is not a requirement: a "user" of the store is normally an
application, so the requirement is per-application keys scoped to buckets — which the mesh
already mints. And the console login it ranked on never worked; the module's own hook comment
records "policy claim missing", a failing login written up as progress.

It also omitted the incumbent's own maintained fork, which changes the question from "which
product replaces it" into two decisions: repoint, or migrate — and if migrating, to which.
Repointing costs an image reference; migrating costs a data copy, two handler rewrites and a
maintenance window. Repointing does not foreclose migrating, which is the argument for taking
it first.

On the corrected requirement Garage ranks first — its per-key-per-bucket model is the
requirement verbatim, its admin API matches how the mesh provisions, and the highest-risk
consumer is first-party documented against it. Its remaining gap (no versioning, no
server-side encryption, partial lifecycle) is unmeasured against the buckets and is the one
thing that could still disqualify it.

Measured and folded in: 230 GiB logical, 82,496 objects, 468 GiB raw at 2.03x, eight drive
directories on one filesystem on one machine. That last fact decides more than any feature —
the erasure coding is not buying independent-drive redundancy, so the redundancy model is
close to irrelevant and only storage overhead remains, which at this volume is a rounding
error against the headroom.

Both errors are recorded at the end of 01 rather than quietly fixed. A configured feature is
not an observed one; and when a dependency dies, "who took it over" precedes "what replaces
it" — searching for alternatives by construction returns things that are not the incumbent.
2026-09-24 18:46:47 +02:00
jochen 999636e2a8 Issue 113: retract the diagnosis table — the original report was right
The diagnosis carried a table headed "claims that could not be substantiated", denying a
module.json, a digest pin, and an all-zeros runtime digest. All three exist. The table is
withdrawn in full and replaced with what is actually true, plus the two claims that remain
genuinely unverified rather than disproven.

The cause: one repository was searched and absence in it was written up as absence. The
catalogue of the mesh being built is a separate repository, not checked out where the search
ran, and all four claims were about that repository. Compounding it, the predecessor's
object-store module and the one being cut over to were treated as one thing — they are
different files in different repositories, one pinning a tag with no sidecar, the other a
digest with two container resources.

Also corrected in the report: located-in named the wrong repository; the "pins a tag" passage
described the predecessor; the open question about pinning by digest is struck, because this
module already does and it made no difference — a deleted digest resolves to nothing either
way. The section on why nothing broke is now scoped explicitly to the predecessor's
machinery.

The lesson kept in the record: "zero occurrences anywhere in the tree" is only as strong as
the tree searched, and a diagnosis must say which tree. A confident rebuttal of a correct
report is worse than no diagnosis — it sends the next person to the wrong place with a
written record behind them.
2026-09-24 18:46:47 +02:00
jochen 48abc36b5b Merge remote-tracking branch 'origin/main' into work/object-store-records 2026-09-24 18:41:53 +02:00
jschoubben 2c5f805467 Merge pull request 'Add hq-defer: park a thought without moving the work off course' (#107) from meta/hq-defer-skill into main 2026-09-24 16:39:22 +00:00
jochen db2c950ba6 hq-defer: make the MEMORY.md pointer an explicit placeholder
Review caught it reading as a real relative link, so a link checker flags
.claude/skills/hq-defer/file.md forever. Angle brackets say placeholder.
2026-09-24 18:38:43 +02:00
jochen 971d0839f5 Add hq-defer: park a thought without moving the work off course
A thought raised mid-task needed remembering but not working on, and there was no
mechanism for that — so it was recorded by hand. This is that, made repeatable.

Records to Claude's persistent memory rather than the repository, deliberately. A parked
thought has no number, owner or status: giving it one asserts triage that deferring says
has not happened. A shared "deferred" document would be a central status file, which
AGENTS.md forbids. And a repository write means a branch, a commit and an MR — the drift
the skill exists to prevent.

Wraps no playbook, because deferring precedes the development cycle rather than being part
of it. It does say which playbook a thought would need if it graduates, and that recording
"undetermined" is the honest answer when the evidence does not say.

The stop condition is the substance: at most two lines, then return to what was in
progress. No plan, no triage question, nothing opened.
2026-09-24 16:40:24 +02:00
jschoubben 5677e97508 Merge pull request 'ADR 0107: persistent data is a directory bind, never a named volume' (#106) from decide/0107-persistent-data-is-a-directory-bind into main 2026-09-24 14:24:17 +00:00
jschoubben a93743708c ADR 0107: persistent data is a directory bind, never a named volume
Records the rule the operator gave directly, mid-session, after checking
that HAL's own postgres and lavinmq both used a directory bind and the
mesh's adoption of them three weeks ago switched to a named volume without
a reason recorded anywhere.

Already built and rolled out on novox (mesh-catalog PR #54) before this
record -- urgent enough to fix first and write down after. Includes the
incident: the new host directories needed the container's own UID, which a
named volume gets for free and a directory bind does not; mesh-store
crash-looped on Permission denied until ownership was matched to what the
original volume already had.

Closes issue 115. Checks pass.
2026-09-24 16:21:51 +02:00
jochen c8cbbcfb8a Research 015: reopen the comparison — the premise for narrowing to one candidate was false
SeaweedFS was scoped as primary because it looked like the only candidate preserving
OIDC console login. Measured: its admin UI is Apache-2.0 but its identity-provider
integration is not — console SSO sits behind the per-TB commercial licence, alongside
point-in-time recovery and automatic EC repair. The free build gives OIDC on the S3 API
via STS and a console authenticated by local username and password.

So the answer to the gating question is that no candidate preserves the current feature
set for free, which this effort had written down as a possible outcome. Reopened across
three candidates with the requirement-by-requirement evidence in 01.

Two corrections to what the overview recorded. RustFS is not a binary-level drop-in
retaining existing data: API and on-disk compatibility are separate paths and the on-disk
one is preview-scoped. And it carries an open defect in the credential path the bucket
provision depends on, which gates it specifically.

Nothing graduates before two measurements named in 01: whether an authenticating proxy
is an acceptable answer to console SSO, and which S3 endpoints consumers actually call —
the latter because Garage does not implement the full span and cannot be ranked until
that is counted.
2026-09-24 16:10:33 +02:00
jschoubben a458f751c6 Merge pull request 'Accept ADR 0038; close issue 091' (#104) from issue/091-ports-are-mesh-assigned-not-manifest-fixed into main 2026-09-24 13:53:36 +00:00
jschoubben 402b798ab6 Accept ADR 0038; close issue 091
The decision (the mesh assigns a container's machine-side port; a module
says only what it needs) was proposed 2026-09-01, and the machinery
already implements it in full -- internal/inventory/ports.go's PortFor,
declaration.go's publishedOn. What was missing was the catalogue actually
complying: 14 of 46 modules baked a machine-side number into their own
manifest anyway. mesh-catalog PR fixes 11 of them (the two defensible
kinds -- foundation, protocol-fixed -- are left alone, per the issue's own
categories). Accepting the decision now that it's actually enforced, and
closing the issue it was blocking.

Checks pass.
2026-09-24 15:52:45 +02:00
jschoubben f10dce4f9e Merge pull request 'Issue 113 and research 015: the object store's images are gone upstream, not access-restricted' (#103) from storage/113-the-object-store-lost-its-upstream into main 2026-09-24 13:45:34 +00:00
jochen 4beb6629db Issue 113: ground the rebuildability point in what the design actually says
Review of my own text found an unattributed claim — "the mesh's claim that a node
can be rebuilt from its declarations" — which is not a stated principle anywhere.
Replaced with the design position that genuinely covers it: to-be 07 chooses
references over payload because "reproducibility comes from pinning the identity of
a thing rather than carrying its bytes". This incident is that choice's failure mode
when the identity stops resolving, which is a sharper point than the one I made.

Scope stated honestly: the passage is about the foundation bundle and this module is
not in it, but pin-identity-fetch-bytes is how every module gets third-party images.

Also names the tension the mirroring question actually carries — mirroring is a move
away from references-over-payload, so it is a decision, not a fix.
2026-09-24 15:44:44 +02:00
jochen d497b37e43 Issue 113 and research 015: the object store's images are gone upstream, not access-restricted
The symptom arrived diagnosed as "the registry disabled anonymous pulls for the
whole vendor namespace". It did not hold: sibling repositories in that namespace
pull normally, the "$disabled" token field appears on every repository including
working ones and describes signing rather than access, and "actions": [] with a
401 is byte-identical to what an invented repository name returns. Both registries'
own APIs establish deletion instead.

Recorded because the correction is the expensive part to rediscover, and because
the instance was harmless while the standing condition is not: no node that does
not already hold the images can ever provision the module again, and nothing
detects that until one tries.

Research 015 scopes the replacement. It is not a redesign — the foundation design
already commits to S3 the protocol rather than the product, and the object store
is an ordinary module, so this instantiates an existing principle. The live OIDC
wiring is the requirement that gates the choice, and it is checked first.
2026-09-24 15:27:03 +02:00
jschoubben 183b22997c Merge pull request 'Issue 112: diagnose — the predecessor's own DNS config already names the carried peers' (#101) from issue/112-diagnosis into main 2026-09-24 12:49:50 +00:00
jschoubben 8d67cf63c5 Issue 112: status located, not diagnosing
Playbook 03 step 2: move status to diagnosing, then located once the
owner is known. located-in is filled with four confirmed packages —
the owner is known.
2026-09-24 14:27:25 +02:00
jschoubben 75c104c355 Issue 112 diagnosis: correct located-in attribution
The carried-peer record (CarriedPeer/TunnelPeer) lives in mesh-controller
internal/inventory, not internal/catalogue. internal/catalogue is the
right package for the zone-generation side of the fix (facts.go's
nodeZones), but a different concern from where the name field itself
would go. Split the two so a decision record doesn't get pointed at the
wrong package.
2026-09-24 13:54:20 +02:00
jschoubben 6abfec7433 Design 25: address first review's four findings before any code
Fixes, each named where it was wrong:

- reload-on is a service field; a container only has restart-on, which
  recreates. Cited precedent (registry-trust-reload) is a service resource,
  not a container. Fix: nats-server's own SIGHUP reload, triggered by an
  in-image entrypoint watching a directory-mounted config file (issue 103's
  recreate-on-change applies to a directly-mounted file, not a directory's
  contents) -- asks nothing new of the host.
- A JetStream-delivered message's Reply field is already claimed by the
  consumer's own ack address, so a responder using it answers nobody. Fix:
  every CONTROL message needing a reply carries its reply subject in its own
  payload; the controller publishes there explicitly, never via Respond().
  Enrolment is the case this design actually depends on, so it's fixed there
  too, not just noted.
- The listed permissions never granted publish on a durable consumer's own
  ack-reply subject -- a module could receive but never ack, so every
  message redelivers forever. Fixed with a scoped grant per module's own
  consumer.
- One account (a deliberate choice, kept) means inbox privacy is the
  permission list or nothing. The design granted 'its reply inbox' without
  scoping it, which read as any user reaching any inbox. Fixed: each user's
  inbox prefix is derived from its own identity and its permissions name
  only that prefix.

New open question from this revision, not closed: whether the in-image
watch-and-SIGHUP shape belongs in mesh-sdk if a second module ever needs it.

Checks pass (records.py, cycle.py, index.py).
2026-09-24 13:53:53 +02:00
jschoubben 6a56738d7b Issue 112: diagnose — the predecessor's own DNS config already names the carried peers
Checked why ADR 0104's forward-to-predecessor shape doesn't transfer to the
resolver the way it did the proxy: DNS is one process on one port, and
assigning the mesh's dnsmasq module replaces it in place, so there is no
predecessor process left standing to forward to.

But /etc/dnsmasq.d/hal-dns.conf's static address= lines for ace/shanks/g14
match the mesh's own carried-peer addresses from overlay show exactly. The
name a carried peer needs isn't a guess the operator has to make under
pressure — it's a transcription of a record the predecessor already has and
has been correctly serving for six days. Located in mesh-controller (no way
to attach a name to a carried peer today) and the dnsmasq module (doesn't
emit a wildcard for a named-but-uncarried peer). Not implemented.
2026-09-24 13:25:44 +02:00
jschoubben 91a5c63d65 Merge pull request 'Name the migration repository in the map, so nobody has to be told it exists' (#99) from meta/name-the-migration-repository into main 2026-09-24 00:02:07 +00:00
jschoubben 545d038198 Name the migration repository in the map, so nobody has to be told it exists
hq cannot hold the migration's operational record — it names machines, addresses and
paths, and this repository is public — but it can say where that record is, which is what
this map is for. Asked for by the operator, who had to be told.
2026-09-24 02:01:41 +02:00
jschoubben 7f438d049f Merge pull request 'Issue 092: genesis publishes to a registry the container runtime does not yet trust' (#79) from issue/092-genesis-registry-trust into main 2026-09-23 23:38:46 +00:00
jschoubben 60e43f9446 Merge pull request 'Issue 091: a module definition carries a machine port' (#78) from issue/091-machine-ports-in-manifests into main 2026-09-23 23:38:40 +00:00
jschoubben 36aa722c2a Merge pull request 'Issues 111 and 112: the resolver was told the wrong set of names, twice over' (#98) from issues/111-112-the-resolver-was-told-the-wrong-names into main 2026-09-23 23:32:42 +00:00
jschoubben f704e2ca64 Issues 111 and 112: the resolver was told the wrong set of names, twice over
111, resolved: the map the control plane hands a resolution holds the machines and the
names the mesh merely serves, and the resolver's zones were given both — inventing names
under a suffix it answers authoritatively for. 112, open: adopting a tunnel gives the mesh
the peers' addresses and none of their names, so taking the resolver before they enrol
stops three machines resolving at all.

Both found by reading the plan before pushing it.
2026-09-24 01:32:04 +02:00
jschoubben e7a90be3ee Merge pull request 'Issue 110: on a converged node a container on the runtime's own network cannot reach the resolver' (#97) from issues/110-the-resolver-and-the-default-network into main 2026-09-23 23:13:48 +00:00
jschoubben 3a8515273d Issue 110: on a converged node a container on the runtime's own network cannot reach the resolver
Found reviewing the resolver's conversion. Nothing fails while the node is adopted; it
fails at the flip, and it is the same split that decided which container survived the
hub's address change.
2026-09-24 01:13:30 +02:00
jschoubben 0e70ca0808 Merge pull request 'Issue 109: a container keeps the address it was made with' (#96) from issues/109-a-container-keeps-the-address-it-was-made-with into main 2026-09-23 23:02:48 +00:00
jschoubben 6ccd138729 Issue 109: a container keeps the address it was made with
Found when adopting the tunnel moved the hub's address: the declaration followed, the
running container did not, and the forge lost its database. Issue 102's rule broken one
level down, and issue 103's fix stopping one input short.
2026-09-24 01:02:16 +02:00
jschoubben 07ee79199a Merge pull request 'ADR 0105: what review settled — the tunnel adoption is implemented' (#95) from decide/0105-implemented into main 2026-09-23 22:39:08 +00:00
jschoubben f660620637 ADR 0105: what review settled — carried peers, the flip, the refusals, and keeping the hub's identity
Implemented in mesh-controller #49 and mesh-host #24. One proposal was rejected on the
record's own terms: converging the hub is not made to wait on other machines' migrations.
2026-09-24 00:38:54 +02:00
jschoubben f61f047a5d Merge pull request 'Issue 102 resolved and verified on the machine; issue 097's orphan was on the host network' (#94) from issues/102-resolved-and-097-worse into main 2026-09-23 22:15:49 +00:00
jschoubben 133e10a738 Issue 102 resolved, verified on the machine with both forwarders gone; 097's orphan was on the host network
The addresses follow: the control plane holds the ports the node gave, a recorded build
holds no address at all, and the two forwarders that had been holding the control plane
together are removed. 097's stranded container turned out to be listening on every
interface and connected to the mesh's store — by its own old database, which is the only
reason nothing was at risk.
2026-09-24 00:02:13 +02:00
jschoubben f3ad60b98c Issue 103 resolved by mesh-host #22 2026-09-23 23:44:31 +02:00
jschoubben 0c433d51ad Design 25: the bus on NATS — subjects, streams, accounts as configuration, enrolment, a person's client, the cutover, the beds
The architecture ADR 0106 asks for, proposed for review before any code.
2026-09-23 23:42:42 +02:00
jschoubben c209a575e1 Merge pull request 'ADR 0106: the bus is NATS; issue 104 resolved' (#92) from decide/0106-the-bus-is-nats into main 2026-09-23 21:40:03 +00:00
jschoubben e022798858 ADR 0106: the bus is NATS — native, built beside the migration, cut over after its core; issue 104 resolved 2026-09-23 23:39:17 +02:00
jschoubben 6f173a7912 Merge pull request 'Issue 108: the registry has no garbage collection, and two doors make it harder to add' (#91) from issues/108-registry-gc into main 2026-09-23 21:32:52 +00:00
jschoubben 73091dcb4c Issue 108: the registry has no garbage collection, and two doors make it harder to add 2026-09-23 23:32:29 +02:00
jschoubben f6ec64ee4e Merge pull request 'Issue 107: a declaration carries no order; rescue on an enrolled node is reconcile, not apply FILE' (#90) from issues/107-declarations-carry-no-order into main 2026-09-23 21:27:24 +00:00
jschoubben 671c2f3881 Issue 107: a declaration carries no order; rescue on an enrolled node is reconcile, not apply FILE
Both from the review of the issue-104 fix: a hand-applied file on an enrolled node is
recorded as carried and would remove the foundation, and nothing on the wire orders one
declaration against another.
2026-09-23 23:27:08 +02:00
jschoubben 2445d80565 Merge pull request 'Research 014: the bus on NATS — decide now, build in the lab, cut over once after the core' (#89) from research/014-nats into main 2026-09-23 21:15:22 +00:00
jschoubben a548b34f5d Research 014: fix the reference to ADR 0039 2026-09-23 23:14:57 +02:00
jschoubben b59907ee08 Research 014: the bus on NATS — decide now, build in the lab, cut over once after the core
Measured: AMQP is spoken in three places of the mesh's own code and in none of the
sdk or the modules; the predecessor's world is AMQP and retiring. NATS answers every
guarantee the bus relies on, durability via JetStream. Recommended: not underneath
the migration, not after it either — in parallel, one rehearsed rollout.
2026-09-23 23:14:39 +02:00
jschoubben 8638ba3a4f Merge pull request 'Issues 102–106 and ADR 0105: what the core migration found, and the hub adopting the predecessor's tunnel' (#88) from core/issues-102-106-and-tunnel-adr into main 2026-09-23 20:54:53 +00:00
jschoubben cb2117f1c4 Issues 102–106 and ADR 0105 from the core migration
Two birth-address outages and a registry that would have been the third; a
container that keeps a stale environment after its file changes; a host command
that applied a converged declaration to an adopted node; the hub and the vault
without seats. And the decision the operator made under it all: the hub adopts
the predecessor's tunnel in place, key and peers and range and port.
2026-09-23 22:50:10 +02:00
jschoubben ee2bdf220c Merge pull request 'Issue 101: taking a service its neighbours reach by container name cuts them off' (#87) from issues/101-a-service-reached-by-name-loses-its-network into main 2026-09-23 18:14:08 +00:00
jschoubben 61e4e971a4 Issue 101: taking a service reached by container name cuts its neighbours off
Found checking the third cutover rather than running it. The first two were safe by
accident — both are reached through a host port, which survives a change of owner.
This is the first constraint found that decides the order of the migration.
2026-09-23 20:13:53 +02:00
jschoubben f4f58e7c30 Merge pull request 'Issue 100: a secret the mesh mints cannot be the one the service it takes over already uses' (#86) from issues/100-a-minted-secret-cannot-be-the-one-the-service-already-uses into main 2026-09-23 00:54:21 +00:00
jschoubben 157c6edd50 Issue 100: a minted secret cannot be the one the service already uses
Found at the second cutover. Carrying a value in works only for a module's own
secrets; a secret answered by the provision is minted, and six catalogue modules
take one that way.
2026-09-23 02:54:09 +02:00
jschoubben e9e5df36bd Merge pull request 'Issue 099: a module's image pin ages into a downgrade, and taking it over is where that is discovered' (#85) from issues/099-a-pin-ages-into-a-downgrade into main 2026-09-23 00:48:02 +00:00
jschoubben 9faa985be3 Issue 099: a module's image pin ages into a downgrade
Three modules in a row on one machine; the first was found by taking it and cost a
three-minute outage. The runbook's answer is a rule a person must remember, which is
the shape this repository says not to settle for.
2026-09-23 02:47:50 +02:00
jschoubben cda7a4e348 Merge pull request 'Issue 094 diagnosed and resolved; 096, 097 and 098 opened from what it uncovered' (#84) from issues/094-diagnosis-and-096 into main 2026-09-23 00:37:27 +00:00
jschoubben 668ce3ad62 Issue 094 resolved: a given port names either end and is answered once
The first pass answered under both ends, which review showed is the same fault seen
from the other side where two mappings share a number. Verified on the machine: the
forge is back on the port its own configuration has always advertised.
2026-09-23 02:37:07 +02:00
jschoubben 0c8615aa8f Issue 098: taking a module replaces a configuration nobody compared
Found reading the second module's cutover rather than running it: the catalogue's
config drops a rule the machine's has, and no step puts the two side by side.
2026-09-23 02:23:58 +02:00
jschoubben 235b9ea0e5 Issues 096 and 097, and 094 diagnosed: a setting stored where it cannot work, and a resource that changed target
094's cause is one blind spot read from two ends, written up in its diagnosis; the fix
answers the first open question and not the other two, which become 096. 097 was found
looking at what the forge's cutover left running.
2026-09-23 02:20:48 +02:00
jschoubben 1b8e5043ee Merge pull request 'Issues 094 and 095, both found in the first module's cutover' (#83) from issues/094-095-from-the-first-cutover into main 2026-09-22 23:57:48 +00:00
jschoubben 515cb8adc1 Issues 094 and 095, both found in the first module's cutover 2026-09-23 01:57:10 +02:00
jschoubben 0d8b683ad7 Merge pull request 'Research 013: the forge and the registries — a seat answers the wrong question' (#82) from research/013-the-forge-and-the-registries into main 2026-09-23 01:03:10 +02:00
jschoubben 43b55d6664 Research 013: the forge and the registries — a seat answers the wrong question; issues 090 and 085 corrected from the code 2026-09-23 00:38:34 +02:00
jschoubben 6c81ea2204 Merge pull request 'ADR 0104: a provision may be answered by an adapter to the predecessor' (#81) from decide/0104-route-adapter into main 2026-09-22 23:57:54 +02:00
jschoubben a23ede495e ADR 0104: a provision may be answered by an adapter to the predecessor; issue 093 located; connectivity says how the proxy hands over 2026-09-22 23:57:38 +02:00
jschoubben a2323ade2e Merge pull request 'Issue 093: the successor proxy cannot serve what the predecessor still serves' (#80) from issue/093-proxy-handover into main 2026-09-22 23:57:04 +02:00
jschoubben d697c6f776 Issue 093: the successor proxy cannot serve what the predecessor still serves, so no web module can migrate one at a time 2026-09-22 23:45:49 +02:00
jschoubben a7b3f7823b Issue 092: it happened twice more — the registry's mesh name, and the mesh's own trust naming the default port 2026-09-22 23:37:05 +02:00
jschoubben cc3af29084 Issue 092: genesis publishes to a registry the container runtime does not yet trust 2026-09-22 22:43:17 +02:00
jschoubben fb9d3035d4 Issue 091: a module definition carries a machine port, measured across the catalogue 2026-09-22 22:26:14 +02:00
jschoubben dd81523eab Merge pull request 'Issue 085 resolved' (#77) from fix/085-resolved into main 2026-09-22 22:00:49 +02:00
jschoubben 5c5821b210 Issue 085 resolved: the packages port is a node setting; two of its open questions stay open 2026-09-22 22:00:38 +02:00
jschoubben c46c508c03 Merge pull request 'Issues 088, 089 and 090, found fixing 085' (#76) from fix/issue-085-followups into main 2026-09-22 22:00:02 +02:00
jschoubben 1728765fe3 Issues 089 and 090: a contributed route does not follow a moved port; the forge module cannot take over the forge genesis raised 2026-09-22 21:54:35 +02:00
jschoubben e609ccdbfb Issue 088: the forge's own address names a port it may not have 2026-09-22 21:41:58 +02:00
jschoubben 60eae92fb1 Merge pull request 'Adoption mode as built: ADRs 0101–0103, issues 084–086' (#75) from feat/adoption-mode into main 2026-09-22 21:02:00 +02:00
jschoubben 7bbbfb158a ADR 0103: a unit is found when an administrator installed it or the machine uses it 2026-09-22 20:02:18 +02:00
jschoubben 87ae893407 Issue 087: the controller cannot tell that a node's host is too old for what it sends 2026-09-22 19:43:01 +02:00
jschoubben c74ea2a4a8 Records say what the build does: 0101 names only measured daemons; 0102 adds to lists and keeps what it writes over; 0103 names every held kind, the found-service rule, conflicting found rules, guards, and what of 0100 it replaces; designs 05, 09 and 17 in step; issue 084's diagnosis in its own file 2026-09-22 19:38:16 +02:00
jschoubben ca1f648973 Issue 086: taking a module narrows a port the predecessor served, without saying so 2026-09-22 19:00:59 +02:00
jschoubben 019184ec0f Issue 085: the packages port given at genesis is not a setting, and a later module can undo it 2026-09-22 18:17:21 +02:00
jschoubben 213ab898d6 ADR 0103: what an adopted node holds and what its guard refuses; the node host and connectivity designs name 0102 and 0103 2026-09-22 17:52:58 +02:00
jschoubben 84761f0600 ADR 0102: the mesh writes into a shared file, never over it; issue 084 located 2026-09-22 17:46:06 +02:00
jschoubben 1901a90d68 ADR 0101 accepted; raising a mesh names it 2026-09-22 17:43:13 +02:00
jschoubben 347bbce633 ADR 0101 proposed: a machine's own resolver does not make it in use, as measured on a fresh machine 2026-09-22 17:38:44 +02:00
jschoubben 6165a7ae02 Issue 084: taking networking on an adopted node restarts every container, and the held runtime file blocks pulling 2026-09-22 17:09:21 +02:00
jschoubben dfadfd23c0 Merge pull request 'ADR 0100 (proposed): a node in use is adopted before it is converged' (#74) from feat/adoption-mode into main 2026-09-22 16:35:32 +02:00
jschoubben 3d4ab23830 ADR 0100: the machine's own traffic is known by its interface, not its source address 2026-09-22 16:35:25 +02:00
jschoubben 37252f9c3e ADR 0100: the guard lets the machine itself through; in use is a non-loopback listener; openings say from where; 09 in step with the flip 2026-09-22 16:34:53 +02:00
jschoubben f3152d827f ADR 0100 after re-review: the bus and registry stay reachable for enrolment; the mesh guards the store in a table that only refuses; a machine in use defined; the flip refuses while a found container is held; held containers and returning to adopted spelled out 2026-09-22 16:32:37 +02:00
jschoubben 02c40bcab4 ADR 0100 after review: found means unrecorded; assigning prepares, taking cuts over; openings through the found firewall on both paths; the mesh guards its own ports; ports kept as node settings; a converged genesis refuses a machine in use; designs 05, 07, 08, 09 and 17 in step 2026-09-22 16:28:31 +02:00
jschoubben 111456abb5 ADR 0100 accepted; the node host, connectivity, the node lifecycle and raising a mesh amended for a node adopted before it is converged 2026-09-22 16:20:27 +02:00
jschoubben 5fc3cbde4c Research 012: migrating a node that is in use, measured on the control-node; ADR 0100 proposed — a node in use is adopted before it is converged 2026-09-22 16:16:28 +02:00
jschoubben 21baf397a8 Merge pull request 'Issue 083 resolved: nothing the control queue carries is lost while the store restarts' (#73) from multiple-fixes into main 2026-09-22 14:43:00 +02:00
jschoubben 4de74880cb Issue 083: the proof, the replay caveat, and what is not closed, as three reviews found them 2026-09-22 14:39:29 +02:00
jschoubben 266ee34b5c Issue 083: the diagnosis describes held messages, the enrolment's order, and what is not closed 2026-09-22 14:24:31 +02:00
jschoubben 2d45aa5f42 Issue 083 resolved: nothing the control queue carries is lost while the store restarts; an enrolment claims its token and spends it last 2026-09-22 14:11:44 +02:00
jschoubben 4e74cf29e4 Merge pull request 'Issues 081 and 082 resolved; 083 opened' (#72) from multiple-fixes into main 2026-09-22 13:53:16 +02:00
jschoubben 0d40918e92 Issue 081: proven by the two-node bed, and what it found about baserow's data directory 2026-09-22 13:53:00 +02:00
jschoubben f6ded3102d Issue 083 opened (other control messages lost while the store restarts); 082's diagnosis carries its review 2026-09-22 13:34:05 +02:00
jschoubben 9a75f2b6a9 Issue 082: a report that arrives while the store restarts was lost; 081's diagnosis corrected on review 2026-09-22 13:25:33 +02:00
jschoubben 5d1a0372cb Issue 081 resolved: neither cache consumer can keep its keys under its login, so neither takes the shared cache 2026-09-22 12:27:12 +02:00
jschoubben f1e3925009 Merge pull request 'Issues 079 and 080, found by running the large mesh bed; 074's addendum' (#71) from multiple-fixes into main 2026-09-22 02:19:17 +02:00
jschoubben 35e44c3e5e Issue 081 opened (a cache consumer does not use its login); 079 and 080 diagnoses carry the review's consequences 2026-09-22 02:04:29 +02:00
jschoubben f760bf632c Issue 080: a cache grant let the consumer flush the server; 079: the names follow the resolver's rule for the private network 2026-09-22 01:39:29 +02:00
jschoubben e31e077fc8 Issue 079: the suffix is handed down, not written twice 2026-09-22 01:27:53 +02:00
jschoubben 17cc36069f Issue 079: every machine named twice over, found by the large mesh bed; what running that bed cost, on 074 2026-09-22 01:13:29 +02:00
jschoubben 446553dd1f Merge pull request 'ADR 0099; issues 077, 078 and 074 resolved; designs 08 and 20 amended' (#70) from multiple-fixes into main 2026-09-21 23:58:50 +02:00
jschoubben 3ea5e47c21 Review: 077 says what closed it and what did not; 078 names module issue; 074's retired fixture; ADR 0099's scope 2026-09-21 23:58:15 +02:00
jschoubben ead8913a74 Issue 078: module issue's orphan account, refused before it is made 2026-09-21 23:46:51 +02:00
jschoubben faa7196ad6 Issue 074: opened date restored 2026-09-21 23:44:22 +02:00
jschoubben 8d159cee84 Issue 074 resolved: the declared list is empty of WEARING; what the last three cost 2026-09-21 23:43:31 +02:00
jschoubben 0e0f0298c6 ADR 0099: a step that runs once names what it reads; issues 077 and 078 resolved; designs 08 and 20 amended 2026-09-21 23:33:47 +02:00
jschoubben c2a81cbb20 Merge pull request 'ADR 0098; issue 076 opened and resolved; ADR 0097's base refusal live; ADR 0096 proven against the public hub; 074 down to one bed' (#69) from multiple-fixes into main 2026-09-21 22:58:26 +02:00
jschoubben 6bc9df4b49 Review corrections: 076 and ADR 0098 say what the authority could and could not do; issues 077 (a fetched fact is fetched once) and 078 (secret accept takes any name) opened 2026-09-21 22:55:11 +02:00
jschoubben 6d3cb60949 Issue 076: the route-forwarding bed proves ADR 0098; what the run taught about the overlay 2026-09-21 22:44:24 +02:00
jschoubben 75af72ca27 ADR 0096: the copy is proven against the public hub by the genesis bed 2026-09-21 22:32:03 +02:00
jschoubben 252c6042e8 ADR 0098: a fact a provider makes at first start is fetched from it; issue 076 resolved; design 08 amended; 074 down to one bed 2026-09-21 22:27:32 +02:00
jschoubben bc373797c7 Issue 076 opened: a served fact made at first start cannot be served; ADR 0097's refusal of an undeclared base is live 2026-09-21 22:16:15 +02:00
jschoubben 559683318b Merge pull request 'Multiple fixes: issues 064, 066 and 020 resolved; 074 narrowed to two beds' (#68) from multiple-fixes into main 2026-09-21 22:13:20 +02:00
jschoubben ac11bea77f Issue 020 resolved: the symptom was the bed's, proven against Pebble and step-ca alike 2026-09-21 22:12:51 +02:00
jschoubben c1a10dc3f8 Issue 066 resolved: a file and its reader are guarded by a gate, proven by the coupled-pair spike; design 20 says so 2026-09-21 22:04:14 +02:00
jschoubben e993233004 Issue 074: six of the ten beds retired rather than converted 2026-09-21 21:53:11 +02:00
jschoubben 2d77911511 Issue 064 resolved: the package half placed by the genesis run, the image half by ADR 0097 2026-09-21 21:52:32 +02:00
jschoubben f11824d299 Merge pull request 'Multiple fixes: ADRs 0094–0097; issues 069, 049, 046 resolved; 064's image half decided' (#67) from multiple-fixes into main 2026-09-21 21:05:13 +02:00
jschoubben 43ca9ce0ae ADR 0097: an undeclared base is said, not yet refused 2026-09-21 20:48:19 +02:00
jschoubben 71072e240d ADR 0097: a vendor image is a declared build input; issue 064's image half decided; design 18 amended 2026-09-21 20:45:36 +02:00
jschoubben 8d981e21b1 ADR 0096: an upstream image is copied between registries; issue 046 resolved; design 18 amended 2026-09-21 20:43:07 +02:00
jschoubben 39a01b7f7e ADR 0095: the control plane is the way to ask a module; issue 049 resolved; design 19 amended 2026-09-21 20:38:23 +02:00
jschoubben acc9824949 ADR 0094: a module may hold several secrets from one provider; issue 069 resolved; design 24 amended 2026-09-21 20:29:29 +02:00
jschoubben a259d1292f Merge pull request 'ADR 0093: a fixture that runs a module's runtime carries its name; 074 diagnosed; 075 resolved' (#66) from feat/mesh-tests-and-runtimes into main 2026-09-21 19:36:44 +02:00
jschoubben 76cd91413e ADR 0093: a fixture that runs a module's runtime carries its name; 074 diagnosed, two beds converted; 075 resolved 2026-09-21 19:30:35 +02:00
jschoubben 6427afe6ae Merge pull request 'Issues 072 and 073 name their merges, the branches being gone' (#65) from chore/merges-named into main 2026-09-21 19:25:06 +02:00
jschoubben d92068be7b Issues 072 and 073 name their merges, the branches being gone 2026-09-21 19:24:28 +02:00
jschoubben 6a35e5a58c Merge pull request 'Multiple fixes: status vocabulary aligned, 065/026/070/007 resolved, 066/069/049/046 located, 064 diagnosed' (#64) from feat/multiple-fixes into main 2026-09-21 19:23:34 +02:00
jschoubben bf4d99843f Merge pull request 'Issue 072 diagnosed: genesis registers the manifest its build produced; design 17 says so' (#63) from feat/one-controller-manifest into main 2026-09-21 19:23:31 +02:00
jschoubben 702289ad7d Merge pull request 'ADR 0089: a bed reads the catalogue it proves; issue 073 diagnosed; issues 074 and 075 opened' (#62) from feat/beds-read-the-catalogue into main 2026-09-21 19:23:16 +02:00
jschoubben 6ee95a3f31 ADR 0090: a failure is the same by resource id, not by the host's words 2026-09-21 19:23:04 +02:00
jschoubben 159a583cf0 Issue 069: the count is the catalogue's, not the report's 2026-09-21 17:51:40 +02:00
jschoubben 62ff9a5bd3 Issues 046, 049 and 069 located, each with the decision its fix needs named 2026-09-21 17:50:42 +02:00
jschoubben 16442ecd23 ADRs 0091 and 0092; issues 026 and 070 resolved; designs 18 and 24 amended 2026-09-21 17:50:41 +02:00
jschoubben ffb7fa52ec Issues 007 resolved, 066 located, 064 diagnosed 2026-09-21 17:48:38 +02:00
jschoubben fcf34727fe ADR 0090: a failure that repeats is said to be stuck; issue 065 resolved
The controller kept one report per machine, replaced, so a resource nothing can
ever apply looked like a failure that had just happened, every few minutes, for
ever. It now counts identical reports and status says stuck after three.
2026-09-21 17:43:02 +02:00
jschoubben a5e351f4cb Nine resolved issues name where their fix landed
The cycle check now asks a resolved issue for its owner, and these had none.
2026-09-21 17:39:34 +02:00
jschoubben e3279f4b5b Issue 054 names where its fix landed 2026-09-21 17:39:09 +02:00
jschoubben d52687d0c8 Issues 067 and 068 name where their fix landed 2026-09-21 17:38:48 +02:00
jschoubben 36d9b38a0d An issue is open, diagnosing, located, resolved or wontfix — nothing else
The playbook, the README and the status skill knew five statuses; the cycle check
knew a sixth, 'fixed', and not 'wontfix'. Eleven issues sat in the sixth for weeks
with their fixes shipped, one step short of closed. They are resolved; the check
refuses the word from now on and accepts the one the playbook allows.
2026-09-21 17:38:17 +02:00
jschoubben da6414c27c Issue 072 diagnosed: genesis registers the manifest its build produced; design 17 says so
ADR 0069 had already placed the controller's manifest in its own repository, and the
raising design called the catalogue copy a thing to remove. The installer's step 3 was
handed that manifest by the build and step 9 read a second copy anyway.
2026-09-21 15:18:32 +02:00
jschoubben 5dcb9dfbf6 ADR 0089: a bed reads the catalogue it proves; issue 073 diagnosed; issues 074 and 075 opened
The end-to-end design held 'the run rebuilds what it tests' for binaries and images
and not for manifests. The beds' inline copies fell into three kinds; only the first
is a stale copy. The other two are named: a mesh test wearing a catalogue module's
name (074) and a stocked runtime image the run never rebuilds (075).
2026-09-21 14:35:08 +02:00
jschoubben 1c90777124 Merge pull request 'Issues 031, 035 and 054 name their merges, the branch being gone' (#61) from chore/migration-blockers-merges into main 2026-09-21 13:50:06 +02:00
jschoubben ba07b43b18 Issues 031, 035 and 054 name their merges, the branch being gone 2026-09-21 13:49:56 +02:00
jschoubben 1db93ffc64 Merge pull request 'ADRs 0087 and 0088; issues 031, 035 and 054 resolved; issue 041 surveyed; issue 073 opened' (#60) from feat/migration-blockers into main 2026-09-21 13:49:09 +02:00
jschoubben ea42b00d0a Issue 041: mongodb names its secrets' owner; the three conversions are proven 2026-09-21 13:47:39 +02:00
jschoubben 198e4db0a3 Design 07: the base ruleset closes the hub's port until the filter module derives one 2026-09-21 13:30:26 +02:00
jschoubben a438cbdac3 Issue 041: the declared exceptions, surveyed and partly converted 2026-09-21 12:53:14 +02:00
jschoubben 5fa35362ac Issue 073: beds carry copies of catalogue manifests 2026-09-21 12:31:04 +02:00
jschoubben ee67f69ef4 ADRs 0087 and 0088; issues 031, 035 and 054 resolved; designs 05, 07 and 18 amended
A seeded file is created once (0087); the foundation filters before anything
listens (0088); a machine becomes the last thing it was told (design 05).
Each says how it is checked.
2026-09-21 12:11:50 +02:00
jschoubben a211aecdfa Merge pull request 'ADR 0086 (a secret reaches a process as a file), the vault shipped, issues 041 and 072' (#59) from feat/secret-not-in-environment into main 2026-09-21 11:59:21 +02:00
jschoubben 5fb8f06a91 Issue 072: the controller's manifest exists twice; decisions index regenerated for ADR 0086 2026-09-21 10:34:17 +02:00
jschoubben ec7b6e0748 ADR 0086: a secret reaches a process as a file, and an exception is declared
Closes issue 041 by decision and by code on the same branch: the catalogue
engine refuses a secret in a container's env, and a secret-carrying env-file
unless the container declares its reason; the controller reads all six of
its credentials from files; design 13 states the rule and how it is checked.
2026-09-21 10:10:33 +02:00
jschoubben 7015bf58ac The vault shipped: as-is 06 describes it, design 24 is implemented, 071 names its merges 2026-09-21 10:10:33 +02:00
jschoubben aa9b5b99e4 Merge pull request 'The secrets vault: handoff, the amended decision, issues 069–071, and the reconciliation of open issues' (#58) from feat/secrets-vault into main 2026-09-21 10:03:28 +02:00
jschoubben 61f61b5d06 Reconcile the open issues against main
An audit of the six code repositories found eleven open issues fixed on main
with commits and beds to show (025, 027, 033, 036, 037, 040, 045, 047, 050,
052, 053), three partly (007, 026, 035), nine not (020, 031, 041, 046, 049,
054, 064, 065, 066) and one whose fix would live outside those repos (006).
Resolved ones name their evidence; partly ones say what remains; 041 records
that the exposure has widened since it was reported.
2026-09-21 01:52:53 +02:00
jschoubben ffd7592687 Design 24: the export says what an earlier key opens 2026-09-21 01:16:54 +02:00
jschoubben 1bce3f33a8 Designs 07 and 21 and issue 071: genesis now makes the root secrets
The installer half of the amended ADR 0085 is built and proven by the
genesis bed's root-secrets step; the foundation design closes its open item
and the installation design says what the installer does and what it still
cannot check.
2026-09-21 00:50:56 +02:00
jschoubben 8607d21110 Design 24: pair credentials are sealed to the operator too 2026-09-21 00:36:30 +02:00
jschoubben 050a2d08e8 Playbook 07: a feature worktree needs the siblings the lab reads
A bed run from .work/<slug>/mesh-lab derives mesh-tools and mesh-sdk by
sibling path and fails at once when the directory holds only the touched
repos. Detached worktrees on main, never symlinks.
2026-09-21 00:27:01 +02:00
jschoubben baa3351552 Amend ADR 0085: the vault is a foundation module and holds the root secrets
Recorded on the record, dated, before anything shipped against the sentences
that change. The vault is installed at genesis like the store and broker, one
per mesh, and holds every secret a module has for itself sealed a second time
to an operator key whose private half never enters the mesh — the break-glass
path the first version left open, without a key one place holds.

Design 24 says how; 07 and 21 say what genesis does not yet do; issue 071
names the fixed credentials the foundation is raised with today.
2026-09-20 23:55:01 +02:00
jschoubben 187389b7d1 Hand off the secrets vault to mesh-catalog; open issues 069 and 070
Design 24 flips to in-progress with mesh-catalog as its owner (playbook 04).
Starting the build surfaced two gaps the decision did not settle: a module
requiring `secret` receives exactly one value (069), and no command can
accept an operator's value into a consumer↔vault pair (070). Both opened as
issues rather than improvised around.

Also fills fixed-by on 067 and 068, which the cycle check refused as resolved
with no reference.
2026-09-20 23:08:28 +02:00
jschoubben 731027c009 Merge pull request 'Graduate issues 067 and 068 — provider scoping and the secrets vault' (#57) from multi-node/harden-and-prove into main 2026-09-20 21:30:12 +02:00
jschoubben fc4ab370d6 Graduate issues 067 and 068 to decisions and to-be designs
ADR 0084 (extends 0027) — a provision is served by a node-scoped provider the
consumer selects, defaulting to co-location; a module may instead carry a private
embedded instance that is not a provision. Design: 01-to-be/23-choosing-a-provider.

ADR 0085 (extends 0031) — a secret is a provision and the vault is the module that
provides it; a module's own local secret becomes an ordinary pair credential that
rotates through the existing machinery, while the controller's provisioning-credential
mint (0048) is unchanged. Design: 01-to-be/24-the-secrets-vault; doc 13 amended to
cross-link the non-pair secret.

Issues 067/068 marked resolved with amended-design set. ADR index regenerated;
records and index checks pass.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-20 21:27:29 +02:00
jschoubben d66579c8f6 Issue 068 — secrets have no owning module
Store and broker seats were made ordinary modules; secret-minting is still a
privileged property of the controller that no module owns. Propose the vault
become a module that provides a secret provision (generate/hold/rotate/backup/
audit), node-scoped like every other provider (issue 067), subsuming the three
secret paths and giving local-secret rotation a home.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-20 21:24:42 +02:00
jschoubben c4ff0478b7 Issue 067 — fold in the embedded-vs-provisioned axis
Naming which provider is only half of how a module gets a database. The other
half: a module may carry its own version/fork-pinned instance, module-network
only, no published port, not a provision — and the model has no word for it.
Add it as a second axis with its own open questions.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-20 21:24:42 +02:00
jschoubben 1ce9ffe79f Issue 067 — a provision cannot name which provider serves it
The mesh models provisions as mesh-scoped (one provider of a kind, a single
mesh-store). But node-specific services delivered to the mesh was the plan from
the start: both nodes already run their own postgres, SQL server, redis and
object store, and identity — currently single — already serves apps on a second
node. The model cannot express which provider serves a consumer, so it collapses
a deliberately per-node fleet to one. Provider scoping is a whole-mesh decision
across postgres/s3-bucket/oidc, not an SSO patch.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-20 21:24:42 +02:00
jschoubben df2c3f63a5 Merge pull request 'Issues 065 and 066 — the apply failure model's two gaps' (#56) from issue/065-066-apply-failure-model into main 2026-09-20 14:28:18 +02:00
jschoubben a9bc0c1cc5 Open issues 065 and 066 — the apply failure model's two gaps
Surfaced while reviewing the host apply loop for the controller/host
split discussion. 065: a permanently-failing resource retries for ever
with no escalation — reported, but never raises its hand as stuck.
066: a partly-applied declaration leaves a mixed state with no rollback,
which is harmless for independent resources and unexamined for pairs
that are only correct together. Both are questions HQ must answer, not
incidents — status open, no fix proposed.
2026-09-20 14:28:04 +02:00
jschoubben 6d28683898 Merge pull request '057 and 058 resolved — fixes merged and proven' (#55) from issue/057-058-resolve into main 2026-09-20 12:55:16 +02:00
jschoubben 82c03a8d10 057 and 058 resolved — their fixes merged and proven
The P1-sweep PR left both at 'located'; the fixes have since merged
(mesh-controller #31, mesh-tools #10) and the bed proves them. Flip to
resolved with their fixed-by.
2026-09-20 12:55:00 +02:00
jschoubben 31c5aee72d Merge pull request 'ADR 0083 and issues 057/058/060/063/064 — the P1 sweep' (#54) from issue/057-058-one-push-patient-runtime into main 2026-09-20 12:53:24 +02:00
jschoubben 05857f71f0 Accept ADR 0083 — one push leaves the mesh consistent 2026-09-20 12:52:55 +02:00
jschoubben 98e9e62fce Issues 060 resolved, 064 opened; 057/058 on this branch too
060: 44 of 70 modules now mesh-buildable (was 8) — the mechanical
majority, proven by direct builds and the bed. The structural
remainders: external-dependency fetch (new issue 064) and route-proxy's
cross-repo build context. 064 records the build environment's isolation
from public npm and Docker Hub.
2026-09-18 02:52:36 +02:00
jschoubben 0d5693cc2c ADR 0083: the cascade compares against what was last sent, not a snapshot 2026-09-18 02:38:00 +02:00
jschoubben 1572d74a18 Issue 063 — the foundation's ports are forwarded, resolved
The broker's port was accepted on input but not forwarded, so a joined
node reached a DNAT'd broker only until it restarted. Fixed in
mesh-controller (25e42b3), proven by the built-store-cross-node bed
adopting the foundation's broker (which restarts it) and the joined
node still receiving declarations.
2026-09-18 02:20:11 +02:00
jschoubben 9fd7e6c458 ADR 0083 proposed; 057/058 diagnosed and located
One push leaves the mesh consistent (the 057 decision, proposed for
acceptance); the shared runtime waits for its broker (058). Fixes on
mesh-control fix/one-push-is-enough and mesh-tools
fix/the-runtime-waits-for-its-broker; the built-store-cross-node bed
enforces both.
2026-09-18 02:10:01 +02:00
jschoubben 5703c86cf6 Merge pull request 'Issue 059 resolved — the broker credential resolves the seat' (#53) from issue/059-broker-address-seat into main 2026-09-18 01:13:13 +02:00
jschoubben d85d74d607 059 resolved — the broker credential resolves the seat
Fixed by mesh-controller PR 30 (79a9e17), rebased onto the merged
registry-reach train and proven by a green fresh run of the no-fake
two-node bed built from that commit.
2026-09-18 01:13:03 +02:00
jschoubben 46efb7ae98 Merge pull request 'ADR 0082 and issues 042/048/061/062 — the registry is reached by name and trusted by the overlay' (#52) from issue/042-048-registry-reach into main 2026-09-18 01:00:55 +02:00
jschoubben 4279e4914b 042/048/061/062 resolved — the network carries the registry trust
One green fresh run of the no-fake two-node bed is the proof: node2's
consumers open the store and broker the mesh built and adopted, over
the overlay. Fixed by mesh-controller PR 29, mesh-catalog PR 26, gated
by mesh-lab PR 34.
2026-09-18 00:26:50 +02:00
jschoubben a0acaad86d Issues 061/062 — two silent-success defects the no-fake bed surfaced
061: the broker module's provisioner never ran; its runtime container
named no command and the image default is the tool host. 062: a failed
artifact-store lookup composed the network without the registry trust,
turning a transient error into permanent silent state. Both located,
fixes on the 042/048 train branches.
2026-09-18 00:23:52 +02:00
jschoubben 08c814aa0a 0082 takes its homes: 042/048 located against it, the delivery design cites it
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 23:02:32 +02:00
jschoubben 32fa6b40f5 ADR 0082: the registry is reached by name, and the overlay is its security
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 23:02:18 +02:00
jschoubben 8df3469813 042 and 048 move to diagnosing — the registry-reach work begins
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:44:39 +02:00
jschoubben e7869d27bd Merge pull request 'Issue 060 — the mesh cannot build most of its own catalogue' (#51) from issue/060-catalog-not-mesh-buildable into main 2026-09-17 22:39:01 +02:00
jschoubben 0b002ef39c Issue 060 — the mesh cannot build most of its own catalogue
Only 8 of 70 modules carry a build section; the other 62 exist through an out-of-band
build script plus placeholder rewriting, so the mesh's own build-and-deliver path has
never run for them. Found by the no-fake bed; blocks it and the migration's delivery
assumption.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:38:48 +02:00
jschoubben 101848242f Merge pull request 'Issue 059 — the broker credential names the hub, not the broker's node' (#50) from issue/059-broker-address-review into main 2026-09-17 22:37:26 +02:00
jschoubben 5395b6a47f Merge pull request 'ADR 0081: decisions are in the chain — no orphan records' (#49) from process/decisions-in-the-chain into main 2026-09-17 22:36:48 +02:00
jschoubben 90e4a368dc ADR 0081: a decision nothing cites is not yet in the chain
Decisions were the one link the cycle checks skipped, and measuring found 19 of 70 records
orphaned — the credential flow and the module-runtime cluster among them, which is how a
stale premise about a settled decision survived in working memory. cycle.py now refuses an
accepted record nothing cites; the 19 got true homes (design frontmatter, the playbook that
implements 0021, META for the process records). The overview names the practice: spec-driven
development with provenance.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:36:33 +02:00
jschoubben d1b5210d83 Issue 059 — the broker credential names the hub, not the broker's node
Findings of an adversarial review of the 055 fix: hub!=broker conflated, membership tested
as has-address, silent staleness, portless silent fallback, two hubs unrefused. One review
claim recorded as disputed against live wire measurements.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:29:25 +02:00
jschoubben 1162f7fe25 Merge pull request 'ADR 0080: the development cycle is checked, not trusted — plus review fixes' (#48) from process/the-development-cycle into main 2026-09-17 22:28:22 +02:00
jschoubben becae7ba51 ADR 0080: the development cycle is checked, not trusted
The flow the process overview draws — idea/symptom -> decision -> to-be design -> code ->
as-is — was enforced by nothing. cycle.py now refuses a to-be design naming no decision, an
in-progress/implemented design naming no owning code, a located/fixed issue with no owner,
a fixed/resolved issue with no fix, and a graduated research overview that does not say what
it became. AGENTS.md carries the cycle and a where-to-look table so a fresh session (or a
cleared context) finds the chain in frontmatter instead of assuming it. Grounding the check
surfaced two real gaps, fixed here: the work-ahead design named no owning code, and research
003 listed one became target twice.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:28:03 +02:00
jschoubben 94fee5d849 Fix what the records review found
The glossary's authority page still named the controller's seat the-controller in two
entries, contradicting its own seat section after ADR 0079; issue 058's heading kept the
pre-renumber 059; 055's fixed-by named branches that stop existing after merge (now merge
commits/PRs) and its located-in listed file paths where the convention wants repos; 056's
located-in named mesh-host, which received no fix, instead of mesh-catalog; and the design
layer never said the one-store/one-broker property is enforced — 07-the-foundation and the
installation table now state the seats.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:24:57 +02:00
jschoubben 5856555390 Merge pull request 'Issue 058 — a provisioner runtime crash-loops until the overlay is up' (#47) from issue/059-provisioner-startup-resilience into main 2026-09-17 21:57:03 +02:00
jschoubben 0ffdf4e581 The issue takes the next free number, 058
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 21:56:49 +02:00
jschoubben dd5bf9cd8c Issue 059 — a provisioner runtime crash-loops until the overlay is up
Observed in the two-node bed: a cross-node runtime exits on 'timed out fetching the
broker's certificate' until the tunnel forms, then settles. The retry belongs in-process.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 21:56:40 +02:00
jschoubben 0e9d4ab86a Merge pull request 'ADR 0079: foundation seats are named after their servers; issue 056 resolved' (#46) from multi-node/foundation-seats into main 2026-09-17 02:16:11 +02:00
jschoubben ee7abe0a8e ADR 0079: foundation seats are named after their servers; issue 056 resolved
The store and broker were "one per mesh" by convention only. Each foundation module now
claims a mesh-scoped seat named after the server it guards — postgres/mesh-store,
lavinmq/mesh-broker — and the controller's seat is renamed the-controller -> mesh-controller
so all three follow one rule. The resolver refuses a second holder, closing 056. Glossary,
the foundation and installation docs, and the decisions index follow the new name.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 02:15:08 +02:00
jschoubben f70973e222 Merge pull request 'Issue 055 resolved and proven; 056 and 057 opened' (#45) from multi-node/harden-and-prove into main 2026-09-17 01:44:59 +02:00
jschoubben 812dc3b303 Issue 055 resolved, and 057 opened for its silent operational edge
055 is proven end-to-end by the two-node bed: a consumer on a joined node reaches the
adopted broker over the overlay, its binding names the control-node, and its vhost is
minted. Three fixes — the broker's amqps port in the firewall, the broker credential
naming the overlay not the public address, and pushing the provider node after the remote
consumer arrives. The last is an operator ordering, not a code fix, and its silent-failure
edge (a cross-node consumer that never provisions and only crash-loops) is opened as 057.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 01:35:47 +02:00
jschoubben 33fc322f4d Issue 055 diagnosed — the module broker URL uses the public address, not the overlay
A two-node bed pins it: the firewall gap (broker's 5671 not in listens) is fixed
in mesh-catalog; the real bug is modules.go:294/build.go:231 building the broker
URL with the genesis public MESH_BROKER_ADDRESS, which a joined node cannot reach.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 01:02:07 +02:00
jschoubben bec1dd0082 Issue 056 — an adopted module assigned to a second node raises a second server
"One postgres, one lavinmq" (ADR 0078) holds only by genesis assigning the
adopted modules to the control-node alone; nothing enforces it. A second assign
raises a divergent second server, silently. Open questions: a mesh-scoped
exclusive seat, or explicit adoption the controller refuses to place elsewhere.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 00:22:18 +02:00
jschoubben a32f7e50af Merge pull request 'Establish the repo for the completed Phase 0-3 build' (#44) from establish/phase-3-close into main 2026-09-17 00:05:20 +02:00
jschoubben 1111bd84d7 Establish the repo for the completed Phase 0-3 build
Settles the design repository now that the self-upgrade build is on main:
- Records the two decisions that shipped without a record — ADR 0077 (the
  controller/foundation/node vocabulary) and ADR 0078 (the store and broker are
  ordinary modules); accepts ADR 0075 and 0076, which shipped work rests on.
- Fills issue 051's amended-design and wires ADR 0078 into 07-the-foundation.
- Sweeps the repo rename (mesh-control -> mesh-controller) into the mutable docs
  now that the forge repo is renamed; updates the glossary note and repos.md.
- Fixes the six broken links from the design-doc renames, indexes the glossary,
  regenerates the decisions reading order.

Both checks (records.py, index.py) are green. Statuses stay honest: the build is
on main and lab-proven but not deployed as the production mesh, so the to-be docs
remain in-progress and the as-is layer (the hal mesh) is unchanged — graduation
to implemented + as-is belongs to deployment, not merge.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 00:04:58 +02:00
jschoubben 873cc4007b Merge pull request 'Glossary, the mesh-controller/foundation vocabulary, ADR 0076, Phase 3 closed' (#43) from issue/047-the-other-half into main 2026-09-16 23:25:51 +02:00
jschoubben 0073e52881 Phase 3 closed — the store and broker are ordinary modules, issue 051 fixed
Marks WBS 3.2/3.3/3.4 done and resolves issue 051: the foundation's store and
broker are adopted in place as the postgres and lavinmq modules, upgradeable
through their stated windows, source-tracked by status. A bare-metal mesh runs
one postgres and one lavinmq, proven 22/22 in the one-node lab. The two follow-up
gaps are tracked as issues 054 and 055.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 23:07:03 +02:00
jschoubben fcf3b6f4d5 Issues 054, 055 — the debt adopting the store and broker leaves
054: the adopted servers bind 0.0.0.0 from genesis but the packet filter is
installed later, so there is a window where they are open with only bootstrap
credentials. 055: the servers bind on the control-node and the one-node bed
cannot prove a consumer on another machine can reach them over the overlay.

Both are follow-ups to issue 051's adoption (WBS Phase 3), tracked rather than
rushed.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 22:04:28 +02:00
jschoubben 43f3565617 WBS: correct 3.1 phrasing — a module gets a database only if it asks
Not "every module's database"; modules request one via requires postgres-database.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 21:20:00 +02:00
jschoubben addfdd6940 WBS: Phase 3.1 done — the store is adopted as the postgres module
One postgres, adopted in place, proven 22/22 in the one-node lab. Notes the
pre-filter exposure window as a follow-up.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 21:14:13 +02:00
jschoubben 33a00d5656 Adopt the glossary's vocabulary in the mutable design docs
"control plane" -> controller and "substrate" -> foundation throughout
03-DESIGN, 00-META and the README, with 06-the-control-plane.md and
07-the-substrate.md renamed to 06-the-controller.md and 07-the-foundation.md.
The immutable 02-DECISIONS records keep their original wording (and links to
them are unchanged) — a term retired here may still appear there, which the
glossary explains how to read.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:48:52 +02:00
jschoubben f9f48fbbf7 Glossary: one name per thing, and the words we retired
Locks the vocabulary that kept drifting in conversation — controller (not
"control plane"), foundation (not "substrate"), node and control-node, seat /
bench / claim, package vs artifact. AGENTS.md points at it as the authority.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:22:31 +02:00
jschoubben b43b60b183 ADR 0076: the SDK is a published package the toolchain resolves by version
Records the decision the package-registry work turns on — the SDK is built on a
public base and published before the toolchain that consumes it, so nothing is
circular; mesh-tools stays the thin toolchain base but resolves the SDK by
version. Reconciles docs 12/17/22 and indexes the record.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 10:28:08 +02:00
jschoubben 17c2e061df Phase 1 was mostly a phantom: correct doc 19 and the WBS
The protocol spec claimed the envelope and grant drifted across implementations.
Inspection showed the wire agrees — envelope required headers match, the two
optional ones are legitimately optional, and the grant wire (the contributions
file) is identical on both sides. The disagreement was in dead types, now removed.

So phase 1's 'make them agree' work is done by deletion and correction. What
remains is a conformance fixture as prevention — pinning the envelope and the
contributions file so a future change that breaks agreement fails a test — and a
full per-capability suite is deferred until a third language actually needs it.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 00:07:26 +02:00
jschoubben a40595fa08 ADR 0074: correct the evidence — the live wire agrees
The record claimed the two implementations already disagreed. Inspection showed
the live wire agrees: the disagreeing grant types were dead (removed), and the
envelope's two extra headers are optional and set when relevant, not missing.

The danger was dead types contradicting the wire, not live disagreement — which
is a sharper reason for specifying the wire and checking against it, not a weaker
one. The model stands; the conformance suite's job is prevention rather than
repairing a present break.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 00:06:49 +02:00
jschoubben 6e1099a905 Reorder the work: implement and unit-test first, run the lab once at the end
The correction the operator pushed: stop running a 20-minute lab against a mesh
mid-transformation, debugging paths the next phase deletes. The base-build hang
is almost certainly the SDK resolving from a git URL inside a docker build (issue
053), which Phase 2 removes — so debugging it on the current shape is debugging
deprecated code.

Phase 0 folds in: the installer's own regressions are fixed and committed;
whether it runs green is the final acceptance test, after the phases that change
its build path are in. Faults that can be reasoned out of the code path are, by
reading rather than running.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 00:02:51 +02:00
jschoubben 2e1ae061b5 The work ahead: four phases, dependency-ordered, each ending at a run
Everything decided this cycle and not yet built. Phase 0 gets the installer
green, because nothing else is testable end to end without it. Phase 1 makes the
protocol one thing and fixes the Go/TS drift the installer's own provisioning
exercises. Phase 2 stands up the private package registry ADR 0014 assumes and
publishes the SDK into it. Phase 3 adopts the substrate so one postgres and one
lavinmq serve everything, which is the hardest and needs all three above.

Order is dependency, not preference. Each phase ends at a run rather than a
paragraph, because a phase that ends at a claim is how things went missing this
cycle without anything complaining.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 23:41:42 +02:00
jschoubben 03fc13b84b 051: one lavinmq, not two — the duplication is running, not latent
The lavinmq module raises its own server container and provisions vhosts on it,
separate from the substrate's mesh-broker. A mesh with the module assigned runs
two LavinMQ servers where one belongs — the exact AMQP twin of the two postgres
containers.

The module's own provisioner already assumes one server: it creates a vhost per
consumer, named for the login, isolated by the vhost boundary — the analog of
postgres's database-per-login. So the mesh's own control traffic is the / vhost
and every consumer's broker is a vhost beside it, all on one server.

Adoption therefore means the module does not run a server of its own: its server
is the substrate broker, adopted, and the module contributes the provisioner,
tools, events and bootstrap against it. The isolation model is already built;
what remains to design is only raising the one broker at genesis and then holding
it as a module, over the broker it is.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 22:31:17 +02:00
jschoubben a062f181ce The twelve-module table names distribution, as the catalogue does
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 22:03:39 +02:00
jschoubben 87f464cc5f A finished mesh holds twelve, and two rows were one module each
The substrate's store and the postgres module are the same thing. They were two
rows only while the substrate was a different KIND of thing — a store raised from
a bundle cannot provide postgres-database, so anything wanting a database needed
a second server. It is visible on any mesh built today: mesh-store and postgres,
two containers, the same image.

Which name survives is settled by the naming rule. Where a consumer speaks a
protocol the interface is the protocol, and "database" is not a capability. The
control plane's own queries use distinct on and on conflict, so the coupling is
to postgres and a "store" module would advertise a swap that fails the first time
anyone tries it.

The broker collapses the same way with a different outcome: amqp IS a protocol
that several implementations speak, so amqp is a legitimate provision and lavinmq
is one provider of it.

So adopting the substrate is not only an upgrade path — it is two rows of a
mesh's module list becoming one, twice. And it leaves nothing that is a specialty
after the pivot, which is the claim the whole design rests on and is not true
today for exactly those two.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 20:26:42 +02:00
jschoubben 5a71e8659f The host's unit exists; the installer just does not place it
Reported as 'the host does not survive a reboot', which reads as a mesh that
cannot come back. mesh-host/packaging/ ships nox-mesh-host.service and two
companions. The installer declines to place them because a unit file is a
packaging decision, and the lab starts the host with --host-in-background, which
says in its own help that it does not survive a reboot.

So the lab run failing this was the lab being honest, and the gap is the step
that puts a shipped unit on a machine — narrower and more fixable than what I
wrote.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 20:15:48 +02:00
jschoubben 1e8273eb19 Correct 0075: co-residence is not the exclusivity that matters
The first draft implied gitea and the registry could not share a machine, because
registry claims the-artifact-store at node scope and I carried that across to
gitea without asking what the claim is for.

A machine running gitea for git and packages alongside a registry serving
artifacts is an ordinary arrangement. They are different ports doing different
jobs, and nothing about one being the mesh's artifact store requires the other
not to exist.

The exclusivity that matters is mesh-wide and already expressed: provides at mesh
scope means two providers are two answers, and the resolver refuses until one is
assigned. Forbidding co-residence adds nothing and forbids something reasonable.

Whether registry should still hold that claim is left open rather than answered
from outside its manifest — it may be protecting something about its port or its
data directory that nobody wrote down.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 19:35:05 +02:00
jschoubben 199c2f20bb ADR 0075 — an artifact store is a provision, and a package registry is another
Two questions circling turned out to be one asked twice: should gitea be the
mesh's registry, and where does the SDK come from. The framing that dissolves
both is that artifact-store is already a provision and the registry already
provides it — so this was never about replacing a component. It is a second
provider of an existing provision, which this mesh has a mechanism for and uses
for certificate authorities already.

So: two provisions, because they are two jobs. artifact-store is content
addressed, pinned by digest, no versions and no ranges — what the mesh delivers
to machines. package-registry is an ecosystem's own, addressed by name and
version — what code resolves when compiled. Conflating them is how a mesh that
pins everything ends up rebuilding one commit into two different things.

The small registry stays the provider genesis installs, not because it is better
but because of what it is: a directory and one container, installable where there
is no database and no control plane. Gitea needs both, and the pivot needs
somewhere to publish before either exists.

Gitea also provides artifact-store, so a mesh may choose it — and choosing it
answers issues 042 and 048 by adopting something that already has accounts and
TLS, rather than reimplementing them in a registry that has neither.

Moving between providers is a designed act with a verification step that is easy
to skip and is the only thing between it and a mesh that cannot restart its own
control plane.

And the bootstrap still has no package registry when the first build needs one.
Named rather than solved, so the next person does not discover it.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 18:06:03 +02:00
jschoubben 94229d95eb The installation, written out in full
Every step from a bare machine to a mesh that maintains itself, in three phases,
with each step named as the installer prints it.

The point of writing it out is the shape it exposes. The installer owns twelve
steps and ends at a mesh that RUNS. Seven more turn that into a mesh that WORKS —
the shared base, a store that is a provider rather than the control plane's own
memory, the catalogue, the replay of what was built before the catalogue existed,
the control plane rebuilt through the module path, the private network with the
node actually placed on it, and the packet filter. None of those seven is the
installer's. They are things somebody types, which is why a test had to be
written to discover they were missing.

Machines arrive last, in phase three, because a machine joining a mesh that
cannot build anything proves enrolment works and nothing else.

And five things that are not yet true are named rather than implied: phase two is
manual, a second machine cannot pull what the mesh built, ADR 0014 assumes a
private package registry that genesis has not installed when the first build
needs it, the host agent does not survive a reboot, and nothing can contradict a
claim that a machine was installed this way.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 16:51:52 +02:00
jschoubben 27b2161379 The SDK's delivery is decided; the git URL violates it
ADR 0014 is accepted and unambiguous: each module consumes its dependencies from
the private registry, the mesh's own shared library included, and a cross-package
change is publish then consume.

So the git dependency at a pinned commit is not a mechanism under consideration.
It is the shared library being consumed a way the record rules out, and the lock
naming a sibling directory is what that looks like when nobody publishes. Issue
053 is reclassified from a question about mechanism to a violation with a
direction.

What stays open is narrower and real: which software serves the private registry,
and that a fresh mesh has none when the first SDK is built — ADR 0014 assumes one
exists, and at genesis nothing has installed it.

Recorded after arguing at length for a bespoke content-addressed alternative,
against a decision that was already made and that I had not read.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 16:50:40 +02:00
jschoubben bc271d4de0 A worked guide: one module, four capabilities, four languages
What the SDK contains, answered by exclusion as much as by inclusion. It is the
protocol and nothing else — no configuration loader, because configuration
arrives as files the mesh wrote; no API clients, because a Plex client changes
when Plex changes and that has nothing to do with any other module; no storage,
HTTP or logging, because the language has those. The test for anything proposed
is ADR 0039's: does editing it recompile unrelated modules, and does it change
often. Both, and it stays out.

Then the worked module: events in TypeScript, tools in Go, a provisioner in Rust,
a scheduled job in Python. Four artifacts, four toolchains, four processes, one
module — and each part is an ordinary project in its language depending on the
mesh SDK the ordinary way, so a laptop resolves what a build resolves.

And publishing a package as a module capability, which makes the SDK unspecial:
it is simply the first module that published a library. A Plex client belongs to
the Plex module because that is the only thing that knows when Plex changed.

Three things left open rather than papered over: which registry (the catalogue
holds verdaccio and a forge usually serves one too, and nothing says which is
ours), who may publish (a credential that does not exist), and what a range means
in a mesh where everything else is pinned by digest — a mesh that can rebuild a
commit and get a different library is a real change, and should be decided rather
than arrived at.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 14:12:01 +02:00
jschoubben b637fe106f The module protocol, specified per capability
A floor every implementation needs and three capabilities independent of each
other, so an SDK can implement the floor and events and be a real thing rather
than an unfinished one.

Written as a specification, which means it says what is required rather than how
anything is arranged — and says plainly where it describes behaviour that is not
yet true. Three places it does:

x-causation-id and x-schema are specified and emitted by nothing; the Go side
writes four headers and the TypeScript side declares six. A module may serve
tools and may not call them, because a caller needs a reply queue its account may
not declare. And the two implementations disagree about what a grant carries — in
TypeScript consumer is the module, in Go it is the node and the module is From.
One word, two meanings, in two halves of one mesh.

Naming those in the specification rather than leaving them for conformance to
discover, because a specification that only described what already works would
have nothing to say about the things most likely to break.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 13:45:25 +02:00
jschoubben 033a5d384c Issue 053 — the SDK is pinned twice and the two disagree, and ADR 0074 accepted
The tool runtime's manifest names the SDK as a git dependency at a pinned commit;
its lock file names a sibling directory that exists on one workstation. It builds
only because the recipe runs npm install, which tolerates a lock disagreeing with
its manifest and re-resolves from the manifest — the one command that hides this.

It matters because a lock exists to make a build reproducible and this one
describes one machine, and because it is the first thing a fresh mesh builds: the
toolchain carries the SDK and everything with code of its own compiles inside it,
so a dependency resolved differently on the build machine than on a workstation
is a difference in every module the mesh will ever build.

And it is about to be copied. Each language's toolchain will carry that
language's SDK the same way, so the shape is worth settling before there are four
of them.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 13:44:02 +02:00
jschoubben 89302aa3e0 The protocol is split per capability, and an SDK implements it
Two corrections, both from the operator and both better than what was written.

An SDK is an implementation of the mesh's module protocol in one language, and
nothing more. The first draft defined it by the test it passes, which describes
how you check one rather than what one is — and leaves it sounding like a library
that helpers could accumulate in.

And the protocol is split per capability, which was missing entirely. A module
that only consumes events uses the event capability; one that serves tools uses
the tool capability; a provider uses provisioning. Nothing about consuming an
event requires knowing how a grant is answered, so an SDK need not implement all
of it to be real.

That has a precedent here: a host declares which resource kinds it can apply, and
a partial host is a real thing rather than a broken one (ADR 0005). An SDK
implementing the floor and events is exactly as legitimate, and a module written
against it is a module that does events.

Which changes what adding a language costs. A Rust SDK doing connection and
events is useful the day it exists, with tools and provisioning following when
something needs them — rather than a language being unsupported until it is
entirely supported, which is what makes adding one a project instead of a
contribution.

Conformance is therefore per capability too: a monolithic pass or fail would make
a partial implementation indistinguishable from a broken one, which is the
distinction the whole thing rests on.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 13:40:15 +02:00
jschoubben 3d693a8735 ADR 0074 — the wire is specified; an SDK is what passes the suite
ADR 0039 settles what belongs in an SDK. It does not say what happens when there
is more than one, and there already is: the contracts are expressed as Go types
in the control plane and host and as TypeScript types in the SDK, and nobody has
felt it because both live in one repository.

They already disagree. The provision's field is "resource" in one and "Provision"
in the other; "consumer" means the module in one and the node in the other; the
envelope declares six headers on one side and emits four on both — the missing
two being x-causation-id and x-schema, the second of which is exactly what a body
needs in order to change shape without silent misreads.

That class of failure does not announce itself. Two implementations disagreeing
about an envelope do not fail to compile — they ignore each other's messages, and
a mesh where a module stops reacting looks like a mesh where nothing happened.

So the decision is to specify the wire rather than share the types, because the
shapes are the easy half. What two implementations actually disagree about is
behaviour: queue naming and durability, which headers are required and what an
unknown one means, taking identity from the sealed credential rather than the
environment, dedup on an id only the emitter can make, pinning a fingerprint
rather than trusting an authority.

And the suite is executable rather than prose, because a specification nobody can
run is a document two implementations drift from while both believe they conform.
The two existing implementations are the first made to pass it — a suite only new
SDKs must satisfy would certify every future language against a disagreement that
is already here.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 13:26:57 +02:00
jschoubben 411f0680b8 Document the whole module surface, and keep it true with a test
A reference table goes stale the day somebody adds a field, so this one points at
mesh-catalog/modules/showcase — a module that uses all of it — and a test that
fails when it stops doing so. Read the module when the table disagrees with it.

Two rows in the coverage survey were stale because of this week's work: systemd
units were a file plus a service, which made every author write unit syntax and
is why "process" exists; and building from source was images only, where a
bundle now names a language and lets the mesh choose the toolchain.

And the two rows at the bottom of the resource table are the interesting ones.
"action" is refused to modules outright — the link may not carry a command, so a
module needing something done ships a program that reconciles. "service"
installs no unit by design, right for software shipping one and wrong for code
the mesh built, which has none until the mesh writes it.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 13:06:16 +02:00
jschoubben 8a7328c282 Correct the design: archives already work, compiling is what is missing
The first draft said the builder refused archives. It does not. An archive is
packed deterministically, hashed, published by digest, fetched by the machine and
unpacked — the whole path exists. Only the local builder used at genesis refuses
one, and deliberately: an archive is bytes that mean nothing until something
serves them, and at genesis nothing does.

What an archive cannot do is compile. Its source is a directory packed as it
stands, so shipping compiled output means compiling somewhere first, which means
a Dockerfile — the burden this document is about. The gap is not the artifact
kind. It is that no recipe both builds and packs.

Found by reading the builder rather than the manifest schema, which is where the
first draft's claim came from.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 01:54:43 +02:00
jschoubben 5d1e6d0b09 Design: building a module, and why the recipe cannot always be a Dockerfile
12-a-module-repository says what a module may build and where it goes. Nothing
said how a build is MODELLED, and the model is the problem: a recipe is implicit,
singular and always a Dockerfile; a toolchain is not modelled at all, arriving as
two build arguments the module hand-writes; a language is not a concept; and an
archive is declared in the manifest and refused by the builder.

The cost is measurable rather than theoretical. Adding a module with its own code
means repeating an incantation - two ARG bases, a specific working directory so
the SDK resolves upward, the compiler invoked by absolute path because the usual
symlink is resolved away when the base is assembled, a second stage, an env var
naming the entrypoints. Most of the catalogue is unconverted, and two conversions
done in one session were each wrong twice with a working example open.

So: recipe becomes explicit with three kinds, and toolchain becomes derived from
a declared language rather than written by every author. A Dockerfile stays, and
stops being compulsory - it is right for software needing a particular base and
wrong for "compile my module's code", which is the same operation every time.

The cost is stated before it is chosen: every language is permanent, and the
contracts are already expressed twice - Go structs and TypeScript types kept in
step by hand. A second language makes that drift. So language-neutral contracts
come first, or the drift gets worse while hiding.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 01:51:42 +02:00
jschoubben b163ed1fcc Issue 052 — the firewall closes the port the mesh runs on
The packet filter generates its rules from what modules declare they listen on.
The broker is not a module, so it declares nothing, so its port is not opened.
Every machine dials that port to enrol and to receive every declaration it is
ever sent.

Invisible where it is assembled and fatal on the next machine: a mesh of one
never dials its own broker across the network, so the ruleset looks right. The
first machine to join is refused at the packet filter during enrolment, several
steps from anything that reports it — and assigning the firewall before joining
machines is both the natural order and the one that breaks.

This is 051 in a second place. That issue says the substrate cannot be updated
because the mesh holds no record of it; the same absence means the firewall
cannot know it exists. Ssh is already a floor for the same reason — a machine
nobody can reach is a machine nobody can repair — and the broker may be the same
class of fact.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:59:33 +02:00
jschoubben f52895a646 Issue 051 — the mesh can update everything except what it depends on
The store and broker come from a bundle the installer writes once, with images
pinned in it, and nothing can change them afterwards: no build, no version to be
behind, no roll-out, and no way to report being out of date, because the mesh
holds no record of them as modules at all.

That is backwards. They are what everything else depends on, so their updates
matter most, and they are the only things with no mechanism to deliver one. A
mesh with a year-old broker reports itself entirely current.

The fix probably already exists: the control plane is carried, raised and then
adopted as an ordinary module pinned to what is running. Nothing in that pattern
is specific to the control plane. It would also remove a duplication visible on
any one-node mesh — the same postgres image running twice, because a store that
cannot be a module cannot provide a database to anything.

Recorded with the two hard parts stated rather than waved at: upgrading a store
the control plane is reading from, and upgrading a broker over the broker.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:44:10 +02:00
jschoubben 5d13f9c83d Issue 050 — the catalogue knows nothing built before it started
A mesh raised from bare metal built six modules and its catalogue reported three:
exactly those built after it began running. Missing were the shared base, the
store, and the catalogue itself.

The hole is never random. On a fresh mesh the modules built before the catalogue
are by necessity the ones it needed in order to exist, so the foundation is
always what is absent, on every mesh, at the moment the graph is first populated.

It breaks the question the catalogue is for: build edges hang off the base, so a
catalogue with no record of it answers 'what must be rebuilt' confidently and
wrongly. And nothing reports the gap, because a catalogue cannot know what it was
never told — this was found by comparing its answer against what the mesh had
just been watched doing.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:16:23 +02:00
jschoubben 0aa8f62e0c Issue 049 — a module serves tools and nothing may call them
A module's broker account is scoped to what it declares it emits and consumes. A
tool call needs a reply queue, which that scope does not cover and should not. So
the account is right, the request is reasonable, and no account exists that can
make it — asking a module its own question, from its own container, with its own
credential, is refused.

It matters because a module's tools are its operator-facing surface: the
catalogue serves the five questions it exists to answer and nothing can reach
them. It is also why a running module keeps being mistaken for a working one — a
test that cannot ask anything checks a container is up, and that substitution has
hidden two faults this week.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:08:40 +02:00
jschoubben ee70d5b451 Issue 048 — a stated rule about the registry is enforced by nothing
The registry says every machine pulls from it and opens its port to the mesh for
that reason. A machine that tries is refused by its own container runtime: the
registry serves plain HTTP and anything but loopback is treated as HTTPS.

It has never failed, and that is the finding. Every proof that a machine can
fetch a mesh-built artifact was a proof about the machine that built it, where
the reference was loopback. The bed that uses a routable address gets away with
it because the harness writes the runtime's configuration before the mesh exists.

Kept separate from 042, which they are easy to confuse: that one is about not
being allowed to pull, this one about not being able to whatever the credential
says. Fixing 042 alone leaves a machine with a valid account for a registry its
runtime will not talk to.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 21:01:32 +02:00
jschoubben 3121e47245 Merge pull request 'Issue 047 — and the half that runs the other way' (#42) from issue/047-the-other-half into main 2026-09-14 15:09:36 +02:00
jschoubben 261064ec63 Issue 047 — and the half that runs the other way
The report covered ports the firewall cannot close. The matching fault is ports
it does close and should not: rules come from what modules declare, a migrating
mesh knows about almost nothing, and anything listening on the host is dropped.

No module declares an ssh port and the generator has no allowance for one. The
session that loads the rules survives on conntrack until it drops, and then the
machine is reached from a rescue console.
2026-09-14 15:09:20 +02:00
jschoubben 3f35153dd5 Merge pull request 'Issue 047 — the firewall does not cover published container ports' (#41) from issue/the-firewall-does-not-filter-published-ports into main 2026-09-14 14:59:31 +02:00
jschoubben 027c41d125 Issue 047 — the firewall leaves published container ports open
Rehearsed on three lab machines before doing it on an anchor: loading the mesh's
rules refuses an ordinary host port and leaves a published container port
reachable, same prober, same second.

It is deliberate — no forward chain, because dropping there would stop every
container — but traffic to a published port never reaches the input chain, so
the firewall is silent about it. The anchor publishes 38 such ports today, all
filtered by the system being replaced, so the cutover would open every one of
them while reporting a firewall that is up.
2026-09-14 14:59:11 +02:00
jschoubben 9359542990 Merge pull request 'Issue 046 — an upstream image cannot be mirrored into the mesh's registry' (#40) from issue/mirroring-an-upstream-image into main 2026-09-14 13:27:59 +02:00
jschoubben 24d52a8380 Issue 046 — an upstream image cannot be mirrored into the mesh's registry
Found while migrating the first module. Five approaches were tried and observed
to fail, including resolving the index and naming the platform at both ends; a
speculative fix was written and reverted rather than shipped, because it did not
make the mirror work.

One module uses this today, which is why it went unnoticed — and it is the shape
most of a migration wants, because the services being moved are third-party.
2026-09-14 13:27:42 +02:00
jschoubben e773bf4c34 Merge pull request 'Installing ends with a builder, and the record says so' (#39) from feat/installing-ends-with-a-builder into main 2026-09-14 12:34:43 +02:00
jschoubben 256884e671 Installing ends with a builder, and the record says so
The design said how the builder arrives was unsettled and that nothing
installed it — the one gap stopping a fresh mesh from producing anything. Both
are now false. What is still true is narrower and worth keeping separate:
nothing asks a raised mesh for the rest of the catalogue, and no bed asserts
that it could.
2026-09-14 12:34:26 +02:00
jschoubben aa35ba1594 Merge pull request 'A module names its base, so the mesh can act on what it notices' (#38) from feat/a-module-names-its-base into main 2026-09-13 23:58:23 +02:00
jschoubben 2eb65a17a3 A module names its base, so the mesh can act on what it notices
The mesh could already say which modules a base change invalidated, and could
not do anything about it: each recipe named one particular copy of the base by
fingerprint, and rebuilding produced the old one. Worse, the copy each named
existed only inside a throwaway lab, so those three modules could not be built
anywhere at all — and the line each replaced had the same fault.
2026-09-13 23:58:06 +02:00
jschoubben 68d83c5b24 Merge pull request 'Two graphs, the builder's arrival, and two findings from building on a live mesh' (#37) from feat/two-graphs-and-the-build-chain into main 2026-09-13 11:17:48 +02:00
jschoubben 8e7fac8bf1 Record what genesis now does, and how each rule is checked
The builder's arrival was the one rule the document said nothing checked. It is
checked now, by both genesis beds — and so is the thing that distinguishes a
built control plane from a carried one, which every earlier assertion accepted
either way.
2026-09-13 06:09:31 +02:00
jschoubben fdc61054e3 Genesis carries a builder, and the document says so
The section saying the change was decided and had not happened now contradicted
the section below it. It also records the argument that failed, because a reader
will otherwise ask the same question and reach the same wrong answer.
2026-09-13 04:26:01 +02:00
jschoubben aabaaf2bb4 Rewrite 0073: the registry does not move, and the argument fails
Written an hour ago claiming a produced image must be published before anything
can fetch it, so the registry had to precede the control plane. The premise is
false: the machine that builds the image is the machine that runs it, and the
temporary control plane names a built image exactly as it names a carried one —
by the digest of its own configuration, which requires nothing to have served
it. Building changes where the bytes came from, not where they are.

Rewritten rather than superseded because nothing has been built on it and
nobody has read it: a record that contradicts itself is a draft, not a decision.
The argument is kept, because it was asked for and a negative answer is the
result.
2026-09-13 03:22:05 +02:00
jschoubben 9f5ac38662 Settle how the builder arrives, and what it publishes into
The two questions the design record named as the one gap stopping a fresh mesh
from producing anything. They cannot be answered apart: a builder with nowhere
to publish has made a file on a disk.

The registry's role did not change — the answer to 'must it precede the control
plane' did, because the control plane's image is now produced rather than
carried, and a produced image must be put somewhere before it can be fetched.
2026-09-13 03:20:07 +02:00
jschoubben 3983188b8a Issue 044 resolved — the mesh builds its own floor
Cloned alone, built by the mesh, every module moved onto it, and the base then
changed for real: all three went stale naming what moved. A comment-only change
correctly makes nothing stale, which found a second bug — staleness compared
commits where it should compare artifacts.
2026-09-13 02:51:08 +02:00
jschoubben 15ac7fd8dc The base has no circularity — the builder does not stand on it
Written as an open question; it has an answer, and leaving it open would have
made the fix look harder than it is.
2026-09-13 02:25:07 +02:00
jschoubben 3a9ed9d3fd Two findings from building the catalogue on a live mesh
The runtime every module compiles against cannot be built by the mesh, so the
one rule that would catch it moving can never fire. And a container keeps the
values it was created with, so two good applies can leave it running on neither.
2026-09-13 01:49:02 +02:00
jschoubben a4866ca899 Merge pull request 'Two graphs, and a build chain that orders itself' (#36) from feat/two-graphs-and-the-build-chain into main 2026-09-12 23:09:42 +02:00
jschoubben 58ad0742d8 Two graphs, and a build chain that orders itself
Corrects ADR 0070, written an hour earlier, which had the control plane consuming
the catalogue in order to compose a declaration. That was written before the two
graphs had been told apart and creates a dependency that need not exist: a
catalogue that is down would leave the control plane unable to compose the thing
that would repair it.

The catalogue links module-versions to each other and does not know nodes exist.
The control plane links module-versions to nodes, and holds capabilities and
claims. They meet only when something is installed, and everything between them
travels as events over the broker, on durable queues, so nothing is lost when a
receiver is away.

Build order is not computed anywhere. The builder never consults the graph and
builds what it is asked for; the catalogue asks for the next build after the
previous registration, so ordering holds by construction. Its rule is a condition
rather than a schedule — rebuild once everything a module was built against is
current — which covers a chain and a diamond alike.

Left open: whether an upgrade is applied or merely noticed, and whether a module
on several machines upgrades on all at once.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 23:09:23 +02:00
jschoubben 90f7bcb209 Merge pull request 'Genesis clones from a mesh, and checks what it got' (#35) from feat/where-genesis-gets-its-source into main 2026-09-12 22:20:28 +02:00
jschoubben 72b0810c53 Genesis clones from a mesh, and checks what it got
ADR 0070 has the init builder clone the source and does not say from where, and
ADR 0067 had rejected building at genesis partly because the forge runs on the
mesh being rebuilt. That objection binds only when those are the same mesh, which
is true exactly once.

So genesis clones from a mesh by name, and if that mesh is lost the name moves to
another that holds a copy — recovery is a name pointing elsewhere rather than a
backup being restored, and every installation adds somewhere it could point.

It names a commit and checks what it got, because the forge a mesh installs from
is the trust anchor for everything that mesh will run. On 2026-09-11 that forge
was running a cryptominer and tampering with git operations in flight; nothing was
altered, but a mesh installing during those hours could not have known that.

What relationship a mesh keeps afterwards is left open and named, so that whoever
writes the init builder does not settle it by accident: a snapshot and then
independence, or a continuing upstream for core modules.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 22:19:41 +02:00
jschoubben 03c73e7698 Merge pull request 'The catalogue owns the module graph, and genesis builds rather than carries' (#34) from feat/the-catalogue-owns-the-module-graph into main 2026-09-12 22:09:38 +02:00
jschoubben 64ae75a480 The catalogue owns the module graph, and genesis builds rather than carries
The graph had no owner: what modules are, what they require, what they claim and
what they are built against all sat in the control plane because that is where it
was written first. The control plane's own test says otherwise — anything a single
machine could answer alone is not its work, and what a module needs requires no
knowledge of any node.

So the catalogue becomes a core module beside the control plane and the builder,
owning the graph and serving tools over it. The control plane consumes it, which
is the opposite of what the tiers suggest and is therefore stated rather than
inferred.

That makes the catalogue a fourth thing that cannot arrive through the ordinary
path, so the claim written this morning that the list was closed at three is
corrected. All four are answered by one mechanism instead: the installer carries
an init builder and the core modules are built on the machine, so what is carried
is a builder rather than a result and nothing is left without a route.

Two things are open and named rather than assumed: where the init builder clones
from, given the forge normally runs on the mesh it would be rebuilding, and what
it publishes into, given the registry is installed later in the order today.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 22:03:05 +02:00
jschoubben 0fa4d14f54 Merge pull request 'A module is a repository and a path, and installing is described to its end' (#33) from feat/a-module-is-a-repository-and-a-path into main 2026-09-12 16:54:22 +02:00
jschoubben 837b5df2f7 Withdraw 043: the capability existed and the wrong verb was used
A build machine was refused the build queue, and this was raised as a gap in what
a manifest can express. It is not: `builder issue` creates exactly that account,
three lines from the code being read at the time.

Kept rather than deleted, for the one real thing in it — the wrong verb succeeds
and reports success, producing an account that authenticates and can do nothing,
so the failure surfaces a layer away as a permissions error that reads like a
missing feature.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 16:51:37 +02:00
jschoubben 9da45c68d6 Propose that the lab takes requests, one at a time, from its own copy
Raising a scenario occupies the machine and the person who started it, and
running in the background against a working copy is worse than waiting: a run
reads that copy as it goes, so editing while it runs yields a result describing a
state that never existed.

Proposed rather than accepted. The load-bearing part is the restriction — a
request names a bed and a commit and nothing else — because a request that could
say what to install and where would make the lab a second way of installing a
mesh, which is the arrangement that just cost a year of late-found faults.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 16:51:37 +02:00
jschoubben ddf104f8aa A module is a repository and a path within it
The design said a module's manifest sits at a repository's root, full stop, which
means one repository per module. Nothing that exists is shaped that way: the
catalogue holds sixty-seven modules one to a directory, no code repository has a
manifest at its root, and the system being replaced has always built a module
from a repository and a path.

So the builder could be asked to build nothing that exists — pointed at the
catalogue it finds no manifest, pointed at a module's source it finds none
either. Recorded as a decision because it moves the core modules' manifests
beside their source, and corrects the design that said otherwise.

Also corrects, in the same document, how the three things the build loop cannot
produce actually arrive. They were written as though all three were carried in.
Only the control plane is: the registry is pulled from the public internet, and
the builder has no route at all — which is now stated as the open one rather than
implied to be solved.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 16:51:37 +02:00
jschoubben 3d939b5c77 Describe how a mesh is raised, because only a test did
The one complete account of standing a mesh up was an integration test, and a
fixture is free to invent what it needs — which is how a registry that exists in
no production hid two faults for as long as the lab existed.

Written from what the installer does, not what it should do: genesis and joining
are separate moments, the lab runs the installer rather than describing
installing, and three things that are not true yet are named rather than glossed,
including one rule nothing checks.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 16:51:37 +02:00
jschoubben c567410687 Merge pull request 'Genesis is a pivot, public routing is name-agnostic, and five issues the fake registry was hiding' (#32) from design/bootstrap-is-a-pivot into main 2026-09-11 21:56:49 +02:00
jschoubben 26bec28c60 Issue 042 — nothing gives a node an account for a registry
Found by deleting the lab's registry and giving the machines a real path out:
public images fetched, the operator's own could not be fetched at all. There is
no provision for a registry credential, no manifest field, and no step in
enrolment that establishes one.

It applies to the mesh's own store too, which today asks for nothing — a
decision that has never been written down as one, and so reads as an absence.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 01:36:00 +02:00
jschoubben cde00e1d5f 0054 and 0055 join the topic their subject already had
Both carried 'model access', which is not one of the six the reading order
knows, so neither had a place in it. 0050 — the record they extend, on the same
subject — is 'what runs on it'.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 01:06:59 +02:00
jschoubben c29accd1be Issue 039 resolved — the operator's images are pinned, and why that is a stopgap
Nine references now name the digest their tag resolved to. What closed is the
immediate fault; the open questions stand, because a digest in a repository is
wrong the moment anybody rebuilds — which is the reason the design wants the
repository to name artifacts and the mesh to hold digests.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 01:06:35 +02:00
jschoubben 47dc3bc8e5 Accept ADR 0066 and ADR 0067
0066 was proven on a four-node bed before it was ratified: one node setting
moved an entire domain, a routed name resolved inside the mesh and was issued a
certificate by the internal authority, and TLS verified against that authority
with no override. 08-connectivity rests on it and could not while it was
proposed.

0067 records what deleting the lab's registry exposed — that pinning quietly
required a registry before the thing that lets a mesh have a registry could
start — and the pivot that resolves it.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 01:04:21 +02:00
jschoubben 423fde8534 Place the two new records in the reading order, and fix a link the merge moved
Both carried a topic outside the six the index knows, so neither had a place to
be read in — 0067 had none at all. Both are 'the tiers', beside 0036 (bootstrap
ends at a usable mesh) and 0007 (connectivity), which is what they extend.

0067 cited 0041 for tier 0's property; on this trunk 0041 is events, and the
record it meant is 0005. A citation that resolves to the wrong record reads as
corroboration, which is worse than a dead link.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:48:20 +02:00
jschoubben 56d0dcc4ff Merge branch 'feat/routing-is-name-agnostic' into design/bootstrap-is-a-pivot 2026-09-11 00:46:04 +02:00
jschoubben 84c95c53af Merge branch 'issue/021-provider-port-published-on-loopback' into design/bootstrap-is-a-pivot
# Conflicts:
#	03-DESIGN/01-to-be/04-lab-installation.md
2026-09-11 00:45:59 +02:00
jschoubben e4a5d9b1d0 Lab installation: the reachability check joins the assertions it belongs with
The doc's own rule is that the lab verifies capability by outcome, never by
reading a setting. A path out is exactly that kind of claim — a route and a
policy can both read correctly while nothing gets through — so it is asserted by
fetching something, in the table with the rest.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:11:39 +02:00
jschoubben 1ce1c3d4e0 Lab installation: a path out, and why the daemon having one is not enough
A container runtime on the same workstation sets the kernel's forwarding policy
to drop, and the virtualisation daemon's own accept rules do not override it —
both are consulted and a drop anywhere is the answer. The machines then get
addresses and resolve names, because the daemon's resolver is on the bridge, and
discard every packet to anything real.

The sharpest 'available is not adequate' yet: nothing is misconfigured, nothing
logs, and it presents as every image pull hanging. A workstation that runs
containers is the ordinary case, so it is a prerequisite — verified by reaching
something, never by reading a setting.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:11:22 +02:00
jschoubben 8f23d4b114 0067: the artifact store may require nothing, not merely build nothing
029 says a module providing the artifact store may not build artifacts. The
pivot shows that is the narrow case: it may not require anything the store is
needed to deliver. A route-label migration gave the registry a public name and
a route requirement, and at genesis nothing provides a route — nor can anything,
since the routing stack needs images and images need the store.

The same cycle through a door the existing wording did not cover, so the rule is
widened where the bootstrap decision states it, with a check that would catch the
next one where it is written.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:10:39 +02:00
jschoubben f9e129f734 Issue 040 — the only description of how a mesh is stood up is a test
There is no installer. The complete account of standing up a mesh is an
integration test in the lab, and a fixture may invent what it needs — this one
raised a registry no production has and rewrote every image reference through
it, concealing both 039 and the fact that a first node outside the lab had no
bootstrap path at all. The bed being green said nothing about whether a mesh
could be installed.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:34:59 +02:00
jschoubben e228355a52 Issue 039 — the lab's registry was pinning what the catalogue left unpinned
Nine images across seven modules name a tag, not a digest. ADR 0006 forbids it
and the host refuses it by name — and the refusal has never fired in a bed,
because the lab pushed every image into its own registry and rewrote every
reference to the digest it had just assigned. The harness was supplying the
property under test. Found by deleting the harness.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:20:10 +02:00
jschoubben cd1653bfe9 ADR 0067 — genesis is a pivot through a temporary control plane
The control plane's image is built from source and pushed nowhere, so it has no
manifest digest; a registry assigns those. Pinning therefore required a registry
before the thing that lets a mesh have a registry could start — a dependency the
rule created by accident. The lab hid it by raising a disposable registry no
production has.

Proposed, not accepted.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:19:17 +02:00
jschoubben 5852b35ab9 Renumber the routing record to 0066 — 0056 was already taken
0056 is 'the authority is the control plane, not a database', drafted on the
in-progress record chain this branch was cut from before those three records
landed. Two files would have collided at merge, which is the kind of thing that
is cheap now and confusing later. The code written against it still says 0056
and is corrected separately.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:17:20 +02:00
jschoubben a4440c9acf Issue 038 resolved — corrected root cause (same-node served-port mismatch, not loopback) + fix reference
The real defect is same-node consumers announced the declared port instead of
the assigned/published one; the loopback observation was a stale pre-0038 build.
Fixed in mesh-control fix/same-node-provider-announced-port (c147a26) with a
regression test.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 23:06:58 +02:00
jschoubben 9ffe9e97c9 Renumber to issue 038 — 021 is already taken (same-node credentials, closed) and issues run to 037 across branches
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 22:46:11 +02:00
jschoubben d16b5294f6 Issue 021: narrow to mesh-assigned ports (bare decl binds loopback, explicit host mapping binds 0.0.0.0)
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 22:35:06 +02:00
jschoubben 315bd7118a Issue 021 — a provider is announced at a name its port is not bound to
A from:mesh provider is announced (per 018's fix) at the node's private-network
name, but its port is published bound to loopback, so consumers dialing the
announced <node>.internal:port reach nothing. Diagnosed from the lab: DB
consumers that require the database at startup crash-loop; the provider is
healthy on loopback only.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 22:29:04 +02:00
jschoubben e2afb3e144 ADR 0056 + connectivity: public routing is name-agnostic, resolved in-mesh, internally certifiable
A route contribution carries a label; the node carries its public domain; the
mesh composes <label>.<public-domain> and holds no name map. A granted route is
published into internal resolution so anything in-mesh (notably an internal ACME
authority) can resolve and reach it. That authority certifies routed names by the
same path a public one would, differing only in issuer and trusted root.

Records the decision as proposed and amends connectivity SS2/SS3/SS5 plus its Open
list with the lab findings behind it.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 22:24:43 +02:00
jschoubben 450a4d8605 Merge pull request 'ADR 0055 — model access is answered by a licence or a node that hosts the model' (#31) from feat/adr-0055-node-model into main 2026-09-07 04:39:25 +02:00
jschoubben 9dfa5ba2ac ADR 0055 — model access is answered by a licence or a node that hosts the model
Extends ADR 0050: model-access, a provision, may be answered by a NODE that
hosts a model (Ollama/vLLM serving an OpenAI-compatible endpoint) as well as by
a licence record. A node-answer delivers an endpoint (base URL + model, and a
key only if the server wants one), rides the ordinary serves/binds path, and
uses no adapter; the resolver already prefers a local model over a licence. One
provision, two answers — the mesh's own model sits behind the same interface as
a vendor's. Names the mesh-scope-alongside-licence limitation and its fix.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 04:39:10 +02:00
jschoubben b7a33a1e26 Merge pull request 'ADR 0054 — model usage is a vendor-neutral record at two grains' (#29) from feat/adr-0054-model-usage into main 2026-09-06 23:39:11 +02:00
jschoubben ec2d73bc06 ADR 0054 — model usage is a vendor-neutral record at two grains (licence + session)
Closes the usage-tracking half ADR 0050 left open. One row shape (0050's), recorded at two
consumer grains: the holding module (licence-level, e.g. Anthropic utilization%) and the agent
session (per-session tokens/cost, since a session IS a consumer per ADR 0026). Produced by the
vendor adapter; read on a schedule (0053); recorded as events the audit-logger keeps (0041/0042)
plus a queryable usage context store (0008); usage is not a credential and is recorded in the
clear. Reuses the session, schedule, event, and store the mesh already has. Accepted per the
user's choice to build the full feature incl. per-session token/cost.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 23:38:54 +02:00
jschoubben a361c86c3c Merge pull request 'ADR 0053 — a scheduled step is a container run on a recurring schedule' (#28) from feat/adr-0053-scheduled-task into main 2026-09-06 13:53:11 +02:00
jschoubben 12a311ba4c ADR 0053 — a scheduled step is a container run on a recurring schedule
The recurring twin of run-once (0052): a schedule modifier on the container shape, reusing
its security bound (no new host action/shape, strictly less than an action) and reversing its
gating rule — a scheduled step runs after convergence, does not gate the apply, and a failed
run is logged, not fatal. Unblocks kometa's sync and pollers. Accepted per direction to build
the primitive now.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 13:52:55 +02:00
jschoubben 897626a9f6 Merge pull request 'ADR 0052 — a step that runs once before a container (issue 037 / run-once primitive)' (#27) from feat/adr-0052-run-once-lifecycle into main 2026-09-06 00:05:04 +02:00
jschoubben a3b0e68ec5 Accept ADR 0052 — a step that runs once before a container
The run-once lifecycle primitive: run-once:true on the existing container shape,
gated by declaration order + exit 0, idempotent by digest. No new host shape, no
arbitrary command — strictly less powerful than an action. Resolves issue 037.
Verified sound and faithful (ADR 0005/0047); status proposed->accepted.
2026-09-06 00:04:21 +02:00
jschoubben 2405d72fb0 ADR 0052 (proposed) — an init step is a container run once to completion
A module can declare state but not a step that runs. mosquitto must seed its
dynsec admin into dynamic-security.json before the broker starts, or the plugin
aborts; the database providers need the same for first-boot migrations and
health gates (04-ISSUES/037). The old event-hook engine that did this was
powerful and flaky; this is the narrowest sound mechanism instead.

A run-once step is an ordinary container marked `run-once: true`: the host runs
it to completion, requires exit 0, and gates the apply on it — so what the
declaration places after it (the broker) starts only once it has finished.
Gating is by declaration order, not a resolved dependency (ADR 0005); the
completion marker is the recorded declaration digest (ADR 0018), so a re-apply
does not re-run it unless the declaration changed. No new host shape and no
arbitrary host command: strictly less powerful than an `action`.

Points 04-ISSUES/037 fixed-by/amended-design at the record; index regenerated;
records.py and index.py pass.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 23:52:26 +02:00
jschoubben c41deb55d2 Merge pull request 'ADR 0051 — shared data is the operator's; a module is granted access (fixes issue 036)' (#26) from feat/adr-0051-shared-data-operator-owned into main 2026-09-05 22:29:44 +02:00
jschoubben d5e12c82aa Accept ADR 0051 — shared data is the operator's; a module is granted access
Your decision, ratified: the media library (and shared/pre-existing data) is
operator-owned and external; a module declares access, not ownership; the host
mounts but owns nothing (no create/chown/reconcile/remove); an absent accessed
path is refused clearly; several accessors co-resolve. Status proposed->accepted;
index regenerated (records + index checks pass). Implementation lives on the code
branches (mesh-control/catalog/host), held for merge after the convergence fix.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 22:29:24 +02:00
jschoubben 220e3c72f5 ADR 0051 (proposed) — shared data is the operator's
Resolves 04-ISSUES/036: eight media modules each declared the shared
library and download directories as their own resources, and the
resolver's duplicate-owner refusal — right in general — would refuse the
stack's only sensible assignment the first time two landed on one node.

The decision, from the operator: shared, pre-existing data is
operator-owned and external. The mesh does not create, chown, reconcile
or remove it. A module declares it needs access to such a path (read or
read-write); the host mounts it and owns nothing. Several modules
accessing one path is normal — the duplicate-path refusal is about
ownership, not use. An accessed path absent at apply is refused clearly,
not created. Extends ADR 0030: the third case the host had no word for,
what it neither made nor configured and must not touch.

Point 036's fixed-by/amended-design at the record; mark it located in
mesh-control, mesh-catalog and mesh-host. Regenerate the decision index.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 22:22:30 +02:00
jschoubben 4430cc1748 Merge pull request 'ADR 0050 (proposed) — model access is vendor-agnostic [awaiting ratification]' (#25) from feat/adr-0050-vendor-agnostic-model-access into main 2026-09-05 21:42:28 +02:00
jschoubben 860d512e91 Accept ADR 0050 — model access is vendor-agnostic
Verified and ratified: model-access stays one vendor-blind provision; per-vendor
adapter keyed by licence.vendor (mirrors public-dns registrar providers); the
sealing-vs-central-rotation carve-out bounded to refreshable-grant vendors /
refresh token / manager node only. Status proposed -> accepted; index regenerated
(records + index checks pass); design doc note updated.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 21:42:20 +02:00
jschoubben 3a9b47918d ADR 0050 (proposed) — model access is vendor-agnostic; amend 14-model-access
Turn the completed vendor-agnostic analysis into HQ design. The model-access
provision stays one vendor-blind interface (extends 0024/0027); the
vendor-specific lifecycle moves into a per-vendor adapter keyed by the licence's
`vendor` field, mirroring registrar-scoped public-dns providers (0044), named at
the consumer's real coupling per 0040.

The crux is the sealing-vs-central-rotation carve-out: for refreshable-grant
vendors only, the manager node holds the refresh token encrypted at rest (a
bounded, declared exception), access tokens sealed per holder, refresh stripped
on delivery. Static-key vendors keep full sealing.

Amend 03-DESIGN/01-to-be/14-model-access.md with the adapter generalisation as a
proposed section (prose + diagram, no code); regenerate the decision index.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 13:43:01 +02:00
jschoubben 3b0aff7585 Merge pull request 'Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work' (#24) from reconcile-init-into-main into main 2026-09-05 12:27:11 +02:00
jschoubben 495db6dc89 Restore initialization's issue 003 — the status flip was based on a wrong premise
Init's 003 was already resolved with its own consolidated attribution
(03-DESIGN/01-to-be/08-connectivity.md); the re-homing overwrote it with the
session's ADR-0045 firewall attribution. Init is canonical and the firewall
decision is recorded in the ported ADR 0045 regardless, so 003 is restored
untouched. No initialization record is modified by this reconciliation.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 12:26:23 +02:00
jschoubben e269f9a185 Re-home this session's new ADRs (0039-0049) and issues (032-037) onto the consolidated scheme; flip issue 003; port repos.md sdk line + feature-branches playbook (07); regenerate index
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 12:24:07 +02:00
jschoubben 546caa31c4 Merge initialization into main — adopt its consolidated structure as canonical
The real work lived on initialization (consolidated decisions 0001-0038, the
fuller issue set 001-031, the control-plane/substrate/node-lifecycle/delivery
design, research 011/012, the checks tooling). main had diverged onto a stale
base and only carried this session's genuinely-new work. This merge makes
initialization's tree canonical on main; this session's 11 new ADRs and 6 new
issues are re-homed on top in the following commits. initialization is recorded
as a parent so its history is preserved in main's ancestry.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 11:59:01 +02:00
jschoubben db8be15a15 Merge pull request 'Issue 013 — a module cannot run its own code at a lifecycle phase' (#23) from feat/hq-issue-013 into main 2026-09-05 11:16:27 +02:00
jschoubben 9c8c4a3ee5 Issue 013 — a module cannot run its own code at a lifecycle phase
Surfaced converting the catalogue: a module can declare things that exist
(dir/file/network/container) but not a step that runs at a point in its
lifecycle. mosquitto's dynsec admin client must be seeded before first start;
the DB providers have nowhere for a migration or health-gate; it is the timing
face of issue 011. Framed as a missing module capability, not a defect. Records
the prior-art event hooks and their real warning — powerful but complex and
flaky — so the resolution avoids rebuilding that. Ends in open questions
(run-once resource vs general lifecycle hook, where the code runs, idempotency).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 11:15:53 +02:00
jschoubben 600a57aafb Merge pull request 'Issues 011 and 012 — what the manifest review could not fix' (#12) from issues/manifest-review into main 2026-09-05 03:52:32 +02:00
jschoubben 2c8616b650 Renumber manifest-review issues 008/009 -> 011/012
Numbers 008 and 009 were taken on main after this branch was opened
(008-provider-runtime-has-no-seal-key, 009-runtime-config-change-does-not-restart,
both merged). Renumber the seed-file-wipe and shared-directory issues to the next
free numbers so merging records two more issues rather than duplicating two.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 03:52:02 +02:00
jschoubben 69fa79c585 Merge pull request 'Playbook 06 — one feature, one branch, one MR per repo' (#22) from feat/feature-branch-workflow into main 2026-09-05 03:47:25 +02:00
jschoubben cabe873503 Playbook 06 — one feature, one branch, one MR per repo
Records the branching-and-merging workflow for code changes across the mesh
repos, written against a failure it names: branches and MRs opened per unit of
thought, treated as done when opened not merged, and named differently per repo,
so they pile up unmerged — one session left sixteen to consolidate by hand. The
rule is one feat/<slug> shared across every repo a feature touches, isolated in
.work/<slug>/<repo> worktrees off main, pushed and opened as one MR per repo only
when the whole feature is done, then merged promptly. Adds the ground-rule
pointer in AGENTS.md and the row in the process overview.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 03:40:39 +02:00
jschoubben 36936dc7a3 Merge pull request 'Issue 009 (fixed) + Issue 008 (resolved via ADR 0053) — module-runtime config & provider contract' (#21) from worktree-issue-provider-seal-key into main 2026-09-05 03:02:34 +02:00
jschoubben bfa08742f8 Consolidate hq: ADRs 0044-0052 merged in, and 0017/0049-0052 accepted
Brings the independent ADR branches (0044-0052) onto one branch so hq lands as a
single MR, and ratifies the five that were still proposed — 0017, and 0049-0052,
which are implemented and green in the lab. With 0053/0054 already accepted here, the
whole ADR chain 0044-0054 is accepted on this branch.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 03:01:06 +02:00
jschoubben b5a74178c3 Merge remote-tracking branches 'origin/design/adr-0044-sdk-boundary', 'origin/design/adr-0045-what-a-module-is', 'origin/design/adr-0046-events', 'origin/design/adr-0047-event-wire-shape', 'origin/worktree-adr-0048-module-broker-account', 'origin/worktree-adr-reachability-dns-firewall', 'origin/worktree-adr-config-is-the-assignments' and 'origin/worktree-adr-module-runtime' into worktree-issue-provider-seal-key 2026-09-05 03:00:24 +02:00
jschoubben 1a83ed28fa Merge pull request '011: measure the graph against every facet a module carries' (#11) from research/011-facet-coverage into main 2026-09-05 02:53:20 +02:00
jschoubben 984194e765 Accept ADR 0054 (slug) and resolve issue 010
ADR 0054 accepted with option E (a declared slug). Issue 010 resolved: the login fits
via the slug, and the minted secret shrinks to 40 chars for S3's secret-key limit —
both halves of an S3 credential now fit the tightest backend.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 02:52:03 +02:00
jschoubben af5c939d16 ADR 0054: add option E (a declared slug) and recommend it over A+B
Implementing A (bound the identity at 20) revealed the readable budget is node+module
<= 14 chars — so tight that the catalogue's own test names (workstation+keycloak, 25)
compact to an opaque hash. B's fallback would fire for the common case, not the rare
overflow, inverting A+B into mostly-opaque identities. Option E — an optional short
slug a module/node declares, preferred over the cleaned name — is the escape hatch B
wanted to be without the opacity: legible because a person chose it, and it makes an
early refusal palatable (refuse on the slug field, not the machine's name). B dropped;
E recommended over a bound of 20, composing with C later if needed.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 02:30:11 +02:00
jschoubben 89b65dd4c0 ADR 0054 (proposed) — a consumer's identity is bounded by the tightest backend
Sketches the options for issue 010: the mesh's identityLimit (63, postgres's) is not
the shortest among the backends the derived name reaches — S3's is 20 — so CheckIdentity
lets an over-long access key through and minio fails at provision time. Options: bound
by the true minimum and refuse at assignment (recommended, with a compact fallback held
in reserve), per-interface bounds, or a provider-generated identity (rejected — breaks
"the mesh says the identity once"). Links issue 010 to it.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 02:17:46 +02:00
jschoubben 75887d0bc5 Issue 010 — the mesh's derived login does not fit every backend's identity rules
Found doing the per-backend provider e2e: redis and postgres accept the mesh's `as`
(mesh_<node>_<module>) verbatim, but minio's S3 access key is capped at 20 chars and
`as` is 22, so the provisioner cannot create the service account. `as` is doing two
jobs — a stable identity the two ends agree on, and a literal identifier a backend
must accept — and those are not always the same string.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 01:58:11 +02:00
jschoubben e5f4af8cf2 Accept ADR 0053 and resolve issue 008 — provider contract implemented and proven
ADR 0053 accepted; adds the scope boundary the umami rework surfaced (credential
provisions vs data provisions — analytics' generated siteId return is left to a
separate decision) and records the lab proof. Issue 008 marked resolved: the sdk
harness and the four adapters are reworked, the symmetric seal removed, and
provider-uses-mesh-credential is green.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 00:27:56 +02:00
jschoubben e27f8b1b7d ADR 0053 (proposed) — a provider creates the credential the mesh minted, and seals nothing
Issue 008's trace confirmed the premise in control-plane code: the mesh already
mints one password per consumer/provider pair and delivers the provider its copy
(SecretFor/SecretsFrom/grantsFor -> Grant.Sealed; the receives contribution carries
As + Secret). The provisioner's symmetric seal is an orphaned, contradictory second
model. ADR 0053 corrects the provider contract in one place (the sdk harness):
providers create the resource with the mesh-supplied login and password and drop
seal/key/return entirely. Reframe 008 as contract-first (every provider, not four).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 00:06:50 +02:00
jschoubben 332c334767 Issue 008 — sharpen: the provisioner's seal model is orphaned, not just undelivered
A cross-repo trace showed nothing writes the provisioner's grant-request files,
nothing reads its sealed credentials, and no consumer unseals — while the mesh
already mints and delivers provider/consumer credentials asymmetrically with no
shared key. The fix is to drop the symmetric seal and have providers consume the
mesh-minted password, a breaking provider-contract change that wants an ADR.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 23:40:20 +02:00
jschoubben 47b0ca5c90 Issue 009 — settings change does not restart a container-hosted runtime
Found rolling the runtime out to the catalogue: a module's runtime reads its
settings-merged config file once at start, but a container is only recreated on a
spec change, and file content is not part of the spec. So updating settings
re-renders the file and nothing re-reads it — ADR 0051's "on the fly" holds only
for config set before first start. Services have restart-on; containers do not.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 23:06:48 +02:00
jschoubben d1aa4254c9 Issue 008 — a provider runtime has no seal key the mesh can deliver
Found building the module-runtime vertical slice: a provider's provisioner
requires a seal key it has no way to receive, and the consumer no way to obtain
the matching one. The runtime cannot come up as delivered.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 22:23:24 +02:00
jschoubben f63eca13b3 ADR 0052 — a module runs its code as its own process, with its own account
The runtime-model gap the review found. A module with tools or events runs one
container — the tool runtime carrying its code — holding the one scoped account
ADR 0048 gave it. A node-wide runtime can't: it would hold the union of every
module's permissions, the isolation 0048 draws. So per-module: one module, one
process, one account. Tools served per key (serve.<tool>) so a caller names a
tool and only its module answers (superseding a shared tools.invoke); events in
the same process under the same account; the runtime image is the tool runtime
plus the module's code (the audit-logger's shape, made the rule). A plain
service module runs no such process. A provider's provisioner is a runtime too —
which is why a provisioner that emits must carry a broker credential or not emit.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 21:24:46 +02:00
jschoubben e898a3ec44 ADR 0051 — a module's configuration is its assignment's, not its manifest
A module is assigned to a node (there is no mesh assignment; 'mesh' is a scope).
The manifest is what the module IS, plus defaults; the configurable values are
settings, carried by the assignment — per-node or mesh-wide, applied at
resolution, changeable live (what a meshboard edits). Extends settings from a
config file's content to the manifest fields marked settable: foremost
listens.from (postgres from:mesh by default, from:anywhere per node — the
firewall follows), and a provider's own config (a registrar's zone/domain/
ingress). Static config in a manifest is config in the wrong place: it cannot
vary per node and cannot change without a rebuild.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 21:05:55 +02:00
jschoubben 9eb5576682 04-ISSUES/003 — fixed: the firewall scope is enforced now
The manifest refuses unknown keys (DisallowUnknownFields), 'from' is the field
that scopes a port and it is rendered to nftables (AsNftables), and the firewall
module applies the rule set. The chain from a declared scope to a dropped packet
is closed. Amended-design: ADR 0050.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 20:56:44 +02:00
jschoubben 5118258ee3 ADR 0049 and 0050 — public DNS and the firewall, two reachability decisions
0049: a public name is provisioned like any capability — a module requires
public-dns and contributes its host; a neutral interface answered by
registrar-scoped providers (cloudflare-dns, route53-dns) that create/remove
the record pointing the name at the mesh's public ingress. Pairs with route
(the proxy) and a public cert (the proxy's ACME).

0050: answers the firewall question. The firewall is NOT a provider like the
proxy — it is a machine's own filter, derived by the host as the sum of what
its modules declare they listen on, with 'from' the whole of public-vs-internal.
Enforced both ways, unknown keys refused — closing 04-ISSUES/003. A public
service is exposed through the proxy (listens from:mesh + requires route), not
by opening its own port.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 20:45:58 +02:00
jschoubben cfa3ad1930 ADR 0048 — ratified: status accepted 2026-09-04 01:22:47 +02:00
jschoubben b622de4fe6 ADR 0048 — a module's broker account is scoped by its emits and consumes
Events (0046) and their wire (0047) left open how a module reaches the
broker. The code has no generic module broker-account: only node and
builder scopes exist, so emits/consumes are enforced by nothing — a
manifest declaring a scope the broker does not draw (04-ISSUES/003).

Decides: on assign, a module gets a broker account whose permissions ARE
the manifest — read on mesh.events + its own queue bound to consumes;
write to mesh.events under module.<self>.* only; nothing else. Consuming
'#' is a deliberate, auditable grant. The account is what makes the
declaration a rule the broker enforces, not a comment.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 01:19:15 +02:00
jschoubben 820ce8fc8c ADR 0047 — the shape of an event on the wire
The wire contract ADR 0046 left open: two topic exchanges (mesh.events,
mesh.rpc, kept apart so # is a clean audit); the routing key as the event
type namespaced by origin (module.*, mesh.*, node.*); metadata in AMQP
headers (required x-event-id/x-source/x-node/x-time/content-type; optional
x-causation-id/x-schema; unknown x- headers ignored) with the body only the
payload; persistent messages; per-consumer durable dead-lettered queues
with prefetch; at-least-once with idempotent consumers (no false exactly-
once). The precedent is ADR 0043 for declarations.

Supersedes the sdk's first cut (metadata in body -> headers); that and the
queue config are code to align in mesh-sdk and mesh-tools.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-03 23:47:58 +02:00
jschoubben 4bb6dd26ed ADR 0046 — events are a relationship, provisioning's lighter sibling
A module emits and consumes events, both declared (emits/consumes),
parallel to provides/requires. Events are 1:many, broadcast, credential-
free — no provisioner, just the broker's topic routing — so most inter-
module reaction should be an event, not a provision. Every event carries
source/node/time so it is auditable; the audit logger is just a module
consuming '#', no privilege. A consumes for an event nothing emits is a
dangling edge and refused, like requires. One per-node runtime serves
tools, provisioning and events alike.

Extends ADR 0045; builds on ADR 0001 (the broker) and 0044 (emit/on are
stable sdk surface; the binding and runtime are not).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-03 23:42:25 +02:00
jschoubben b2cb481681 ADR 0045 — what a module is
A module is one self-contained piece of software the mesh installs and
manages; the software is its identity, and capabilities/seats/provisions
are the relationships between modules, not what a module is. Records the
three relationships (shared seat, exclusive seat, provide/require), that
interfaces are mesh-owned and providers adapt to them, and the naming
rule: draw the interface at the consumer's real coupling — neutral where
the coupling is thin (analytics), protocol-scoped where the consumer
speaks a protocol (postgres/mssql/mongodb), never false genericity.

Supersedes 0017 (domain grouping — wrong axis), refines 0002, generalises
0027's protocol-not-product rule.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-03 22:50:52 +02:00
jschoubben 50d61398b6 ADR 0044 — what the SDK holds, and what it refuses
Supersedes ADR 0030's 'types, not behaviour' line for mesh-sdk. The
boundary is change-frequency, not kind: the SDK holds the stable spine
(tool-serving harness, messaging/event framework, contracts, core
primitives) and refuses per-module clients, per-module tool code, and
anything volatile — because those are what turned hal/sdk into constant
maintenance and made every edit rebuild every module.

States the rule (frequent AND cascading is the disease), why the root
cause was intra-module feature-sharing leaking into inter-module
coupling, and where per-module shared code lives instead (in the
module — a shared file, or a module-local sdk for the few large ones).
Updates repos.md's canonical mesh-sdk description to match; leaves 0030
untouched (immutable).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-03 21:41:08 +02:00
jschoubben 7b4a664c50 Issues 008 and 009 — what the manifest review could not fix
Both from the 2026-09-02 review of the catalogue examples, and both
design gaps rather than defects in a file: a seed file the host
reconciles back to empty over the grants that grew in it, and a module
stack refused co-assignment because six manifests each own the
directories they exist to share. Fixing either in place would have
been picking an answer the records do not yet hold.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-02 01:07:05 +02:00
jschoubben 9638aa02c5 031 — a machine becomes each thing it was told, in turn
Declarations queue and the host applies all of them, oldest first, so a
machine pushed five things in a minute spends five applies becoming the
last one. Correct every step — each declaration is the whole machine —
and wasted in all but the final step.

Only visible since a report names its declaration: the reports arriving
were about ever-older ones, while timestamp comparisons used to happen
to pass. The fix is consumption order, not the queue; the open question
is what a superseded declaration's report should say, because silence
reads as disobedience and "applied" would be a lie.
2026-09-02 01:06:12 +02:00
jschoubben b245f5e254 011: measure the graph against every facet a module carries
The five declarations cover relations; a module is more than its
relations. Add the facet-by-facet coverage table so the effort cannot
conclude while tools, verification and contributions are unplaced —
and weigh each candidate gap rather than adopting it: contributions
probably dissolve into declared resources, mandatory verifiers risk
trivial ones, and the tool surface is the one facet with no home.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-01 23:14:03 +02:00
jschoubben 6bdd3f2bfa 030 — asking what a machine should be re-signed its certificate
Composing a declaration signed the machine's certificate anew each
time, and a signing carries a fresh random serial — so what the mesh
would send differed from what it had sent by one byte, for ever, and
every machine carrying a certificate stood eternally waiting.

Third find of the same rule: issued once and kept. The port had it, the
secret had it, the certificate composed fresh on every asking — and the
keeping column had existed since the serving-key migration, written by
nothing, the same shape ReleasePorts was found in.

Found by keeping the scenario standing and diffing two plans seconds
apart: one line, where four theories had none.
2026-09-01 22:28:26 +02:00
jschoubben 6385d52bc0 029 fixed — the store is added, not built, and the cycle is refused
The registry module names its image by digest, the way the bundle names
the three a first node starts from, and goes in as a manifest. The lab
now walks that path and passes.

A manifest that provides the store and also builds artifacts is refused
at parse. The lab had been passing only because its Docker Hub stand-in
quietly received the push — a prop covering for the thing under test.
2026-09-01 21:43:42 +02:00
jschoubben 013dbce4a9 029 — the artifact store cannot be delivered by the artifact store
Installing the module that provides `artifact-store` requires something
that provides `artifact-store`: its image is mirrored in, mirroring
publishes to the store, and the builder refuses to run without one.

Never seen, because the lab always has a registry standing before the
mesh asks for one, and so does any mesh built on a machine that already
had one.

What it blocks is larger than a registry. The store holds images and
packed archives both — it is the module catalogue in artefact form — so
until it exists a mesh can run only what its bundle already carries.

The substrate record already answers it. It asks of each candidate
whether it can grant itself the thing it provides: the store cannot
create its own database, the broker cannot create its own virtual host,
and the registry cannot grant itself a repository. So the registry
module names its image and is never built.

With a limit worth saying out loud rather than discovering: a module
providing the store may not build artifacts of its own, a UI or a tool
server included, because there is nowhere to put them until it runs.
Such a registry is two modules.
2026-09-01 21:15:05 +02:00
jschoubben e75683c01f 026 reopened — the rule is right about data, wrong about facilities
Enforcing "declare what you mount" refuses the builder, which mounts the
container runtime's socket. That socket is not the builder's data: it
exists already, the machine owns it, and declaring it as one of the
module's directories would be a lie the host would act on.

Two kinds of mount are spelled identically today — the directory my data
lives in, and a machine facility I was granted. Until a manifest can say
which, the fourteen declared mounts are right by coincidence, which is
what this issue was opened about.

Recorded rather than decided: separating them is new vocabulary, and
inventing it to turn a check green is how a mechanism nobody chose ends
up load-bearing.
2026-09-01 19:45:55 +02:00
jschoubben bd8f09d647 026 fixed — and the coincidence turned into a rule
The fourteen undeclared mounts are declared. More to the point, a
manifest that does not declare one is now refused: they were right by
coincidence, and a checklist nothing enforces is a checklist that is
true until the next commit.

Refused in the control plane, because the machine cannot tell the
difference — asked to mount a path that does not exist, it makes the
directory, which is a thing it is perfectly able to do.
2026-09-01 19:36:15 +02:00
jschoubben 0bcfeb4e80 028 fixed — the mesh assigns the port, and knows what it cannot move
A module now says its port once, in `listens`, and the container's
mapping, the rule set and what a consumer is told are all derived from
one assignment. The three hand-written copies that agreed only because
one person wrote them are gone.

The half that made this an issue rather than an inconvenience was that
the substrate is not a module: nothing in the mesh had heard of its own
store, so it handed a database module the port the store already had.
The machine now says what it carries, and the mesh assigns around it.

Ports the protocol fixes became claims, which needed no new mechanism —
the mesh already had one for what is singular on a machine.

Left open, and unchanged by any of this: whether a module should publish
to the machine at all. Assignment makes publishing safe without making
it necessary.
2026-09-01 18:32:58 +02:00
jschoubben c2a37ab7f8 38 — the mesh assigns the port, and a module does not care
Jochen's call, and the right one: a module cannot choose a port well,
because it is written once and assigned anywhere. Any number it picks is
a guess about a machine it has never seen, and two modules guessing the
same number is not a mistake either of them made.

Writing it up turned up something the issue had missed. The same number
appears three times in every module — the rule set, what a consumer is
told, and what the runtime publishes — and nothing checks that they
agree. They agree today because one person wrote all three. A module
whose `serves` said one thing and whose container published another
would resolve, compose, apply, and hand every consumer a port that
answers nothing.

So the decision is one source with the other two derived, and an
assignment made once and kept, as a credential is.

The part that needed thought is ports that cannot move — mail on 25,
submission on 587. Those become claims, which is what the mesh already
has for what is singular on a machine. Two modules wanting 25 is the
same shape as two wanting the seat, and gets refused by name at
assignment rather than by a container runtime at apply. That makes this
mostly a matter of pointing an existing mechanism at ports.

Left open: whether a module should publish to the machine at all.
Assignment makes publishing safe without making it necessary.
2026-09-01 17:42:27 +02:00
jschoubben 6330abce5f 028 — two things want one port, and nothing says so until the machine
Found by fixing 027 and pushing again. The declaration is now accepted
and the database container still cannot start: the mesh's own store
holds 5432 on that machine, and the module publishes 5432.

Nothing catches it because the substrate is not a module. It arrives
from the bundle before there is a mesh to ask, so the control plane has
never heard of the store and does not know it holds a port. Resolution
can compare modules with each other and cannot compare one against what
the mesh is built on.

Nor does it compare modules with each other. A port is exclusive on a
machine in exactly the way a claim is, and the mesh has a mechanism for
that which ports do not use.

It has been met before: the end-to-end test that exercises a real
database publishes 5433 rather than 5432, inline, with nothing saying
why. That is how a constraint becomes folklore.

The open question is bigger than the bug. Whether a module should
publish to the machine at all decides how a consumer reaches it, and
changes what `serves` means.
2026-09-01 17:21:31 +02:00
jschoubben 03266fd4a2 027 — a container cannot follow a file, and a rotated credential is the case
Found by the forge failing to start. I had put `restart-on` on nine
containers so they would pick up a rotated credential; it belongs to a
service, and the host refused the whole declaration.

Removing it fixes the modules and leaves the reason I reached for it.
The mechanism is written against exactly this, in the host's own words:
a running service does not re-read its configuration, so replace the
file, find it already running, do nothing, and the machine keeps
behaving as before while every check passes. Every word of that applies
to a container, and nearly everything the mesh runs is one.

The cost is concrete. Rotation replaces the file and tells the provider
to accept the new credential. A provider reconciles, so it takes it. A
consumer is usually a container, so it does not — and the two ends hold
different passwords, which is the fault ADR 0001 records costing two
days. The test that proves rotation works uses a consumer that reads the
file on each attempt, so it does not meet this.

Two things a fix has to keep: it stays declared state rather than a
command, because the link may not carry an action; and where an env-file
changed, the honest verb is recreate rather than restart, because a
container's environment is fixed at creation.
2026-09-01 17:19:43 +02:00
jschoubben 3f00d3b413 026 filed; 025 corrected; the image store is a module
Three corrections, two of them to things I wrote today.

026 is the serious one. Four modules mounted fourteen host paths nothing
declared — the mail spool, the databases, the object store's data. The
runtime creates those as root, so owner and mode go unapplied, and the
rule that keeps a directory holding data the mesh did not put there is
written in terms of declared directories. It reached the configuration
and missed the data. The cause was carrying compose files across: a
container shape that can express one gets filled in like one.

025 claimed nothing turns a tag into a digest. That is false, and the
answer was designed and built before I wrote it. A module names an
artifact, not an image, and `kind: upstream` mirrors somebody else's
image into the mesh's own registry, pinned by the digest it lands with.
The two-document split the issue described as the shape of a fix is the
design. Pinning twelve images by hand was treating the symptom, and left
them pointing at a public registry rather than the mesh's.

And the image store was written up as something the mesh does. It is an
ordinary module — considered for the substrate and removed, because the
test is whether the control plane needs it before its first instruction,
not whether it can grant itself one. So somebody's own registry is the
same module as the mesh's.
2026-09-01 16:10:13 +02:00
jschoubben befbb1d978 The five modules could not have run, and now the forge does
Recorded against phase 3, because the phase note said running them
needed images stocked and provisioners built — and missed that not one
of them named an image that exists. Sixty-four zeros where a digest
belongs, eighteen times, parsing and resolving perfectly.

The forge now runs: on a database another module provides, with a
password it did not choose and a connection string it could not have
written itself. First of these descriptions to be started rather than
planned, and it exercises the whole of the credential work.

What remains is the mechanism rather than the data. Nothing turns a tag
into a digest as part of the mesh's own work, so it was done by hand —
which is what the issue says a person should not be asked to do. Asking
a registry takes a second and pulls nothing, so the main argument for
leaving it undone is gone.
2026-09-01 15:40:52 +02:00
jschoubben 2bbdc52440 025 — half done: nothing unrunnable reaches a machine now
The refusal landed, and the examples pin images that exist. What is
still missing is the part that makes it unnecessary: nothing in the mesh
turns a tag into a digest, so it was done by hand — which is precisely
what the issue says a person should not be asked to do.

Recorded because the resolution mechanism turned out to be trivial:
asking a registry what a tag points at takes about a second and pulls
nothing. That removes the main argument for leaving this open.

Also records the two faults that fell out of pinning for real. The mail
system named seven repositories that do not exist, because it publishes
to a different registry than the manifest assumed, and one of the seven
had been renamed upstream. Nothing checking only the shape of a
reference could have found either.
2026-09-01 15:13:52 +02:00
jschoubben fc7717f6b1 021 closed — the record was left open after the fix landed
The code and its tests went in hours before the record was touched, so
an issue that read `located` had been fixed all along. That is the exact
failure the frontmatter exists to prevent: status is meant to be
answerable from the record rather than by reading the code.

Closed with the commit that did it, and cross-referenced to 022 and 023,
which came out of the same mistaken instinct — treating the machine as
a boundary, then as an identity, then finding a consumer had a password
and no name to present with it.
2026-09-01 15:04:36 +02:00
jschoubben 0871e6ec11 37 — where a module lives, proposed
The module descriptions sit in `examples/` inside the control plane, and
that name has been doing harm: everything there reads as a sketch, and
one shipped naming a container image nothing builds. A directory called
the catalogue would have made "does this work" the obvious question.

The shape of the answer turns on one measurement. Of the 126 modules in
the system being replaced, 47 are software in their own right — the
largest is 182 source files, and a speech-capture module carries a whole
daemon. Another 44 ship helper scripts. Only 35 are a description and
nothing else.

So a catalogue cannot be a folder of manifests, because two thirds of
modules are programs. That splits them four ways, and only two of the
four belong in a catalogue: things the world made that we describe, and
packages with some files. What the mesh is made of stays in the
repositories that build it. What we wrote keeps its description beside
its code, in the same commit, because nothing else can stop the two
drifting.

The mesh's list of modules is a table, not a repository, and it already
records where each module came from and at which commit. Nothing needs
inventing for modules from anywhere; a repository of ours is just the
source we curate.

The check that a description is valid should move to a command on the
control plane's binary. Today a test reaches into the control plane's
internals to parse manifests, and another reads its build file to check
images exist — two jobs tangled. A command would also give the same
check to somebody describing their own application, which is the case
that matters most and has none.

Left open: how a provisioner's image gets published and pinned, and
whether thirty-five install-a-package modules deserve to be modules at
all.
2026-09-01 14:24:53 +02:00
jschoubben f6b1834ea9 024 fixed — the registry was addressed by hand and nothing else was
The cause was one line. Machines get a systemd-networkd unit with a
static address, so networkd finishes and reports the link configured.
The registry ran `ip addr add` inline, which leaves networkd waiting to
configure something it was never told about — and
systemd-networkd-wait-online has an infinite timeout.

So network-online.target was never reached and everything ordered after
it never started. On these machines that is Docker, so `docker load`
blocked on a socket whose daemon was queued behind a target that would
never come, and three bounded timeouts stacked to thirty-five minutes.
These machines have no DHCP by design, so that wait was never going to
end.

The hypothesis in this record was wrong and the record now says so.
Stocking had just been changed, so stocking looked guilty; stocking
takes 34 seconds and always did, timed directly before changing
anything.

Fixed with two things that made it cost hours instead of minutes: an
image placement now waits for the runtime and refuses after 120s naming
what systemd is waiting on, and the end-to-end test passes onProgress —
the raise reported every step and the test discarded it, which is why
thirty-five minutes and four minutes of silence looked the same.

The suite then ran to completion, 23 of 24, the one failure a check of
its own flagging a path as a credential because `/` is in the base64
alphabet.

Also recorded: a redirected log lags, because Node block-buffers stdout
to a file. Read as a stall twice, the second time right after the real
fix — where a buffering artifact argues the fix did not work.
2026-09-01 09:53:08 +02:00
jschoubben 407416e6d0 024 — a run stalls before the host is placed, and says nothing while it does
Seen twice today. Once mid-run: thirteen passes, then the process ended
with no summary, no failure and no receipt. Once from the start: the
first test ran 35 minutes against a measured 4.5 and was still running
when it was stopped.

Ruled out rather than assumed: not memory (84 GiB free, no OOM), not the
daemon (the stalled machine answered `incus exec` immediately), and not
the changes under test — the anchor VM had no host log and no
containers, so the run never reached placing the host.

What changed just before is that the rebuild went from two artifacts to
six, and every one of them is pushed into the scenario's registry, which
is the step the second stall sat in. Recorded as what changed, not as
the diagnosis.

The reason this is an issue and not a slow test: the suite prints
nothing between starting a scenario and finishing its first test, so
four minutes and thirty-five look identical from outside, and the only
recourse is to guess. That is how a workstation was left unbootable in
August. And a run that ends silently after thirteen passes is a run
somebody may believe.
2026-09-01 03:20:40 +02:00
jschoubben 39916b26e9 What stands between 3.1 and a running identity provider is a program
023 is fixed, so the design faults are gone and one concrete thing is
left: the realm provisioner does not exist. Its manifest named an image
nothing builds and no program backs, which has been removed — a manifest
describing a program nobody wrote is the same mistake as the credential
files that could never be read.

Keycloak's manifest now says what is true today, and the gap is loud: it
no longer claims to provide oidc-client, so a consumer asking for one is
refused by name at plan time instead of resolving cleanly and waiting
for a client nothing will create.

The provisioner should be written against a real Keycloak in the lab
rather than from the API documentation. The object store's took three
corrections that only a running server produced.
2026-09-01 03:13:11 +02:00
jschoubben 0760280bc6 The largest gap in the coverage list is not a gap
Tool servers — 56 modules, over half — were written up as the biggest
missing thing. They are expressible with what exists, and the first
framing was wrong in a way worth keeping: a module provides `tools` and
the session requires them does not work, because a requirement has one
answer and 56 modules offering tools would be 56 answers.

Turned around it fits exactly. The session provides `tool-host`; every
module offering tools requires it and contributes where its tools are.
Many-to-one is what `contributes` has always been, and the session
receives all of them in one file. Verified by resolving it rather than
by reading the code.

It only became possible today: until 022, several modules on one node
requiring the same thing was refused outright. Worth noting because it
means the credential fix bought more than credentials.

What remains is a decision about what a tool server is, which is work
rather than a missing shape.

The entry stays in the list rather than being deleted — a checklist that
quietly loses its biggest item reads as though nobody looked.
2026-09-01 03:09:11 +02:00
jschoubben f583502fc9 023 fixed — the mesh says who a consumer is, and what it is bound to
Both halves had one cause: the mesh knew something and did not say it.

Who a consumer is now comes from one derivation, sent to the provider in
its grant and to the consumer in its binding, so the two agree by
construction. The provisioners use the name they are given and refuse to
invent one, because a name of their own would create a login the
consumer could never guess while everything reported success.

Bound values reach the file that needs them through the symmetric twin
of the sealed placeholder — simpler, because they are not secret, so the
control plane fills them in and the host gains nothing.

The lab run meant to prove this failed in a way that looked like the fix
being wrong: rotation could not authenticate against a real database.
The cause was the suite rebuilding the control plane's image and not the
provisioner's, so an image built that minute ran against a provisioner
built the day before. That is 005's family and is recorded with the
issue, because the misleading part is worth more than the fix.
2026-09-01 03:05:43 +02:00
jschoubben 4adc656b54 Where Phase 3 actually stands
All three modules have manifests, all three parse, resolve and plan, and
none of them can start. Worth writing down before it reads as progress
or as failure, because it is neither.

The vocabulary held. Nothing in 3.1–3.3 needed a new shape — including
the mail system's several containers on a private network, which was the
one expected to break it. That was the question this phase was designed
to answer.

What did not hold was underneath: 022, now fixed, and 023, open. Both
are about credentials rather than about what a module can say.

The third fault was in the manifests, not the design: a secret declared
at a path named .env and read as one, when a sealed file holds a
password and nothing else. That is what a manifest checked only by a
parser buys, and it is why there are now two tests reading the manifests
on disk.

023 is the whole of what remains before the identity provider runs.
2026-09-01 02:56:54 +02:00
jschoubben dea46c3c6c Correct what the coverage document said about actions
I wrote that a module cannot declare an action. It could — the parser
accepted one, and the refusal only came on the machine. The claim was
wrong in the direction that matters: it read as "the design prevents
this", when what prevented it was a check at the far end that nobody
would connect back to the manifest.

Health checks are still the gap most worth closing, but the shape of the
answer is different from what I wrote. An action is not available to a
module at all, so a health check needs a way to say ask this and expect
that without saying run this — closer to a listens entry than to an
action.

Also records the finding itself, because it is a recurring shape here
and not a one-off: a rule enforced only at the far end is enforced and
unusable.
2026-09-01 02:56:05 +02:00
jschoubben 9073d3f2df Say plainly that env-file never points at a sealed secret
The playbook offered `env-file` and `${secret:name}` as alternatives,
and that reading is what produced the bug every example module shipped
with: own-secrets pointing at a path named `.env`, mounted as env-file,
holding a bare password. The container starts with no password set —
which is a service running on the wrong credential, not a failure.

They are not alternatives. A sealed file holds a password and nothing
else, so env-file points at a file the module declares whose content
leaves a hole, and the host fills it on the machine. A provisioner is
the exception, because it reads a password file.

Written out as the three lines a module needs, with the failure it
prevents named, since the abstract version was already there and was
read the other way.
2026-09-01 02:52:55 +02:00
jschoubben 2baf22ac43 A pair is a module and a provider, not two machines
Amends the credentials page, which said "every pair has its own
credential" and meant two machines. Built that way, it was wrong in a
way that only shows on a real node: a machine running several services
against one database server had one credential between them, so the
provider refused to plan at all and the consuming node quietly gave the
first module a credential and the rest nothing.

The page already argues the case against itself — one credential with
many holders is the first of the three faults it was written to remove.
It just drew the boundary at the machine.

Two modules on one node are as separate as two on different nodes, and
one login opening both is what this page exists to prevent. It is also
what makes withdrawal possible: one role per machine cannot say that
this module has lost its login and the others still have theirs.
2026-09-01 02:46:37 +02:00
jschoubben 80f18caf03 022 fixed; 023 filed — a password is not a connection
022 turned out to have a silent half worth recording: the provider
refuses loudly and names the modules, which reads as a decision, while
the consuming node does not refuse at all. Three modules wanting one
database produce one need, so two of them get no credential file and
each starts and fails to authenticate with nothing saying why.

023 is what remained after fixing it. A consumer now gets its own
password, in whatever shape its configuration wants, and still cannot
connect: the user name is invented by the provisioner and recorded
nowhere in the mesh, and the host and port sit in a JSON binding that an
application reading KEY=value cannot use.

The asymmetry is backwards and the coverage document now says so. The
secret is the hard case, because the mesh must not be able to read it,
and the secret is the part that arrives. The host and port are ordinary
facts the mesh holds in the clear, and they are the ones stuck.

Keycloak, Gitea, Mailu and MinIO all parse and resolve and none of them
can start. This is what stands between the module set and a running one.
2026-09-01 02:40:40 +02:00
jschoubben c86adbe3cc 022 — a credential belongs to a node, so a second consumer refuses
Found while checking whether the module vocabulary covers real use
cases. A node running three modules that all want a database cannot be
planned at all:

  anchor has 3 modules asking for "postgres-database" and they would
  share one credential: gitea, keycloak, umami

The refusal is right about what it says and wrong about what it implies.
They would share one credential, and sharing is worse than refusing —
but the arrangement being refused is the ordinary one, and the node this
mesh exists to take over runs eight modules against one database server.

The cause is the key: a credential is keyed by provision, consumer node
and provider node, so `consumer` is a machine. The provisioner inherits
it and names the role `mesh_<node>`. The refusal is not a check that
caught something; it is the only honest thing that function can do with
a key that cannot tell two consumers apart.

It is the same mistake as 021 with a different face. There the machine
was treated as a trust boundary; here it is treated as an identity, as
though "who is asking" is answered by naming a host. Two modules on one
node are as separate as two on different nodes.

Worth stating plainly: without the refusal, gitea's login would have
opened keycloak's database, and nothing would have said so — from the
provisioner's side it created exactly what it was asked to create.

Not a local fix. It crosses the control plane, the grant file naming and
every provisioner that names something after a consumer.
2026-09-01 02:28:39 +02:00
jschoubben 36d342f176 What a module must be able to say, measured against 127 that exist
Every manifest in the system being replaced was read and every key
counted, then set against what the new one can express. Three findings
worth more than the table.

**The most-used key was already covered and I expected a gap.**
Depending on another module — 65 manifests, the commonest thing any of
them says — is a requirement naming a module, which already means that
module rather than anything providing the name.

**The largest real gap is tool servers: 56 modules, over half.** A
module can already run one; what is missing is anything saying it offers
tools. That is plausibly a provision rather than new vocabulary, which
would need nothing added — not yet decided, and recorded as undecided.

**The gap most worth closing is health, at seven modules.** The mesh
knows a container is running, which is not whether it answers, and this
project has paid for that distinction twice. An action with a verify is
exactly the right shape and may not arrive over the link, so a module
cannot declare one.

Two things are missing deliberately and say so: stage hooks, because the
link may not carry an action and a module needing setup ships a program;
and flavours, retired in favour of claims.

Config merging is missing and should stay missing. A mechanism that
understands TOML gets asked for YAML, then INI, which is how the thing
being replaced became unholdable.

Also records what the survey found that is not about coverage: manifests
that had stopped matching what was actually brokered, one fact derived
in two places giving two answers, and a live listing returning
credentials in plaintext.
2026-09-01 02:20:34 +02:00
jschoubben 1ee62392b9 Playbook 06 — writing a module, from doing it once
Written after porting the first real workload end to end. Every step
exists because skipping it cost something, and the ratio is recorded
because it is the lesson: six attempts, one real bug, and the mesh was
right every time.

The rule worth carrying out of it: read the host's log before
theorising. A declaration that was sent and not applied says so there
and nowhere else — it took an hour to look, and the answer was one line.
2026-09-01 01:46:46 +02:00
jschoubben ecfc2c215e Point the plan at the survey, and name the likelier failure
The conversion's detail is operational and names machines, so it lives
in the mesh's knowledge base rather than in this repository:
`migration/where-service-data-lives` for where every service's data
actually sits, and `troubleshooting/db-password-frozen-at-first-init`
for the lockout. This document says the rule; those say the specifics.

The lockout is the finding worth carrying here, because it is worse than
the one this plan was already guarding against and it is likelier. A
database image consumes its password variable only when its data
directory is empty. Everything keeps data on a persistent directory, so
the role holds whatever password it was created with for ever;
regenerate the variable and the application moves on while the database
does not, permanently, because nothing reconciles it.

Eight modules are in that state today and work only because nobody has
regenerated their credential since their data directory was created.
It was already documented in the knowledge base and my survey had missed
it — found by searching, which is the argument for the knowledge base
existing.
2026-08-31 21:46:53 +02:00
jschoubben 122405df8e 020: the server version is not it either
Pinned 2.5.0 rather than latest, on the suspicion that its draft
profiles extension was involved. Identical failure, so that is ruled out
and recorded — two of the three guesses in this issue have now been
tested and both were wrong, which is the useful half.

The scenario keeps the pin regardless; it should have had one from the
start.
2026-08-31 21:42:13 +02:00
jschoubben acd5a0d80b File 020 — a certificate is issued and never collected; close Phase 1
Against a real ACME server the proxy orders, the challenge is answered
at the name on port 80 through the proxy itself, the authorisation goes
valid, finalisation is accepted, and the authority issues a certificate.
The client then posts to an empty URL to collect it, and never does.

Read from the authority's own log rather than inferred. Across one run
it issued two certificates and accepted finalise three times: the client
reaches issuance every attempt and fails at the same step after it.

Ruled out and recorded, so nobody repeats it: the directory is complete;
the authority's API certificate covers the address; the challenge path
works. A hand-written server config was suspected and was wrong —
replacing it with the server's own default, changing only the challenge
port, gives the identical error.

Filed rather than pursued because what remains is interop between two
libraries against a server that exists to be a test server, and may say
nothing about a real authority. What the mesh needed to show, it showed:
a routed name gets a certificate ordered from a configured authority,
and an unrouted one gets nothing — that second assertion passes.

Phase 1 closes with this one item partly open. Two of its four tasks
needed no code at all, the network shape was built, and the next thing
to learn comes from moving a module rather than a fourth lab run.
2026-08-31 21:40:03 +02:00
jschoubben f4347e2f14 Correct an overstatement: a sealed secret is readable on its node
An earlier paragraph implied a secret becomes unrecoverable once
accepted. It does not. It is sealed to the node, which holds the private
half and writes the plaintext into the module's own file at 0600 — the
value is there, on the machine, as an ordinary file.

What does not exist is a way to ask the mesh what a secret is. That is
the property worth having and it is narrower than what was written.

The reason to capture the old system's environment first is simply that
adoption means supplying those values, not that they become
unrecoverable.
2026-08-31 21:34:50 +02:00
jschoubben b68103d198 A module is adopted with the credentials it already has
Nothing is rotated during the conversion. A service keeps the password
it is already using, because minting a new one is how a running service
stops being able to reach its own database mid-migration.

The mesh has both paths already: generate-and-seal for a new module,
accept-and-seal for an adopted one. Adoption needs the second, and it is
built.

Rotation becomes a separate act afterwards, once everything works — the
machinery is proven, and it is a thing to do deliberately rather than as
a side effect of moving a service between systems.

Records the step that has to come first and is easy to miss: read the
current environment out of the old system while it can still be read.
Once accepted, the mesh cannot show a secret back, and once the old
system is gone neither can that. A password nobody wrote down is a
service nobody can adopt.
2026-08-31 21:31:26 +02:00
jschoubben f801b4b3e2 The board is published publicly, which decides everything else about it
Assumed throughout and stated nowhere — the wrong way round for the most
consequential fact about this component.

A board reachable only over the private network would sit inside the
boundary 0004 already calls the security boundary, and a login there
would guard a room whose door is inside the building. This one faces the
internet, so its login is a perimeter rather than defence in depth.

Which makes the identity provider the mesh's outermost gate. The board
presents the control plane, and the control plane's networked surfaces
can change the mesh (0035) — so whoever that provider admits can assign
modules, from anywhere. Written flatly because it is easy to arrive at
one reasonable step at a time and then be surprised by.

What follows is not the board's own design: who may log in is a decision
about the mesh rather than about an application; a public name needs a
certificate from an authority the world trusts, which is why that work
exists; and the provider going wrong in the permissive direction is a
mesh-wide exposure with no local symptom.

The command line is unaffected and is why this is tolerable — it
authenticates through nothing and answers to the machine's own login, so
the mesh stays operable by somebody standing at it whatever happens to
the gate. That is the property to protect if the rest is ever traded
away.
2026-08-31 21:26:06 +02:00
jschoubben a1a10e9ed2 Bootstrap ends at a usable mesh, and the first credential comes from a person
Bootstrap stopped when the control plane started — a mesh that runs and
cannot be used by anybody not standing at the machine, since the
networked surfaces need an identity provider and no module has been
assigned yet. It now runs through the provider and the first login.

The obstacle was not incidental. The mesh has never held a readable
secret: Make generates and seals, keeping no readable copy. An initial
administrator's credential is the first value a person must read.

Generating it and printing it once was the convenient option and is
refused. It would give the control plane a plaintext secret for the
first time — briefly, and to one terminal, but the capability would then
exist, and an exception made for one case does not stay one. The next
awkward credential gets printed too, and "a copy of the database is a
copy of nothing" stops being checkable by reading the code.

So the operator supplies it, on standard input, not echoed — the path
that already exists for a model-access key. What is created is an
account in the identity provider, not a user of the mesh; there is still
no user model.

Unattended bootstrap remains possible and the value still comes from
outside: automation supplying it is the operator supplying it. What is
refused is the mesh inventing one, so an unattended bootstrap with
nothing provided yields a mesh with no administrator — correct rather
than broken.
2026-08-31 21:23:55 +02:00
jschoubben 0ef6d3f574 One implementation, several surfaces, and what that costs
The mesh is operated from a command line and must be operable from a
browser and from a model's tools, without becoming three systems. The
pattern is already in the code and was unnamed: `board` serves HTTP by
calling the same functions the CLI calls, holding nothing.

Takes the decision 0034 said had to be taken deliberately rather than
arrive with a feature: the HTTP surface is not read-only, so a browser
login now carries authority over the mesh.

Names the dependency by protocol — an OAuth2 identity provider — as the
mesh does for AMQP, S3 and OCI. Keycloak is what fills the role; what
the control plane knows is that it validates a token, and replacing the
provider is a migration rather than a redesign.

Says what this must not become, because it is the failure the project
was started over: a kernel every module imports, 155 files of code from
every context. Shared surfaces are not a shared library. Three adapters
calling the same functions is not the same as logic leaving the context
that owns it.

And records the loop it creates. The networked surfaces depend on a
module the control plane assigns, so when identity is down nobody can
authenticate — including whoever is trying to fix it. The way out is the
command line, which authenticates through nothing and is available to
the account that owns the machine. Hence the rule: no capability exists
only behind an authenticated surface, because that is a capability which
disappears exactly when identity does.
2026-08-31 21:17:17 +02:00
jschoubben fccac61e58 The board is a web application, not a category
Supersedes 0032, which decided the right thing and described it wrongly.
The decision is unchanged: the account that installed the host owns the
mesh, and there is no user model.

What was wrong was inventing "a surface that delegates authentication"
for the board. It is a web application with a login, in the way every
web application has a login. That is a fact about an application, not a
property of the mesh.

The cost was not cosmetic. It made the identity module look like part of
the mesh's authority — something the mesh depends on to know who anybody
is — when the mesh knows nothing about people at all and one of the
applications running on it happens to have a login.

Keeps the line that is worth writing down, and states it more plainly:
signing in to an application must not become authority over the mesh.
Today it cannot, because the board reads and does not act. The moment it
can assign a module, whoever it lets in has mesh authority — and it
would arrive as a feature rather than as a decision. So a surface that
can change the mesh is a change to who owns the mesh, and is taken as
one. Not forbidden; just not something that turns up in a pull request
titled "add assign button".
2026-08-31 21:08:39 +02:00
jschoubben 10fa7c76d7 The substrate is a store and a broker
Third correction to one table today, found the same way as the other
two: by asking whether both halves of the test were answered, or only
the easy one.

0006 admits the registry because "it cannot grant itself a repository" —
true, and the second half. Nothing established that the control plane
needs one in order to run. Counted rather than argued: the bundle raises
twelve resources and no registry is among them. The registry arrives
afterwards as an ordinary module, which is exactly what the lab asserts.

0006 half-said this already, calling it "substrate by role and ordinary
by delivery, provisioned once there is a control plane to do it". A
member provisioned by the thing it supposedly precedes is not a member;
that phrase was carrying a contradiction rather than resolving one.

The registry is a closer call than the object store and the difference
is worth keeping: the control plane never touches an object store at
all, but it genuinely uses the registry. So the registry is a real
dependency of the mesh operating and not of the control plane starting —
and it is the second that the word means.

The substrate is now exactly what the bundle raises, which is the
strongest form the list can take: checkable by counting rather than by
reading an argument, and the two cannot drift.

The finding is not about substrates. A test with two conditions is a
test only when both are asked.
2026-08-31 20:30:19 +02:00
jschoubben a028337490 The local account owns the mesh; a surface delegates to a module
Answers what 0031 left open, and a question it did not ask — who owns
the mesh at all. There was no answer, and the absence was invisible
because every operation so far has been run by the person sitting at the
machine, so nothing had to say whether that was the design or the
circumstance.

The account that installed the host owns the mesh on that node. No user
model, no roles, nothing to administer. It follows from 0004 rather than
adding to it: there is no authorisation between nodes because every node
is the operator's own, so a user model inside that boundary would guard
nothing — anyone it could stop could read the node's key off the disk.

The board is different, and the difference is the network. A surface
reachable by a browser has to know who is asking, because those people
are not by construction people with a shell on the machine. So it
delegates to an OAuth provider, which is a module.

That does not make identity substrate. A surface delegating
authentication is not the control plane delegating it: the control plane
runs, applies declarations and reaches nodes with no identity provider
in existence. Only the board needs one.

Records the cost plainly: anybody with a shell on a node has full
authority there, and there is no way to give somebody authority over one
node without giving them a login on it.
2026-08-31 20:23:42 +02:00
jschoubben e3934e4449 The control plane authenticates nobody, so identity is a module
Closes the last open question about what the substrate contains. 0006
left an identity provider conditional — substrate only if the control
plane delegated authentication — and said the decision had not been
taken. It is now: it delegates to nothing.

The conditional was never about machines. A node proves itself with a
keypair it generated over a broker account issued at enrolment, and
declarations are verified by signature; none of that involves an
identity provider. It was only ever about whether a person signing in to
a mesh surface would be authenticated by something else.

So the substrate is three — a relational store, a message bus, an image
registry — and with 0028 having removed the object store, no member is
conditional and every one is there for the same reason.

It does not settle how a person signs in to a surface, deliberately.
What is settled is that whatever answers that is not something which
must exist before the mesh does, so it can be decided late or replaced —
which being substrate would have prevented.
2026-08-31 20:18:03 +02:00
jschoubben c570c687f6 The handover: switch off the old brain, leave the services running
The conversion method, recorded because it decides everything else and
was not written down.

The old control plane is stopped — provisioning, coordinator, syncs, the
pipeline, anything that decides or writes. The workloads it was managing
keep running, because nothing is managing them. The new mesh then takes
ownership one module at a time.

Nothing is ever unassigned in the old system. Unassigning is how it
removes things and removing is how data is lost; it is asked to stop
having opinions, never to take anything away.

Disabled rather than merely stopped, which is the part easy to get
wrong: those units are enabled, so a stop lasts until the next reboot. A
reboot mid-conversion would bring the old control plane back to
regenerate managed files underneath the new one — the one situation
where two systems really would fight over a machine.

A service left running with nothing managing it is the safe state: it
has its data, its configuration is on disk, and nothing will change
either. The risk in a conversion is in the managing, not the running.

Also records why taking ownership piecemeal is safe: the new host's
orphan removal is per-origin, so it only removes what it recorded
itself. Services it was never told about are not orphans to it.
2026-08-31 19:55:35 +02:00
jschoubben d1ab2dc0b4 Data outlives the mesh that declared it, and the conversion starts where it lives
0030, found by asking what the conversion actually needs rather than by
reviewing anything. The host deleted a directory and everything under it
when it stopped being declared — which happens when a module is
unassigned, or when a manifest is edited to move a data folder, which is
the exact operation this plan needs. A database's files, a mail spool.
The report said "removed".

A directory still holding something is now kept and said so. No flag and
nothing to remember: emptiness is the test, and it works because the
removal order was already right — the mesh's own contents are gone by
the time the directory is reached, so what remains is by definition
something nobody declared.

The plan now says data outranks its own ordering: copy, read back
through the service that owns it, and only then point anything at the
new location. Never move and then check.

And it records where this starts — the node holding all the production
data — with what that costs stated rather than argued with. Everything
proven so far was proven on machines that could be destroyed and raised
again. A scenario proves the mechanism, not the state on that machine.
2026-08-31 19:54:41 +02:00
jschoubben 9ad0ec35e0 The conversion is done by hand, and that removes work from this plan
Recorded because it is load-bearing and was not written down: moving
from the current system to this one is a person at a command line, not a
migration program.

What that removes is larger than what it adds. Nothing in this plan
needs an importer, a translation layer, a compatibility shim, or a way
of keeping two systems agreeing while both are live — each of which
somebody would otherwise reasonably build, use once, and maintain for a
year.

It also settles what "safe" means for the system being retired: a fix to
it must be safe on its own, because there is no careful rollout to
sequence it into. A change needing three steps in the right order is a
change that will be half-applied. That reversed a certificate default I
had chosen this morning.
2026-08-31 19:46:21 +02:00
jschoubben e4327a3a5e Phase 1.3 done: ordering was already there, the network was not
Ordering needed no change for the third time running — resources apply
in the order declared and nothing sorts them — and is now asserted,
because sorting them for any sensible reason would have passed every
other test.

Separates ordering from readiness, which the task had run together: a
container started is not a container ready. Nothing waits, and what
needs something usable retries. That is deliberate and more robust than
start ordering, since a dependency can restart long after apply.

The network was the first thing in Phase 1 that genuinely needed
building, and the first that needed a decision: 0029 records why a shape
rather than an action, and the vocabulary is nine.
2026-08-31 18:55:22 +02:00
jschoubben ce486fd5d2 Phase 1.2 done, and it is the same surprise as 1.1
A session as a licence consumer needed no change either: the two
sessions are two modules, so the existing (node, module) binding already
names them apart. 14-model-access.md's "a step toward it and not it" is
true of a worker and not of a session, and the difference is that there
is one session per node rather than many per machine.

Records what stays open: the worker half of that gap is real and
unaffected, and belongs with 0003, which is unbuilt.

Two tasks in a row that were already possible. Both were written from
the design rather than from the code — the review's own finding arriving
in the plan it produced. The remaining Phase 1 items should be checked
against the code before being started rather than after.
2026-08-31 18:41:19 +02:00
jschoubben fdd909ec40 Phase 1.1 done, and it was not the task that was written down
An object-store provision, proven against a real store with seven
assertions.

The finding is worth more than the task: the control plane
special-cases nothing. provides, requires, contributes and grants are
name-agnostic, so asking for a bucket needed no change to the mesh at
all. What was missing was a provider and the last step on the machine —
"add an object-store provision" was never mesh work, and the breakdown
now says so rather than leaving the next person to rediscover it.

Named s3-bucket by 0027: the coupling is to the API, not the product,
because swapping one store for another does not break a consumer. A
database is the other case and names its engine.

Records the assertion a database does not need, because it is the one
that will be forgotten when somebody writes the next provider: one store
holds every bucket behind one endpoint, so isolation is a policy rather
than a property, and a policy granting everything passes every test that
only checks a consumer can reach its own bucket.
2026-08-31 17:53:02 +02:00
jschoubben cb1954e7b5 Accept 0024, and rewrite the work breakdown around what is actually being done
**0024 accepted.** Model access was decided, built, and proven in the
lab, and two design documents rest on it; only the status had never
moved. The gate is green again.

**The work breakdown rewritten.** It planned a decomposition of the
existing system in place — extract contexts, declared features, shrink
the shared library. That is not the work. A replacement is being built
beside it, and only the old Phase 0 survived contact with reality, so
the one document meant to say what happens next was describing a system
being retired.

Now ordered by what "modules move across one at a time until the old
registry is off" actually requires:

- Phase 0 is marked done against the twenty-two lab assertions, **and
  carries its own limitation**: every module exercised was written to
  test the mechanism, so the vocabulary was shaped by its own fixtures.
- Phase 1 is the vocabulary gaps found by asking what real modules
  need — an object-store provision, a session as a licence consumer, a
  network shape with ordering, public certificate issuance.
- Phase 2 is one module, then a week of running it, because the point of
  going first is to find what Phase 1 missed.
- Phase 3 picks modules that each prove something the first did not; the
  mail system is last because it is the one that may send work back into
  the declaration language.
- Phase 4 is switching the registry off, named as a phase so it is not
  mistaken for the goal.

Keeps the rules of engagement unchanged — they were about how work is
done, not what it is — with one addition: stop and ask before anything
that touches a machine outside the lab.

Adds a section on keeping the list true, since the document it replaces
was wrong for weeks and nothing said so. A claim here is counted, not
reasoned, and a phase is done when the lab says so.
2026-08-31 17:25:27 +02:00
jschoubben 1b5308c9cc Review of the to-be layer: check what the documents claim against what runs
First pass of a design review, done by reading documents against code
and against a raised mesh rather than against each other. Every error
below was invisible to a proofread.

**Statuses were stale, and nothing checked them.** Ten to-be documents
said `designed` while naming working, lab-proven code — several with a
*What was built* or *Raised, and observed* section. Added a
`status-vs-code` check: naming a file is a claim that the file
implements this, so a document that points at one has stopped being
merely designed. It failed on all ten before it passed, per the rule
this folder sets for its own checks.

**The bundle carries three images, not two.** 07 reasoned about which
substrate services go in and overlooked that the control plane is in
there too — it is what the substrate exists to start, and there is
nothing to fetch it with yet. Counted, not deduced.

**The bootstrap uses four shapes, not six.** It listed `file` and
`directory`, which substrate-first-node.lock never asks for. The claim
that mattered — nothing is blocked on the host — was true either way,
which is why the wrong count survived.

**The eight capabilities were documented nowhere.** Implemented in
internal/profile/detectors.go and enumerated in no document, including
the one about the host that detects them. A vocabulary modules write
against, readable only by reading the code. Now written down, with the
seat/graphical-session distinction that is wrong in both directions if
collapsed.

**MinIO swept out of the to-be layer** per 0028.

The gate now fails on one thing left deliberately: ADR 0024 is
`proposed` while two documents rest on it and the feature it decides is
built and lab-proven. Accepting a decision is not mine to do.
2026-08-31 17:20:47 +02:00
jschoubben cbcbba8099 A provision names its engine; the substrate supplies only the control plane
**0027 — provisions.** A module written against PostgreSQL could be
matched to a provider of SQL Server, resolve as satisfied, and fail on
its first query. The name said the role, so nothing distinguished
engines. Refusing on ambiguity could not help: with one provider of
each name nothing is ambiguous. Enforced at parse rather than
documented, because the old naming was the documentation.

**0028 — the substrate.** 0006 admits an object store on the grounds
that it cannot grant itself a bucket. That answers the second half of
the test and assumes the first: the control plane does not need one.
Verified — no S3 client in mesh-control, and internal/builder/registry.go
records the deliberate choice to put artifacts in the OCI registry as
content-addressed blobs. The row was inherited from the system being
replaced, where an object store distributed module tarballs, and was
never re-tested against the definition above it.

So an object store is an ordinary module, and a mesh with nothing
needing one runs none. Migrating it is module work, not substrate work.

0028 also states what 0006 left unsaid: a substrate service and a
module of the same product are different instances. The substrate is
raised from the bundle before any mesh exists, so it is not in the
module graph — a workload depending on it would depend on something the
graph cannot see, cannot rotate a credential for, and cannot move, and
would put workload data in the store the control plane keeps its own
state in.

Both records were found by reading code against design rather than
design against itself, which is the review that should have happened
sooner.
2026-08-31 17:13:07 +02:00
jschoubben 046a990198 Several sessions at once is a surface property, not a session one
Answers the question 15 raised: a board showing many sessions leaves
one-per-node untouched, because each is still one conversation. Only
concurrent conversations with the same session would touch 0004.

Records soulstream and herdr as the prior art to draw from, and marks
it explicitly off the provisioning path so it stays a note rather than
becoming the work.
2026-08-31 16:52:15 +02:00
jschoubben fbf5b1d04b A session's memory is its own, and it is not declared
Settles the question 15 left open: the mesh session holds its own
memory in the mesh root, rather than assembling a view over the node
sessions. Memory follows the rule the rest of the design already uses —
the context root is the whole of what makes one session a different
agent, and memory is part of what makes it that agent.

The control-plane node is what makes this load-bearing rather than
tidy. Two sessions share that machine; if memory belonged to the
machine instead of the root they would share it too, and the mesh's
recollection would be indistinguishable from that node's own — the
collision 0026 exists to avoid, arriving through the back door.

Also corrects an error made writing it up: memory is NOT declared
state. The engram and tools are — the mesh says what they are and the
host writes them (0011). Memory is written by the session itself and
declared by nobody, so a mechanism that regenerates the root wholesale
would erase it on the next heartbeat, silently, while reporting
success. The root is not uniformly managed and which parts are has to
be explicit.
2026-08-31 16:41:06 +02:00
jschoubben 3c6c16abdf The mesh has a session of its own, and it is the node session's mechanism
A session for the mesh itself, addressed as the mesh, differing from a
node's in exactly three things: the context it starts in, its engram,
and its licence binding. Not a new kind of agent — the same mechanism
pointed at a different root. Two implementations of one mechanism drift,
and the vocabulary collision 0001 exists to undo began exactly that way.

It runs on the control-plane node, and the reasoning is easy to get
backwards: not "the important agent on the important machine", but that
this node is already the one place excepted from "compromise of a node
is compromise of that node". Placed anywhere else it would create a
second such place.

It is an addition to per-node messaging and never a replacement. 0001
holds that losing the control plane costs change, not operation — and a
mesh whose only conversational surface lived there would lose the
ability to ask anything while every machine kept running perfectly.

Writing it up exposed that the node session's setup was never designed
at all. 0004 gives behaviour and stops: nothing said how a session
starts, where its context lives, or how a broker message becomes a
prompt. That gap was invisible until something had to be built *like* a
node session. 15-the-agent-session.md covers both as one mechanism.

It also makes "a consumer that is not a machine" undeferrable. The
control-plane node now hosts two sessions that must hold different
licences, and a per-machine binding cannot express that at all. Noted in
14-model-access.md against the gap it was already recorded as.

Also completes the to-be index, which stopped at 10 and omitted four
documents. Pre-existing broken ADR references in the older rows are left
alone rather than guessed at.
2026-08-31 16:32:59 +02:00
jschoubben e823cc1cc5 The design record is read where it is written, never copied to be found
Decides the question 006 narrowed to. An agent reads this repository
directly and the search consults it, so these documents surface beside
ordinary results instead of only when somebody already suspects they
exist.

A scheduled sync into the mesh's memory was the option that works with
what exists today, and lost on the ground this repository can least
afford: it makes a second copy, and the copy that is searched quietly
stops matching the copy that is edited. A design record that has
silently diverged from the reasoning it claims to carry is worse than
one that cannot be found — the first misleads, the second merely fails.

Amends what 0019 promised rather than satisfying it: these documents
will not be indexed, they will be read. The commitment that survives is
the one that mattered — that a searcher finds them without already
suspecting they exist.

Gated on an agent that does not exist yet, so 006 stays open on the
build with a decided shape. What closes it is a check that fails today
by design: search the mesh's memory for a phrase that appears only in a
design document here, and require it back.
2026-08-31 15:45:00 +02:00
jschoubben c192810fba 004 and 008 resolved; 006 narrowed to the decision it actually needs
**004 — certificate issuance.** The resolver declared no authority at
all, so the client fell to its built-in production default: there was no
setting set wrongly, there was no setting. It is now a node property
defaulting to staging, which answers the first open question. Staging by
default rather than production-with-an-override, because the alternative
leaves the safe path depending on remembering to opt out of it — 005's
lesson, in a second place. The rollout is ordered and the order is the
dangerous part; recorded, not performed.

**008 — node rescue.** Read back from running nodes as the report asked,
and one of its own claims was wrong in a way that matters: the health
timer does exist and does fire. It simply never calls the rescue script.
A trigger that exists and does not do what the script claims survives a
halfway check, which makes it worse than the absence the report
described. Resolved by making the documentation true, not by
implementing rescue — the replacement host already supervises recovery,
and wiring unattended restart into the fleet being retired is a
deliberate decision rather than a tidy-up. Two "self-healing" claims
narrowed to what they actually do.

**006 — deliberately not closed.** Re-checked today: the indexing still
does not exist. What is gone is the reason it was an issue — the claim
is no longer load-bearing, because the README names the gap and the
decision's reasoning never invoked indexing. A signpost now points here
from the knowledge base, and was measured rather than assumed: it is
reachable, it is not surfacing. Closing it while the indexing does not
exist would be this repository's own named failure, one folder from
where it names it.
2026-08-31 15:20:11 +02:00
jschoubben 345bbe0552 005 resolved: a suite that cannot run on every push says when it last ran
Retired in favour of the lab rather than repaired — that answers the
first open question. The second finding is the one that generalises:
"nothing runs it, and nothing reports that nothing runs it" is not a
fact about that harness, it is a fact about any suite too expensive to
run on every push. The replacement inherited the fault it was replacing.

Records the three rules that now hold, and what the fix taught twice:
the remedy rebuilt the symptom inside itself, and the code that counts
results passed every test while reading nothing.
2026-08-31 15:02:26 +02:00
jschoubben f3ffdae909 Record the resolver as built, and the two things it must not do
A service is reached at <service>.<node>.internal, so what resolves is anything
under a node's name. The mesh writes the data and runs no daemon; two roles,
two claims, because systemd-resolved cannot serve a wildcard at all.

Both prohibitions were found by a machine rather than by reasoning: an address
systemd already held, and reading resolv.conf for upstreams that now point at
itself.
2026-08-31 14:23:02 +02:00
jschoubben 6ecd03694b Issue 019: a comment asserting a fact about a machine, which nothing checked
Twice in one file, a statement about a machine that read as reasoned and was
wrong — and the module's unit tests all passed while the daemon could not
start. That is what a unit test is: it confirms the assertion was made, never
that it is true of any machine.

003 in prose rather than in a manifest key.
2026-08-31 14:11:34 +02:00
jschoubben a036bac47b Record why the module's own secret is named for whose it is
The field was called needs, beside secrets, and both were name-to-path holding
something secret. What separates them is whose, not how secret — so that is
what the name says now.
2026-08-31 13:47:21 +02:00
jschoubben aafeb5c9df Three issues resolved: one closed by evidence, two answered by the replacement
012 named its own closing condition — a scenario with four images coming up —
and the scenario now stocks seven and has raised cleanly many times at the
memory the wrong diagnosis had raised.

001 is answered by the host reading the package database back after installing.
002 was NOT answered and was present here too, so it is a fix rather than a
note: a stale index is now named instead of reported as a failed install.
2026-08-31 13:00:03 +02:00
jschoubben 573a94e102 Correct the record: a limitation that no longer exists, and one that was never written
The connectivity design still said a hub cannot be filtered — a gap recorded in
the morning and closed in the afternoon, left standing as though it were
current. Worse than a stale date: it would send somebody away from something
that works.

`restart-on` was described nowhere, including the part added today that lets a
service reflect a file another module put on the machine. A rule the host
enforces and no document mentions is a rule nobody can rely on.

And nine of fifteen design documents claimed an `updated:` older than their last
change, some by a week. That field is what cross-cutting views are generated
from, so it is not decoration.
2026-08-31 12:33:35 +02:00
jschoubben f4e81074d3 Which resolver is a claim, and was decided before it was asked
A resolver takes over /etc/resolv.conf, which is a singular resource — ADR 0009
lists it in the table beside the seat and pid 1. So choosing between resolved,
dnsmasq and unbound is assigning a module, per machine, and the mesh refuses
two rather than letting them fight over the file.

Recorded because it was treated as an open question two days after being
decided, which is the argument for that table being a table.
2026-08-31 12:15:12 +02:00
jschoubben a3cee17d48 Record what a container can see of the mesh's names, and what it cannot
Found by a container failing to resolve a name every machine could: a container
gets its own hosts file holding only its own hostname, and on the machine it
always worked, which is what made it easy to miss.

Declared containers are given the names. A container somebody starts by hand is
not the mesh's to configure — which is a second, different reason to want a
resolver, recorded beside the first rather than folded into it.
2026-08-31 11:29:25 +02:00
jschoubben 35ce23fb76 Record that a machine coming back is ordinary, and what waking now does
Asked whether a machine that drops off needs re-adopting: it does not, nothing
expires, and the only thing that forces re-enrolment is losing its own key.

The gap was the twenty or thirty seconds after a resume in which a node
believes it is in a mesh it has left — recovering on its own, which made it a
quality gap rather than a fault, and still a machine waiting to be told
something it already knew.
2026-08-31 10:21:50 +02:00
jschoubben 6374c1eb60 Record what "behind" means now
It meant failed-or-refused, so the question this record says must not be lost
was answerable only for the machines that broke. Out of date, never told, and
not worked out are kept apart: the remedy is the same push and they read
differently to whoever is looking.
2026-08-31 05:28:11 +02:00
jschoubben 7c0be6968e Record what a machine says about itself and what the mesh keeps
The yes gates an assignment and the detail carries a value; they are one fact
read two ways, and only one read was being kept.
2026-08-31 04:54:07 +02:00
jschoubben 7f314eb399 Point three design documents at the code that exists for them
Their subject matter has been built and proven for days and their frontmatter
still said code: [] — which is what the cross-cutting view is generated from,
so it was claiming nothing existed for the substrate, the node lifecycle and
delivery.
2026-08-31 04:51:04 +02:00
jschoubben 5ad3641cbf Record the board, which is the last designed document with no code
One reading answered three ways, holding nothing and touching no context's
store — which is the constraint the whole document is about, and the thing the
board being replaced gets wrong.
2026-08-31 04:50:45 +02:00
jschoubben 0bc4b7774f Record what four more pieces of the mesh became
Rotation and the provisioner contract; model access as a provision answered by
a record, with ADR 0024's other two gaps left as gaps; exposure, which closes
the open question about revoking a route; and the delivery loop, which closes
the gap ADR 0010 left when it replaced a pipeline with a comparison.
2026-08-31 02:56:50 +02:00
jschoubben 6b1c80b442 What a builder-as-a-module can and cannot reach
The broker's fingerprint travels with its credential, and the machine's
filesystem does not travel at all — it runs in a container, which is the
arrangement working rather than a limitation to route around.
2026-08-31 01:53:35 +02:00
jschoubben 8448219de1 Issue 018: a provider on the same machine was never announced to its consumer 2026-08-31 01:35:54 +02:00
jschoubben 00e98f1f92 The vocabulary has no word for a unit that runs and exits
Found by the firewall: every packet filtered as declared, and the machine
reported as not doing what it was told, because the unit that loaded the rules
had finished. Stated as a gap rather than worked around silently.
2026-08-31 01:20:58 +02:00
jschoubben 0ac99f0d68 An action's own idea of being finished must be its verify's
Otherwise it succeeds into a state its verify rejects, and the host's report is
accurate and names nothing. Recorded where the vocabulary is described, because
it is a rule about writing an action rather than about one action.
2026-08-31 01:03:50 +02:00
jschoubben 1ade18209d Issue 017: an action succeeded into a state its own verify rejects 2026-08-31 01:01:17 +02:00
jschoubben a9cd3de5be Issue 016: anything after the declaration in a file was ignored 2026-08-31 00:51:46 +02:00
jschoubben 6e7e77acbe Issue 015: the harness read a swallowed answer as success 2026-08-31 00:48:16 +02:00
jschoubben 890c3ee3fc State the one rule the derivation does not yet reach
A hub needs its overlay port open and a node that is not a hub does not, and
they are the same module — so listens, a static manifest field, cannot express
it while the overlay module's resources are computed per node. Written down
rather than left as an oversight for whoever first puts a firewall on a hub.
2026-08-31 00:42:39 +02:00
jschoubben a6872ac099 A key that is present and unusable, and what the certificate work became
Issue 014: the node's serving key was stored in the host's own encoding, so
every check that reads the file passed and no server could start. Same shape as
013 — two halves of one mechanism designed separately, each correct about its
own half. Where a file exists so a third party can read it, the format is the
interface.
2026-08-31 00:42:13 +02:00
jschoubben 778efaba8b The mesh runs its own registry, certifies its own names, and computes its own filtering
Issue 003 is answered in both halves: manifests are parsed strictly, and a
module says what it listens on and from where rather than carrying a key
nothing reads. The design records what was built and how each part is checked.

Issue 013 is new, found by reading while writing the first module that has
both a computed file and a service that needs it. The file arrived second.
It failed, then the next reconcile fixed it, which is why nothing caught it.
2026-08-31 00:37:34 +02:00
342 changed files with 30237 additions and 551 deletions
+78
View File
@@ -0,0 +1,78 @@
---
name: hq-defer
description: Use when a thought is raised that should be remembered but NOT worked on now — an aside during other work, a "we should look at X someday", a known gap nobody is assigning yet. Triggers on "defer this", "park this", "register this thought", "note this for later", "don't work on it, just remember it". Records and returns to whatever was already in progress.
---
# hq-defer
Parks a thought so it is not lost, **without moving the work off course.** The defer is the
point: the thought is recorded and the previous task resumes.
This skill wraps no playbook, because deferring is not part of the development cycle — it is
what happens *before* something enters it. A parked thought has no number, no owner and no
status, and that is correct.
## Where it goes, and why not the repository
Record it as **one memory file** in Claude's persistent memory directory for this project (the
path is given in the session's memory instructions), with `metadata.type: project`, plus a
one-line pointer in `MEMORY.md`.
**Not** in `04-ISSUES`, `01-RESEARCH` or anywhere else in the repository:
- A parked thought is not an issue or a research effort. Giving it a number asserts it has been
triaged, which is exactly what deferring says has not happened.
- A shared "deferred" or "someday" document is a **central status file**, which
[`AGENTS.md`](../../../AGENTS.md) forbids. Status lives in frontmatter on real records, and a
parked thought has no real record yet.
- A repository write means a branch, a commit and a pull request — drift, which is the one thing
this skill exists to avoid.
## Steps
1. Write the memory file. Slug is kebab-case and descriptive of the thought, not of the act of
deferring.
```markdown
---
name: <kebab-slug>
description: Deferred note — <one line>
metadata:
type: project
---
Raised and deliberately deferred on YYYY-MM-DD: **<the thought, in the user's own terms>**
**Why:** what was being worked on when it came up, and that deferring was intentional so
that work was not pulled off course.
**How to apply:** treat as an open thread, not an assignment. Do not start on it
unprompted. If it graduates it needs an HQ home first — an issue under playbook
[03](../../../00-META/process/03-issues.md) if a stated behaviour does not happen, or
research under playbook [01](../../../00-META/process/01-research.md) if it is still an
idea. Say which is undetermined, if it is.
```
2. Append one line to `MEMORY.md`: `- [<Title>](<kebab-slug>.md) — deferred YYYY-MM-DD; parked, no HQ
record, do not start unprompted`.
3. Convert relative dates to absolute before writing. "Last week" is worthless in six months.
4. Check for an existing memory covering the same thought and update it instead of adding a
duplicate.
## Then stop
Reply in **at most two lines** — what was recorded, and that it is parked — and **return to
whatever was in progress before.** Do not summarise the parked thought back at length, do not
propose a plan for it, do not ask which playbook it belongs to, and do not open anything.
If nothing was in progress, say only that it is recorded.
## Do not
- Do not create an issue, a research effort, a decision record or a design document.
- Do not create a branch, commit or pull request.
- Do not start investigating the thought, however cheap the first check looks.
- Do not name nodes, domains, addresses, absolute paths or usernames in the memory file — the
thought may later be quoted into this repository, which is public.
- Do not decide whether it is an issue or research when the evidence does not say. Recording
"undetermined" is the honest outcome and costs nothing later.
+1
View File
@@ -9,6 +9,7 @@ works — plus the engineering practice that holds across everything Novox build
| [`context.md`](context.md) | The environment — conditions, not aspirations | | [`context.md`](context.md) | The environment — conditions, not aspirations |
| [`effect.md`](effect.md) | What is different when the work is done | | [`effect.md`](effect.md) | What is different when the work is done |
| [`how-we-build.md`](how-we-build.md) | The rules that hold across the mesh, each one earned. **The source of the mesh constitution** — the governed page the mesh injects into design sessions is derived from it. | | [`how-we-build.md`](how-we-build.md) | The rules that hold across the mesh, each one earned. **The source of the mesh constitution** — the governed page the mesh injects into design sessions is derived from it. |
| [`glossary.md`](glossary.md) | One name per thing — the authority on vocabulary, and the words that were retired |
| [`repos.md`](repos.md) | Where implementation lives, and what each repository owns | | [`repos.md`](repos.md) | Where implementation lives, and what each repository owns |
| [`process/`](process/) | The playbooks — how work moves through this repository, for engineers and agents alike | | [`process/`](process/) | The playbooks — how work moves through this repository, for engineers and agents alike |
+8
View File
@@ -26,6 +26,7 @@ indistinguishable from one that cannot.
| `numbering` | the number in the filename is the number in the heading | — | | `numbering` | the number in the filename is the number in the heading | — |
| `topics` | every record names a topic the index knows | — | | `topics` | every record names a topic the index knows | — |
| *(index.py)* | the written reading order matches what the records say | — | | *(index.py)* | the written reading order matches what the records say | — |
| `status-vs-code` | a to-be document naming specific code is not still `designed` | **ten documents**, several with a *What was built* section, describing lab-proven code |
## What is deliberately not checked ## What is deliberately not checked
@@ -47,3 +48,10 @@ what to do.
State what incident it would have caught, and make it fail before you make it pass. A check State what incident it would have caught, and make it fail before you make it pass. A check
whose failure has never been observed is a guess about its own correctness. whose failure has never been observed is a guess about its own correctness.
## cycle.py
The development cycle, checked ([ADR 0080](../../02-DECISIONS/0080-the-development-cycle-is-checked.md)):
a to-be design names a decision, an in-progress/implemented design names its owning code, a
located/fixed issue names its owner, a fixed/resolved issue says what fixed it, a graduated
research overview says what it became. `python3 00-META/checks/cycle.py`
Binary file not shown.
+188
View File
@@ -0,0 +1,188 @@
#!/usr/bin/env python3
"""The development cycle, checked.
The knowledge flow (00-META/process/00-overview.md) says work moves idea -> research ->
decision -> to-be design -> code, and symptom -> issue -> diagnosis -> fix. Those are rules,
and a rule states how it is checked (AGENTS.md) -- this is how. Everything here reads only
frontmatter, because status lives in frontmatter and nowhere else.
What is enforced:
design every 03-DESIGN doc parses, carries `layer:` matching its directory, and a
known `status:`. A TO-BE doc names at least one decision (`decisions:`) -- no
design without a decision -- and once `in-progress` or `implemented` it names
its owning code (`code:`) -- no development without a design that says where.
issues a known `status:`; once `located`, `located-in:` names the owner;
once `resolved`, `fixed-by:` says what fixed it (prose counts --
"nothing, the capability existed" is an answer).
research a known `status:`; a `graduated` overview says what it `became:`, and every
target it names exists.
decisions every accepted record is REACHABLE from the cycle: cited by a design doc's
frontmatter, a research overview, an issue report, a 00-META document, or another
record's extends/supersedes chain. A decision nothing points at is one nobody will
find by following pointers -- which is how records go stale in people's heads.
Deliberately NOT enforced: `resolved` issues may leave `located-in` empty (a symptom that
turned out not to be a defect has no owner), and as-is docs need no decisions (they
describe what exists, not what was decided).
python3 00-META/checks/cycle.py
"""
import glob
import os
import re
import sys
ROOT = os.path.normpath(os.path.join(os.path.dirname(__file__), "..", ".."))
DESIGN_STATUSES = {"proposed", "designed", "in-progress", "implemented", "abandoned"}
ISSUE_STATUSES = {"open", "diagnosing", "located", "resolved", "wontfix"}
RESEARCH_STATUSES = {"active", "graduated", "abandoned"}
def rel(path):
return os.path.relpath(path, ROOT)
def frontmatter(path):
"""The YAML block between the first two --- lines, as {key: raw-value-string}.
Minimal on purpose, like records.py: enough for the fields these checks read. A list
value (block or inline) is joined into its items; a scalar stays a string.
"""
text = open(path, encoding="utf-8").read()
m = re.match(r"^---\n(.*?)\n---", text, re.S)
if not m:
return None
front, out, key = m.group(1), {}, None
for line in front.split("\n"):
item = re.match(r"^\s+-\s*(.+?)\s*$", line)
if item and key:
out[key].append(item.group(1))
continue
kv = re.match(r"^([A-Za-z-]+):\s*(.*)$", line)
if not kv:
continue
key, value = kv.group(1), kv.group(2).strip()
if value.startswith("[") and value.endswith("]"):
out[key] = [v.strip() for v in value[1:-1].split(",") if v.strip()]
elif value == "":
out[key] = [] # a block list may follow; stays [] if nothing does
else:
out[key] = value
return out
def listy(front, key):
v = front.get(key)
if v is None:
return []
return v if isinstance(v, list) else ([v] if str(v).strip() else [])
def main():
failures = []
def bad(path, why):
failures.append(" %s: %s" % (rel(path), why))
# ---- design ------------------------------------------------------------------------
for layer, name in (("00-as-is", "as-is"), ("01-to-be", "to-be")):
for path in sorted(glob.glob(os.path.join(ROOT, "03-DESIGN", layer, "*.md"))):
if os.path.basename(path) == "README.md":
continue
front = frontmatter(path)
if front is None:
bad(path, "no frontmatter")
continue
if front.get("layer") != name:
bad(path, "layer is %r; this directory is %s" % (front.get("layer"), name))
status = front.get("status")
if status not in DESIGN_STATUSES:
bad(path, "status %r is not one of %s" % (status, sorted(DESIGN_STATUSES)))
if name == "to-be":
if not listy(front, "decisions"):
bad(path, "names no decisions -- no design without a decision")
if status in ("in-progress", "implemented") and not listy(front, "code"):
bad(path, "status %s but code: names no owner -- no development "
"without a design that says where" % status)
# ---- issues ------------------------------------------------------------------------
for path in sorted(glob.glob(os.path.join(ROOT, "04-ISSUES", "*", "00-report.md"))):
front = frontmatter(path)
if front is None:
bad(path, "no frontmatter")
continue
status = front.get("status")
if status not in ISSUE_STATUSES:
bad(path, "status %r is not one of %s" % (status, sorted(ISSUE_STATUSES)))
if status in ("located", "resolved") and not listy(front, "located-in"):
bad(path, "status %s but located-in is empty" % status)
if status == "resolved" and not listy(front, "fixed-by"):
bad(path, "status %s but fixed-by says nothing" % status)
# ---- research ----------------------------------------------------------------------
for path in sorted(glob.glob(os.path.join(ROOT, "01-RESEARCH", "*", "00-overview.md"))):
front = frontmatter(path)
if front is None:
bad(path, "no frontmatter")
continue
status = front.get("status")
if status not in RESEARCH_STATUSES:
bad(path, "status %r is not one of %s" % (status, sorted(RESEARCH_STATUSES)))
if status == "graduated":
became = listy(front, "became")
if not became:
bad(path, "graduated but became: names nothing")
for target in became:
if not os.path.exists(os.path.join(ROOT, target)):
bad(path, "became names %s, which does not exist" % target)
# ---- decisions -------------------------------------------------------------------
records = {}
for path in sorted(glob.glob(os.path.join(ROOT, "02-DECISIONS", "[0-9]*.md"))):
front = frontmatter(path)
records[os.path.basename(path)] = (path, (front or {}).get("status"))
cited = set()
sources = (glob.glob(os.path.join(ROOT, "03-DESIGN", "*", "*.md"))
+ glob.glob(os.path.join(ROOT, "01-RESEARCH", "*", "00-overview.md"))
+ glob.glob(os.path.join(ROOT, "04-ISSUES", "*", "00-report.md"))
+ glob.glob(os.path.join(ROOT, "00-META", "**", "*.md"), recursive=True))
for path in sources:
text = open(path, encoding="utf-8").read()
if os.sep + "03-DESIGN" + os.sep in path:
# A design doc's governing citations live in frontmatter; a prose mention is
# commentary, not a home.
m = re.match(r"^---\n(.*?)\n---", text, re.S)
text = m.group(1) if m else ""
for m in re.finditer(r"([0-9]{4}-[^\s\)\],#]+\.md)", text):
cited.add(os.path.basename(m.group(1)))
for name in records:
front = frontmatter(records[name][0]) or {}
for key in ("extends", "supersedes", "superseded-by"):
v = front.get(key)
if isinstance(v, str) and v:
cited.add(os.path.basename(v))
for name, (path, status) in records.items():
if status == "accepted" and name not in cited:
bad(path, "an accepted decision nothing in the cycle cites -- give it a home in a "
"design doc's decisions:, a research became:, an issue, or 00-META")
checked = (
len(glob.glob(os.path.join(ROOT, "03-DESIGN", "0*", "*.md")))
+ len(glob.glob(os.path.join(ROOT, "04-ISSUES", "*", "00-report.md")))
+ len(glob.glob(os.path.join(ROOT, "01-RESEARCH", "*", "00-overview.md")))
+ len(records)
)
if failures:
print("cycle: %d document(s) break the development cycle:" % len(failures))
print("\n".join(failures))
return 1
print("cycle: %d documents checked, the chain holds" % checked)
return 0
if __name__ == "__main__":
sys.exit(main())
+103
View File
@@ -284,6 +284,107 @@ def check_numbering(failures, records):
) )
def check_progressive_insights(failures, records):
"""A correction made inside a record is marked and dated, or it is a silent rewrite.
A record may be corrected in place when a *fact* in it went stale and the decision still
stands (`02-DECISIONS/README.md`, "Progressive insight"). The whole safety of that allowance
is that the correction is legible in the record rather than only in a diff nobody reads, so
the form is what is checked here: every mention of an insight is the marker, the marker
carries an ISO date, and that date is not earlier than the decision's own — an insight
predating the decision it corrects is a copied marker, not a correction.
What this cannot check is an edit made with no marker at all. Nothing mechanical can; that
one is the reviewer's, reading the diff. The check keeps the *marked* path honest so that an
unmarked change stands out as the anomaly it is.
"""
phrase = re.compile(r"progressive insight", re.I)
# Both patterns stay on one line: a bold run does not span paragraphs, and `[^*]*` across
# newlines will happily join an unrelated `**` far above to the marker below, reporting the
# whole span between them. It did exactly that the first time this ran.
# Trailing words after the date are allowed — "— 2026-09-26, correcting the one above." — so
# an insight can say what it relates to. Only the date's presence and position are fixed.
marker = re.compile(r"\*\*Progressive insights?[ \t]*[\u2014\u2013-][ \t]*(\d{4}-\d{2}-\d{2})[^*\n]*\*\*")
loose = re.compile(r"\*\*[^*\n]*[Pp]rogressive insights?[^*\n]*\*\*")
iso = re.compile(r"^\d{4}-\d{2}-\d{2}$")
for number, record in sorted(records.items()):
text = record["text"]
if not phrase.search(text):
continue
decided = str(record["front"].get("date", ""))
good = [(m.start(), m.end(), m.group(1)) for m in marker.finditer(text)]
for m in loose.finditer(text):
if any(s <= m.start() and m.end() <= e for s, e, _ in good):
continue
# A bold run carrying a link is discussing an insight — usually another record's —
# rather than marking one. A marker never needs to cite anything.
if "](" in m.group(0):
continue
failures.add("insights", rel(record["path"]),
"a progressive insight is not in the dated marked form "
"'**Progressive insight \u2014 YYYY-MM-DD.**': %s" % m.group(0))
for _, _, stamp in good:
if decided and iso.match(decided) and stamp < decided:
failures.add("insights", rel(record["path"]),
"a progressive insight dated %s predates the decision (%s)"
% (stamp, decided))
covered = [(s, e) for s, e, _ in good]
for m in phrase.finditer(text):
if any(s <= m.start() and m.end() <= e for s, e in covered):
continue
line = text.rfind("\n", 0, m.start()) + 1
end = text.find("\n", m.end())
whole = text[line:end if end != -1 else len(text)]
if whole.lstrip().startswith("#"):
continue
# A line that also carries a link is discussing the rule, not marking a correction:
# a marker never needs to cite anything, and a record that reasons about the policy
# must be able to name it. Bare prose with no citation is the informal marking this
# is here to catch.
if "](" in whole:
continue
if loose.search(text, line, text.find("\n", m.end()) + 1 or len(text)):
continue
failures.add("insights", rel(record["path"]),
"'progressive insight' appears unmarked; a correction is marked and "
"dated, or it is a silent rewrite")
def check_status_against_code(failures):
"""A design document naming specific code may not still call itself `designed`.
**Naming a file is a claim that the file implements this**, so the two fields have to agree.
They drifted: ten to-be documents named working, lab-proven code — several with a *What was
built* or *Raised, and observed* section — while still saying nothing had been built.
Deliberately weak, and that is the point of it being mechanical. It cannot tell whether the
prose is true, only that a document has stopped claiming to be unbuilt once it points at
something. `code: [mesh-controller]` — a repository with no path — is a plan and stays
`designed`.
"""
for path in markdown_files():
if not rel(path).startswith("03-DESIGN/01-to-be/") or path.endswith("README.md"):
continue
front = frontmatter(read(path))
if front.get("status") != "designed":
continue
for entry in front.get("code") or []:
named = re.sub(r"\s*\(.*\)$", "", entry).strip().split(None, 1)
if len(named) > 1:
failures.add(
"status-vs-code",
rel(path),
f"`designed`, but names {named[1]!r} in {named[0]}. Naming a file claims "
f"it implements this — use `in-progress`, or `implemented` once it is "
f"defensible from that repository's main branch.",
)
break
def main(): def main():
failures = Failures() failures = Failures()
records = load_records() records = load_records()
@@ -293,6 +394,8 @@ def main():
check_supersession_symmetry(failures, records) check_supersession_symmetry(failures, records)
check_numbering(failures, records) check_numbering(failures, records)
check_topics(failures, records) check_topics(failures, records)
check_status_against_code(failures)
check_progressive_insights(failures, records)
print(f"records: {len(records)} decision records checked") print(f"records: {len(records)} decision records checked")
return failures.report() return failures.report()
+1 -1
View File
@@ -30,7 +30,7 @@ named, and nothing should be designed around a particular one existing.
- **A hosted model provider** supplies the thinking for non-human agents, drawn from a - **A hosted model provider** supplies the thinking for non-human agents, drawn from a
shared pool of subscriptions — which is why budget pacing is a first-class concern. shared pool of subscriptions — which is why budget pacing is a first-class concern.
- **Long-lived user services** rather than an orchestrator. No cluster scheduler, no cloud - **Long-lived user services** rather than an orchestrator. No cluster scheduler, no cloud
control plane. controller.
Defaults, not mandates. A second model provider is anticipated by design; nothing in the Defaults, not mandates. A second model provider is anticipated by design; nothing in the
domain may assume one vendor's credential lifecycle. domain may assume one vendor's credential lifecycle.
+82
View File
@@ -0,0 +1,82 @@
# Glossary — the words this repository uses, and the ones it stopped using
One name per thing. This page is the authority; where an older record says something else, that
record is being superseded, not this page. It exists because the terms kept drifting in
conversation — control plane / controller / master / hub for one thing, substrate / foundation for
another — and a mesh you cannot name precisely is a mesh two people describe differently.
## The mesh and its machines
- **node** — a machine in the mesh. There are 0..n of them, and each runs the host agent. A node is
just a machine that has joined; being one implies nothing about what it runs.
- **control-node** — the one node that also holds the `mesh-controller` seat. There is exactly one
per mesh. "control-node" is not a separate kind of machine — it is a node that additionally runs
the controller (and, today, the foundation). Lose it and the other nodes keep running what they
were last told; they simply cannot be told anything new.
- ~~master / slave~~, ~~hub / peer~~ — not used. The relationship is *controller and nodes*, and no
node is subordinate: a node applies declarations on its own and survives the control-node dying.
## What runs the mesh
- **controller** — the component that decides what each node should be, holds the mesh's records,
and tells nodes over the broker. Replaces **"control plane"** (borrowed from networking's
control-plane/data-plane, and opaque here).
- **mesh-controller** — the module that runs the controller. It **claims** the `mesh-controller`
seat at mesh scope, which is what makes it singular. Replaces the module name **`mesh-control`**.
(The git repository has been renamed `mesh-control` -> `mesh-controller` on the forge; the module,
container and image it produces are `mesh-controller`.)
- **foundation** — the store and the broker, raised at genesis before any module system exists.
Replaces **"substrate"** (a biology metaphor that landed for no one). The foundation is not a
third thing beside the store and broker — it *is* those two, named together.
- **store** — the one postgres server. It holds the controller's own context databases
(`inventory`, `identity`, `licences` — a context owns its store, [ADR 0008](../02-DECISIONS/0008-a-context-owns-its-store.md))
and every module's own database. One server, many databases — never one shared "mesh database".
- **bus** — the mesh's own nervous system: NATS, one per mesh, carrying every link the mesh has —
control, declarations, builds, events, tool calls
([ADR 0106](../02-DECISIONS/0106-the-bus-is-nats.md)). A module reaches it by requiring
`mesh-bus` ([ADR 0128](../02-DECISIONS/0128-the-mesh-bus-is-required-not-ambient.md)); one that
does not require it has no account on it. Held by the `mesh-broker` seat, which is named after
the *role* rather than the server, so the server can change without the seat doing so.
- **the deprecated broker** — the lavinmq module. It was the mesh's bus and is not any more. It
keeps running as an **ordinary provider** of the `amqp` provision, for modules that need a
message broker of their own the way something needs a database
([ADR 0127](../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md))) — no seat, not foundation,
never raised at genesis, and a mesh that never installs it is complete.
Say *the deprecated broker*, not "the compatibility broker" (it serves the mesh's own modules,
not only the predecessor's) and not "the AMQP broker" (naming it after a protocol invites
describing the bus by contrast with it, which is backwards: the bus is the mesh's nervous
system and this is a module).
## What the mesh stores and serves
- **package** — what code resolves when it is **compiled**: an npm/cargo/pypi dependency, by
**version**. Served by the **package-registry** (gitea). Only a builder talks to it.
- **artifact** — what the mesh delivers to a machine to **install and run**: an OCI image, by
**digest**. Served by the **artifact-store** (distribution). Every node pulls from it.
- These are two protocols, not one store being weak — see [ADR 0075](../02-DECISIONS/0075-two-stores-and-which-provides-what.md).
## How modules relate to the mesh
- **seat** — a named role at a scope (node / site / mesh), held by a module assignment, from a
**closed set** the mesh defines: a claim naming a seat outside the set is refused. A seat may
**deliver a provision**, and its holder is then the mesh's answer for it when several modules
provide it ([ADR 0126](../02-DECISIONS/0126-a-module-declares-its-own-seats.md) (superseding [ADR 0110](../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md))).
The set, with who holds each seat, is the overview of what a mesh has
([26 — The seats](../03-DESIGN/01-to-be/26-the-seats.md)). A seat has a **capacity**: a
capacity-1 seat is exclusive (one holder); a higher-capacity seat is a **bench** (several holders
coexist).
- **claim** — a module taking a spot on a seat. `claims: [{name, scope}]` in a manifest. A
mesh-scoped exclusive claim is how the mesh says "there is one of me". A foundation seat is
named after the server it guards: the `mesh-controller`, `postgres` and `lavinmq` modules claim
the `mesh-controller`, `mesh-store` and `mesh-broker` seats ([ADR 0079](../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md)).
- **provision** — a service one module `provides` and others `require`; the mesh resolves a provider
and wires the two with an endpoint and a credential. A provision is a service you offer, a seat
is a role you occupy, and the two meet where a seat delivers a provision: occupying the seat is
what makes a module *the* provider of it.
## How this page is kept
A new name for an existing thing lands here first, in the same change that introduces it in code. A
record under `02-DECISIONS/` keeps whatever word it was written with — those are immutable — so a
term retired here may still appear there, and the mapping above is how to read it.
+18 -1
View File
@@ -1,6 +1,10 @@
# Process — overview # Process — overview
How work moves through HQ, and who may do what. Every other document in this folder is a How work moves through HQ, and who may do what. In industry terms this is **spec-driven
development, with provenance**: the decision is the why, the design doc is the spec, `code:`
names the implementation, and the lab beds are the conformance tests — and unlike the common
form, the chain itself is checked ([ADR 0080](../../02-DECISIONS/0080-the-development-cycle-is-checked.md),
[0081](../../02-DECISIONS/0081-a-decision-nothing-cites-is-not-yet-in-the-chain.md)). Every other document in this folder is a
playbook: trigger, who runs it, steps, outputs. Engineers and agents follow the same playbook: trigger, who runs it, steps, outputs. Engineers and agents follow the same
playbooks; agents must not act outside them. playbooks; agents must not act outside them.
@@ -49,6 +53,19 @@ the expensive half.
| [03](03-issues.md) | Issues | Something is wrong — often with the owner unknown | | [03](03-issues.md) | Issues | Something is wrong — often with the owner unknown |
| [04](04-build-handoff.md) | Build handoff | A design is ready to be built | | [04](04-build-handoff.md) | Build handoff | A design is ready to be built |
| [05](05-constitution-sync.md) | Constitution sync | `how-we-build.md` changed a rule the mesh enforces | | [05](05-constitution-sync.md) | Constitution sync | `how-we-build.md` changed a rule the mesh enforces |
| [06](06-writing-a-module.md) | Writing a module | Something that runs today must run on the mesh |
| [07](07-feature-branches.md) | Feature branches across repos | Work that changes code, in one repo or several at once |
## The cycle is checked
The flow above is a rule, and a rule states how it is checked:
[`00-META/checks/cycle.py`](../checks/cycle.py) refuses a to-be design that names no
decision, an `in-progress`/`implemented` design that names no owning code, an issue marked
`located`/`fixed` with no owner or `fixed`/`resolved` with no fix, and a `graduated`
research overview that does not say what it became. Run it with `records.py` and `index.py`
before any HQ merge. What the checks cannot see — that code work actually started from a
handoff — is held by playbooks [04](04-build-handoff.md) and [07](07-feature-branches.md):
a feature branch exists because a design or an issue sent it.
## Status lives in frontmatter ## Status lives in frontmatter
+6 -2
View File
@@ -33,7 +33,10 @@
A design changes only through a decision. A design changes only through a decision.
1. Write the decision record. If it reverses an earlier one, the earlier record's `status:` 1. Write the decision record. If it reverses an earlier one, the earlier record's `status:`
becomes `superseded-by: 02-DECISIONS/NNNN-....md` — **its text is never edited**. becomes `superseded-by: 02-DECISIONS/NNNN-....md` — **its reasoning is never rewritten**. If the
earlier record is sound and only a *fact* in it went stale, that is a **progressive insight**,
corrected in place and marked in the record rather than superseded
([`02-DECISIONS/README.md`](../../02-DECISIONS/README.md)).
2. Edit the to-be design document and set `updated:` to today. 2. Edit the to-be design document and set `updated:` to today.
3. If the amendment came from an issue, set that issue's `amended-design:` to the document 3. If the amendment came from an issue, set that issue's `amended-design:` to the document
path. path.
@@ -54,4 +57,5 @@ Implementation state is a third axis, independent of both design and decision.
- Do not move a to-be document into `00-as-is/`. Write the as-is document; both stand. - Do not move a to-be document into `00-as-is/`. Write the as-is document; both stand.
- Do not edit an as-is document to describe an intention. That is what the to-be layer is for. - Do not edit an as-is document to describe an intention. That is what the to-be layer is for.
- Do not change a decision record's meaning. Supersede it. - Do not change a decision record's meaning. Supersede it. Correcting a fact it got wrong, while
the decision stands, is a progressive insight — marked and dated in the record, never silent.
+3
View File
@@ -1,5 +1,8 @@
# Playbook 05 — Constitution sync # Playbook 05 — Constitution sync
Implements [ADR 0021](../../02-DECISIONS/0021-hq-is-the-source-of-the-constitution.md): HQ is
the source of the constitution, and the knowledge-base page is derived, never edited.
**Trigger.** [`how-we-build.md`](../how-we-build.md) changed a rule that the mesh enforces at **Trigger.** [`how-we-build.md`](../how-we-build.md) changed a rule that the mesh enforces at
runtime. runtime.
+107
View File
@@ -0,0 +1,107 @@
# Playbook 06 — Writing a module
**Trigger.** Something that runs today must run on the mesh, or a new capability must be
declarable.
**Who runs it.** Whoever is porting or writing it.
*Written 2026-09-01 from doing this for the first time end to end. Every step below exists
because skipping it cost something.*
## Before anything: read what runs
**A module is written from the thing, not from memory of the thing.** For a port, that means its
current compose file, its environment, and where its data actually sits. Assumptions about any of
the three have been wrong every time they were not checked.
Three questions, answered from the machine:
| | why it decides something |
|---|---|
| **what containers, and how do they find each other?** | more than one means a `network`; names between them must match what the software is configured to dial |
| **where is its data?** | a bind mount moves with a path; a named volume does not; an anonymous volume is already losing data on every redeploy |
| **which values are secret, and which are merely settings?** | a secret goes in `own-secrets` or a grant; a setting goes in the manifest and may be overridden per node |
## The steps
1. **Name what it provides and requires**, if anything. A name is what a consumer is coupled to,
not the role it plays ([ADR 0027](../../02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md)):
`postgres-database`, not `database`. Most modules provide nothing and require nothing — an
application is usually a leaf.
2. **Declare capabilities, not dependencies, for facts about the machine.** `container-runtime`,
`package-manager`, `seat`. A capability is detected and refused against; it is not something a
module can install.
3. **Write the resources in the order they must happen.** They are applied in the order written
and orphans are removed in reverse, so a `network` is written before the containers that join
it and removed after them.
4. **Put every secret in a file, never in `env`.** A declaration travels over the broker in plain
text: a password in `env` is a password the broker sees. **The mesh delivers parts; a module
that needs them combined combines them.**
**A sealed file holds the password and nothing else** — no key, no `=`, no newline that means
anything. So `env-file` must never point at one. It points at a file the module *declares*,
whose content leaves a hole:
```
own-secrets superuser → /var/lib/postgres/superuser.secret the password, alone
a file /var/lib/postgres/superuser.env, mode 0600,
content: POSTGRES_PASSWORD=${secret:superuser}
the container env-file: [/var/lib/postgres/superuser.env]
```
The host fills the hole on the machine, which is the only place both halves exist — the mesh
discarded the value
([credentials and their rotation](../../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md)).
A **provisioner** is the exception: it reads a password file, so it mounts the `.secret`
directly.
Every example module in `mesh-controller` had this wrong and shipped: `own-secrets` pointing at a
path *named* `.env`, mounted as `env-file`, holding a bare password. Docker reads that as a
malformed line and the container starts **with no password set at all** — not a failure to
start, a service running on the wrong credential. They parsed and they resolved. Two tests in
`examples/modules` now refuse both halves of it.
Add `restart-on` naming the env file, or the container keeps the credential it started with
through every rotation.
5. **Pin every image by digest.** A tag moves. The manifest in a repository names artifacts; the
manifest the mesh holds names digests, and they are not the same document.
6. **Decide generate or accept.** A new module's credential is generated. **An adopted one keeps
the credential it already has** — `secret accept` — because minting a new password for a
database that already exists locks the application out of its own data.
7. **Add a provisioner only if the software cannot read a file.** A proxy that watches a
directory needs nothing. PostgreSQL needs `CREATE ROLE`, an object store needs a bucket and a
policy, an identity provider needs a realm and a client — those need a small program beside
them. It reads what the mesh granted and reconciles; it does not decide anything.
8. **Prove it in the lab, against the real software.** Not that a container started — that the
thing works: the credential authenticates, a wrong one is refused, the containers reach each
other, the data survives a restart.
## What the first port actually cost
Six attempts, one real bug. Recorded because the ratio is the lesson: **the mesh was right every
time and the scaffolding was not.**
- A shape existed in the language and no host implemented it, so every declaration carrying one
was refused whole — correctly, and the host said exactly that. **Nobody was reading the host's
log.** Read it first; it is the only place that says why a machine did nothing.
- A blind find-and-replace renamed a provision in quotes and missed the same word bare.
- A command was tested only for the invocations that should fail, so it rejected every real one
and the suite stayed green.
- A test asserted on a helper rather than on the code that calls it, three separate times. **A
test that cannot fail when the behaviour is deleted is not defending the behaviour.**
## Rules
- **Read the host's log before theorising.** A declaration that was sent and not applied says so
there and nowhere else.
- **A failing test is kept, not skipped.** It is the reproduction.
- **Never rotate during an adoption.** Rotation is a separate act, afterwards, deliberately.
- **A data directory is never removed by the mesh** ([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)),
and that protects against the mesh only — not against a disk or a mistaken command.
+69
View File
@@ -0,0 +1,69 @@
# Playbook 07 — Feature branches across repos
**Trigger.** Work that changes code — in one code repo or in several at once (`mesh-sdk`,
`mesh-controller`, `mesh-catalog`, `mesh-host`, `mesh-lab`, and `hq` when a decision rides along).
**Who runs it.** Anyone who writes code, engineers and agents alike. Agents follow it exactly —
it is the guard against the failure it was written for.
## The failure it prevents
A feature was worked as a branch-and-MR per *unit of thought* — one per decision, one per
stacked increment — and each MR was treated as finished when it was *opened*, not when it was
*merged*. Across repos the same feature took a different branch name in each. The MRs piled up
unmerged: one session left **sixteen** stacked intermediate MRs that had to be consolidated and
closed by hand. An MR is a review checkpoint, not a scratchpad.
## The rule
One feature is **one branch name**, **one worktree per repo**, **one MR per repo**, opened
**once, at the end**.
1. **Name the feature once.** `feat/<slug>`. The *same* branch name in every repo the feature
touches — never a different name per repo, never a fresh branch per increment within the
feature.
2. **Isolate each repo.** One git worktree per touched repo under `.work/<slug>/<repo>`, branched
off `main`:
```
git worktree add .work/<slug>/<repo> -b feat/<slug> origin/main
```
Parallel features never collide, and no shared checkout is edited.
**Add the untouched siblings the lab reads.** The lab beds find the other repositories by
sibling path from the lab checkout (`../mesh-tools/module.json`, `../mesh-sdk`, …), the way
the main layout has them. A `.work/<slug>/` directory holding only the touched repos fails a
bed at once with `no manifest for mesh-tools at …/.work/<slug>/mesh-tools/module.json`, after
genesis has already passed. Give the directory those repos as **detached worktrees on `main`**
— never symlinks:
```
git worktree add --detach .work/<slug>/mesh-tools main
```
3. **Commit as you go — locally.** Increments land on the one branch. Nothing is pushed and no
MR is opened mid-feature.
4. **Finish, then publish.** When the whole feature is done — every repo, tests green — push
every branch and open **one MR per touched repo**, together.
5. **Merge promptly, once approved.** Every merge into `main` is notified and approved
([ADR 0023](../../02-DECISIONS/0023-approval-is-the-checkpoint.md)); once it is, merge —
do not leave it sitting. The branch is deleted on merge.
6. **Leave nothing behind.** After the MRs merge, no `feat/<slug>` branch and no `.work/<slug>`
worktree survive.
## What this is not
- **Not a licence to batch unbounded work.** A feature is a *bounded* unit; if it sprawls for
days, end-of-feature bloat merely replaces per-increment bloat. Split it into features, each
its own branch and MR.
- **Not a second trunk.** Every repo branches off `main`. There is no longer an `initialization`
trunk.
## How it is checked
The end state is visible, and its absence is the smell:
- After a feature merges, `git branch -r | grep feat/<slug>` and `git worktree list` return
nothing for it. A surviving branch or worktree means step 6 was skipped.
- A bed run from `.work/<slug>/mesh-lab` that fails naming a `.work/<slug>/<repo>/…` path it
cannot find is a missing sibling worktree, not a mesh fault.
- More than one open MR in a repo that share no feature name, or a stack of MRs none of which is
merged, is the failure this playbook exists to prevent — stop and consolidate before opening
more.
+10 -7
View File
@@ -1,6 +1,6 @@
--- ---
status: canonical status: canonical
updated: 2026-08-23 updated: 2026-09-24
--- ---
# The Novox repositories # The Novox repositories
@@ -17,21 +17,24 @@ and a forge address is an operational detail (see [`README`](../README.md)).
| `hal` | The monorepo — the node runtime, the module catalogue, the delivery machinery, and the bootstrap scripts. Every core module lives here. | | `hal` | The monorepo — the node runtime, the module catalogue, the delivery machinery, and the bootstrap scripts. Every core module lives here. |
| `hq` | This repository, under the company organisation — mission, research, design, decisions, issue diagnosis. Company-scoped ([ADR 0019](../02-DECISIONS/0019-how-this-repository-works.md)); the mesh is its first product. The source of truth for *why*. Carries no implementation. | | `hq` | This repository, under the company organisation — mission, research, design, decisions, issue diagnosis. Company-scoped ([ADR 0019](../02-DECISIONS/0019-how-this-repository-works.md)); the mesh is its first product. The source of truth for *why*. Carries no implementation. |
| *(one per application)* | Every standalone application, site or side-project gets its own repository, with `module.yml` at the root. Registered with the mesh as a build source; built and deployed by the same pipeline as anything in the monorepo. | | *(one per application)* | Every standalone application, site or side-project gets its own repository, with `module.yml` at the root. Registered with the mesh as a build source; built and deployed by the same pipeline as anything in the monorepo. |
| `migration` | **Private.** The record of one installation replacing the predecessor mesh with this one: the runbook, a dated log of every step and what it cost, the per-service data procedures, the readiness checks, and the scripts. Private because it is the opposite of this repository in every way that matters — it names machines, addresses, ports and paths, because a procedure that cannot be followed is not one. Where hq asks *what did we decide and why*, that repository answers *what happened on the machines, in what order, and what to do next*. Its `HANDOFF.md` is where somebody picking the work up starts. |
## What the mesh becomes ## What the mesh becomes
[ADR 0019](../02-DECISIONS/0019-how-this-repository-works.md) records the repositories the [ADR 0019](../02-DECISIONS/0019-how-this-repository-works.md) records the repositories the
monorepo decomposes into. **`mesh-lab`, `mesh-host` and `mesh-control` exist so far** — the lab is built first monorepo decomposes into. **`mesh-host`, `mesh-controller`, `mesh-catalog`, `mesh-lab`, `mesh-sdk` and `mesh-tools` exist
([ADR 0016](../02-DECISIONS/0016-the-lab.md)); the rest are the so far** — the lab was built first ([ADR 0016](../02-DECISIONS/0016-the-lab.md)). The tiered
target, not the present. decomposition below is the planned shape; the repositories built to date do not map onto it
one-for-one — `mesh-catalog`, `mesh-sdk` and `mesh-tools` exist where the table names
`mesh-foundation` and `mesh-surfaces`, and reconciling the two is itself still ahead.
| Repository | Tier | Holds | | Repository | Tier | Holds |
|---|---|---| |---|---|---|
| `mesh-host` | 0 | **exists.** The node host — one statically linked binary, requiring nothing present ([ADR 0005](../02-DECISIONS/0005-the-node-host.md)) | | `mesh-host` | 0 | **exists.** The node host — one statically linked binary, requiring nothing present ([ADR 0005](../02-DECISIONS/0005-the-node-host.md)) |
| `mesh-substrate` | 1 | the four pinned services, as declarations | | `mesh-foundation` | 1 | the four pinned services, as declarations |
| `mesh-control` | 2 | **exists.** The control plane and its contexts — one of seven built ([ADR 0006](../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)) | | `mesh-controller` | 2 | **exists.** The controller and its contexts — one of seven built ([ADR 0006](../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)) |
| `mesh-surfaces` | 3 | tools, web, cli | | `mesh-surfaces` | 3 | tools, web, cli |
| `mesh-sdk` | — | contracts shared across tiers | | `mesh-sdk` | — | the stable spine modules build against — the tool-serving harness, the messaging/event framework, the contracts and core primitives. Holds nothing per-module and nothing volatile ([ADR 0039](../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md)). |
| `mesh-lab` | — | **exists.** The lab — scenario lifecycle, networking, placement. Ships to nobody; runs on a workstation. | | `mesh-lab` | — | **exists.** The lab — scenario lifecycle, networking, placement. Ships to nobody; runs on a workstation. |
Tier 4's shape is open, and deliberately so: see ADR 0019 and Tier 4's shape is open, and deliberately so: see ADR 0019 and
@@ -3,7 +3,6 @@ status: graduated
initiated: 2026-08-22 initiated: 2026-08-22
touches: [03-DESIGN/00-as-is/05-runtime-and-installation.md] touches: [03-DESIGN/00-as-is/05-runtime-and-installation.md]
became: became:
- 02-DECISIONS/0005-the-node-host.md
- 02-DECISIONS/0005-the-node-host.md - 02-DECISIONS/0005-the-node-host.md
- 03-DESIGN/01-to-be/05-the-node-host.md - 03-DESIGN/01-to-be/05-the-node-host.md
--- ---
@@ -90,5 +90,5 @@ the catalogue where modules genuinely change together under one intent. The skel
| One repository per tier, or per context? | Already open from ADR 0001 as "catalogue destination — one repository or many". The skeleton assumes per tier and does not settle it. | | One repository per tier, or per context? | Already open from ADR 0001 as "catalogue destination — one repository or many". The skeleton assumes per tier and does not settle it. |
| ~~Does an unprivileged node earn a place in the inventory, or only a presence?~~ | **Answered 2026-08-25** by the operator: a node is a *managed machine inside the mesh*, not an unprivileged something — and a disconnected node is still a node, in a different situation. The question posed a class distinction; the answer is that there is none, and what varies is **state**. Recorded as [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md). | | ~~Does an unprivileged node earn a place in the inventory, or only a presence?~~ | **Answered 2026-08-25** by the operator: a node is a *managed machine inside the mesh*, not an unprivileged something — and a disconnected node is still a node, in a different situation. The question posed a class distinction; the answer is that there is none, and what varies is **state**. Recorded as [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md). |
| ~~Does absorbing overlay, filtering, packages, supervision and the container runtime make the host too large?~~ | **Answered 2026-08-25** — [`host-size.md`](host-size.md). Measured: the absorption is smaller than the machinery that already applies state, and eight of ten adapters already carry no dependency. The risk is not size but direction, and it is two modules wide. The claim survives with its scope corrected — the host carries one concern, *apply declared state on this machine*, of which the six are instances. Recorded as [ADR 0005](../../02-DECISIONS/0005-the-node-host.md), designed in [`05-the-node-host.md`](../../03-DESIGN/01-to-be/05-the-node-host.md). | | ~~Does absorbing overlay, filtering, packages, supervision and the container runtime make the host too large?~~ | **Answered 2026-08-25** — [`host-size.md`](host-size.md). Measured: the absorption is smaller than the machinery that already applies state, and eight of ten adapters already carry no dependency. The risk is not size but direction, and it is two modules wide. The claim survives with its scope corrected — the host carries one concern, *apply declared state on this machine*, of which the six are instances. Recorded as [ADR 0005](../../02-DECISIONS/0005-the-node-host.md), designed in [`05-the-node-host.md`](../../03-DESIGN/01-to-be/05-the-node-host.md). |
| ~~Four substrate services or five?~~ | **Answered conditionally**, which is the honest form — [`07-the-substrate.md`](../../03-DESIGN/01-to-be/07-the-substrate.md). The substrate is *what the control plane consumes and cannot grant itself*. The identity provider qualifies only if the control plane delegates authentication; if it authenticates natively it is an ordinary hosted service. The count follows from a decision not yet taken, and asserting four was asserting that decision. | | ~~Four substrate services or five?~~ | **Answered conditionally**, which is the honest form — [`07-the-foundation.md`](../../03-DESIGN/01-to-be/07-the-foundation.md). The substrate is *what the control plane consumes and cannot grant itself*. The identity provider qualifies only if the control plane delegates authentication; if it authenticates natively it is an ordinary hosted service. The count follows from a decision not yet taken, and asserting four was asserting that decision. |
| Does `feature` survive? | The skeleton splits it in two and argues the conflation is what makes the delivery pipeline hard to reason about. Unproven. | | Does `feature` survive? | The skeleton splits it in two and argues the conflation is what makes the delivery pipeline hard to reason about. Unproven. |
@@ -4,8 +4,8 @@ initiated: 2026-08-25
became: became:
- 02-DECISIONS/0009-modules-and-the-graph.md - 02-DECISIONS/0009-modules-and-the-graph.md
- 02-DECISIONS/0008-a-context-owns-its-store.md - 02-DECISIONS/0008-a-context-owns-its-store.md
- 03-DESIGN/01-to-be/06-the-control-plane.md - 03-DESIGN/01-to-be/06-the-controller.md
- 03-DESIGN/01-to-be/07-the-substrate.md - 03-DESIGN/01-to-be/07-the-foundation.md
touches: touches:
- 02-DECISIONS/0009-modules-and-the-graph.md - 02-DECISIONS/0009-modules-and-the-graph.md
- 02-DECISIONS/0009-modules-and-the-graph.md - 02-DECISIONS/0009-modules-and-the-graph.md
@@ -7,6 +7,9 @@ touches:
- 03-DESIGN/01-to-be/05-the-node-host.md - 03-DESIGN/01-to-be/05-the-node-host.md
- 03-DESIGN/00-as-is/05-runtime-and-installation.md - 03-DESIGN/00-as-is/05-runtime-and-installation.md
- 01-RESEARCH/011-the-module-graph/00-overview.md - 01-RESEARCH/011-the-module-graph/00-overview.md
- 02-DECISIONS/0078-the-store-and-broker-are-modules.md
- 02-DECISIONS/0088-the-foundation-filters-before-anything-listens.md
- 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
--- ---
# 012 — The minimum viable node, and adopting what is already there # 012 — The minimum viable node, and adopting what is already there
@@ -180,6 +183,16 @@ same makes a node where something the mesh needed never happened indistinguishab
where a log level differed. Whether a failed line still lets adoption complete is therefore where a log level differed. Whether a failed line still lets adoption complete is therefore
reopened by adding severity, and is not decided here. reopened by adding severity, and is not decided here.
## Adoption as a mode, for the migration
*2026-09-22.* The migration from the predecessor mesh gave the middle state a length. A machine
running the predecessor is **adopted** when the mesh comes up on it — the predecessor's control
stopped, the machine's firewall and files kept in force, the mesh opening what it needs through
them — and stays adopted while its modules migrate one at a time, until the operator **converges**
it. Measured on the control-node, and weighed against the alternatives, in
[*migrating a node that is in use*](migrating-a-node-in-use.md); decided in
[ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md).
## Open questions ## Open questions
| Question | Why it is open | | Question | Why it is open |
@@ -191,6 +204,7 @@ reopened by adding severity, and is not decided here.
| Can a module say which of its settings are load-bearing? | The question that dissolves the conflict rule rather than choosing a side. A setting the module *requires* cannot be kept from the machine without producing something installed and broken; a setting it merely *prefers* should always yield. Until a module can say which is which, adoption is defaulting in the dark. Belongs with the graph. | | Can a module say which of its settings are load-bearing? | The question that dissolves the conflict rule rather than choosing a side. A setting the module *requires* cannot be kept from the machine without producing something installed and broken; a setting it merely *prefers* should always yield. Until a module can say which is which, adoption is defaulting in the dark. Belongs with the graph. |
| Does a `failed` line still let adoption complete? | *Flags inform, they do not block* was decided about conflicts, where the mesh chose and the machine works. A failure is *we could not*, which is different in kind — and treating them alike hides the worse one behind the commoner one. | | Does a `failed` line still let adoption complete? | *Flags inform, they do not block* was decided about conflicts, where the mesh chose and the machine works. A failure is *we could not*, which is different in kind — and treating them alike hides the worse one behind the commoner one. |
| How is a flagged conflict reconciled, and by whom? | The briefing hands it to a session. What that session is empowered to change, and whether the resolution is recorded so the next adoption does not re-raise it, is undecided. | | How is a flagged conflict reconciled, and by whom? | The briefing hands it to a session. What that session is empowered to change, and whether the resolution is recorded so the next adoption does not re-raise it, is undecided. |
| ~~How long is a machine adopted?~~ | **Decided** ([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)) — for as long as it is being migrated: a mode per node, recorded, ended by an explicit and previewed flip. |
| Where does the kept original live, and for how long? | Whether it is recorded in the node's state so adoption is visibly reversible, and whether it is returned when the mesh stops managing the thing. | | Where does the kept original live, and for how long? | Whether it is recorded in the node's state so adoption is visibly reversible, and whether it is returned when the mesh stops managing the thing. |
| What shape is a briefing? | Structured enough to be acted on, prose enough to be read. It is the first thing a session on a new node sees, which makes it an interface rather than a log. | | What shape is a briefing? | Structured enough to be acted on, prose enough to be read. It is the first thing a session on a new node sees, which makes it an interface rather than a log. |
| Does owning a package mean owning its version? | Owning configuration and owning the package are different scopes. The second means the mesh decides which version is installed, and that decision then has to survive the machine's own package manager updating it. | | Does owning a package mean owning its version? | Owning configuration and owning the package are different scopes. The second means the mesh decides which version is installed, and that decision then has to survive the machine's own package manager updating it. |
@@ -0,0 +1,154 @@
# Migrating a node that is in use — adoption as a mode, not a moment
*2026-09-22. Measured on the machine that will be the control-node, which is running the
predecessor mesh today. Nothing below names it; the counts are its own.*
## The question
The mesh replaces a predecessor mesh that is running, on the same machines, with the services
people use. The control-node is decided: it is the machine that already carries the predecessor's
broker and build pipeline. So the question is not *where* the mesh starts but **how a machine
running the predecessor becomes a node of the mesh without its services noticing** — and then how
each of the other machines follows.
## The operator's proposal
Proposed by the operator, 2026-09-22, and the shape this document tests:
1. **Stop the predecessor's control on a machine** — its daemons that write configuration: the
network and firewall configuration above all, which decide what is reachable and what is
blocked. Its services keep running; only the control over their configuration stops.
2. **Bring the mesh up on that machine in adoption mode.** It takes custody of those files and
keeps what it finds in force.
3. **Migrate the modules one at a time**, data preserved, per the cutover procedure.
4. **Move to the next machine and repeat** — adopted first, keeping its local configuration, then
migrated.
5. **When every machine is migrated, flip adoption mode**, and the mesh takes full control of the
configuration it has been holding.
This is the research above made concrete. It keeps the conflict rule already decided here — *on
conflict, what is on the machine stays* — and gives the middle state, *adopted*, a length: not a
one-time import before generating starts, but a mode that lasts for as long as the machine is
being migrated, ended by an explicit act.
## What the machine actually looks like
Measured, read-only:
| | |
|---|---|
| Containers running | 60, all the predecessor's services and their stores |
| The predecessor's control | user-level daemons, separate from the services; none of the 20 running system services is the predecessor's control |
| Firewall | the predecessor's, active: 54 incoming rules and 52 forwarding rules, each served port allowed explicitly — how a default-deny firewall reads |
| Files the predecessor's configuration sync writes | 12, of which 3 are system files (an ssh server drop-in, the package manager's configuration, one service's configuration); the rest are the operator's shell and agent files |
| Per-service configuration | environment and composition files per service, written by the predecessor's service tooling |
**Stopping the predecessor's control stops nothing that serves.** The services are containers and
system units that run without it; what stops is the rewriting of their configuration. Nothing
changes on the machine until something else writes.
## What collides, measured rather than assumed
An earlier note assumed the mesh's foundation could not stand beside the predecessor because both
want the store's and the broker's standard ports. **The measurement says otherwise.** The
predecessor publishes its own store on a non-standard port and its broker on another; the standard
ports the foundation binds for its store and its bus are free.
What does collide:
| The mesh wants | Held by | When |
|---|---|---|
| the registry's port | the predecessor's registry | at genesis — the foundation raises a registry |
| the broker's management port, on loopback | the predecessor's broker | at genesis |
| the private network's port | the predecessor's own tunnel | at genesis on a control-node that is the private network's hub; otherwise when the node is placed on it |
| the web ports | the predecessor's reverse proxy | when the route proxy is assigned |
| the resolver's port | a resolver the predecessor runs | when a resolver module is assigned |
On the control-node, which is the private network's hub, three of these fall at genesis and cannot
be deferred; two only when a particular module is taken, the moment its predecessor stops anyway.
Today the foundation's ports are **fixed**: written in the installer's bundle and in the
catalogue's manifests, so a collision is found when a container fails to bind, not before — and a
port changed at genesis would be changed back when the foundation is adopted as modules, since the
applier recreates a container whose declared spec differs. The private network's address range
must also stay clear of the range the predecessor's tunnel uses; on the machines measured they are
distinct.
## What would break if the mesh came up as it is today
**The firewall.** Genesis loads a base ruleset — drop anything undeclared — in its own table
([ADR 0088](../../02-DECISIONS/0088-the-foundation-filters-before-anything-listens.md)). The
predecessor's firewall is a different table. The kernel runs every base chain registered at the
same hook, in priority order: an accept ends only its own chain and the packet goes on to the next,
and a drop in any is final — whether the other firewall's chains are nftables or legacy iptables.
So the
mesh's base ruleset would drop everything the predecessor's firewall allows and the mesh has not
declared — every web, mail and database port in the table above — the moment genesis ran.
**The files.** The host writes a declared file whatever it finds at the path, reporting it as
updated. A file the predecessor left — the ssh server drop-in, the resolver's configuration — is
replaced the first time a module declaring that path is assigned, before that module's service
has moved.
**The container names.** The applier keys a container on its name. A catalogue module whose
container carries the same name as the predecessor's service it replaces takes that container over
the moment it is assigned: today, *assigning a module is migrating it*, never a preparation.
**The published ports.** The foundation publishes its ports on every interface, and a published
container port reaches the container through the forwarded path, not the incoming one. A firewall
that filters only incoming traffic never sees it. The base ruleset is what keeps the store
unreachable from outside today — and it is the thing that cannot be loaded on this machine.
## The pipeline freezes while the control-node migrates
The predecessor's build pipeline and its coordinator run on the control-node. Stopping its control
there stops the predecessor's updates for **every** machine it manages, until the migration is
done. Their services keep running; they receive nothing new. That is the price of the proposal and
it is worth stating, not a reason against it: the migration is the period in which the predecessor
is being replaced, and it does not need to keep changing.
## What adoption mode has to mean
For the proposal to hold, *adopted* must be a **state the mesh records per node**, not an
intention, and each thing the mesh would otherwise take must say what it does in that state:
- **What is found is kept until its module is taken.** *Found* is precise: present at a declared
path or name with no record in the host's store. A found file or container is held, its original
recorded, until the operator **takes** the module on that node — the cutover, done when the
module's data has moved. Assigning prepares; taking migrates. Without the distinction the rule
never fires: the host only ever sees what assigned modules declare.
- **The firewall found on the machine stays in force.** The mesh loads no table on an adopted node
that drops by default or holds an accept. What it needs open it declares as openings the host converges
*through the found firewall*, on the incoming and the forwarded path, marked as the mesh's and
re-checked on every reconcile so a reload or reboot does not lose them. An accept in a table of
its own would not help: the found firewall's drop would still be final. What a table of its own
*can* do is refuse, and a refusal is final too — so the mesh guards the store and the broker's
management port from everyone but the private network and the machine itself, in a table that
only refuses, ahead of the container runtime's redirect — which the found firewall does not do
and cannot undo. The bus, the registry and the hub's port stay open to
anywhere: a node enrols before it has a private-network address.
- **The foundation's ports are the node's to give** — set at genesis, checked free, and kept as
that node's settings, read everywhere they are used, so adopting the foundation as modules does
not move them back.
- **The flip is per node and previewed** from what is actually reachable — listening sockets and
published ports, not the found firewall's allow list, which does not see what a container runtime
forwards.
## Options weighed
| Option | Verdict |
|---|---|
| Cut the machine over in one go: stop the predecessor's store, broker, registry and proxy, raise the foundation in their place | Rejected. Every predecessor service goes down until it has migrated, and the predecessor's other machines lose their broker. The rollback is restarting the predecessor, which is a recovery, not a step. |
| A separate machine as the control-node | Rejected. Contradicts the decision that the control-node is the machine that already carries the predecessor's broker and pipeline. |
| Converge on joining, as the mesh does today | Rejected. The base ruleset closes every predecessor port at genesis, and found files are replaced before their services move. |
| Make only the foundation's ports configurable and otherwise converge | Rejected as insufficient. It solves the bind collisions and none of the firewall or file ones. |
| Treat assigning a module as migrating it | Rejected on review. The rule that keeps found files would never fire — the host only sees what assigned modules declare — and an assignment on an adopted node would be an outage rather than a preparation. |
| **Adoption as a mode, per node, ended by an explicit flip** | The proposal. It is the conflict rule already decided here, given a duration. |
## What this leaves open
Answered by the decision record this feeds — [ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)
— only for the migration's needs. The general questions above stay open: whether a module can say
which settings are load-bearing, what shape a briefing takes, how a flagged conflict is reconciled.
One is newly sharp: the mesh opening ports through a firewall it did not install needs to speak that
firewall. There is one kind on the machines measured; a machine with another is not covered until
someone writes for it.
@@ -0,0 +1,54 @@
---
status: active
initiated: 2026-09-23
touches:
- 02-DECISIONS/0075-two-stores-and-which-provides-what.md
- 02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md
- 02-DECISIONS/0071-genesis-builds-from-a-mesh-that-already-exists.md
- 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md
- 04-ISSUES/085-the-packages-port-given-at-genesis-is-not-a-setting/00-report.md
- 04-ISSUES/090-the-forge-module-does-not-take-over-the-forge-genesis-raised/00-report.md
---
# 013 — The forge and the registries: what a seat is for, and what it is not
**The question, asked during the first migration:** the mesh runs a forge that serves git and
packages, an image registry, and a bootstrap forge genesis raises before any module exists. Should
there be a mesh-scoped **git seat**, the way there is one for the store and the broker?
**The short answer: no, and the question points at a different hole.** A seat answers *is there
exactly one of you*. Every symptom around the forge is about **an address and a credential nobody
resolves**, which a seat does not answer. The mechanism that would answer it — a provision, with
`serves` and a grant — already exists, is already used for the image store, and is already
declared for packages. It simply has no consumer: the one module that needs it carries a literal
instead.
See [the survey](the-survey.md) for what the code does today, with evidence.
## What was found
- **A seat does exactly two things**: it refuses a second claimant at resolution, and in one place
it answers *where is the broker*. It is a resolution-time predicate; no host ever hears of it.
- **Four mesh-scoped seats exist**, all named after a singular server the mesh runs **for its own
working** — controller, store, broker, catalogue. A forge is an application the world made, not
part of that set.
- **Nothing in the mesh requires git.** There is no provision for it, and the forge's git service —
over https and ssh — is declared nowhere. Only its npm half is a provision.
- **`package-registry` is a fully built provision with zero consumers.** The builder, its only
real consumer, bypasses it with a file naming the module, the address `127.0.0.1` and the port.
- **The builder is told where source lives, per build, by a person.** The clone URL is an opaque
string; each module remembers its own; no credential is ever attached; nothing polls a forge.
- **A seat would forbid something the mesh should allow**: a second forge — one for the mesh's own
source, one for something else — which ADR 0075 already argues for against a node-scoped claim.
## What this leaves open
- Should the forge's git service become a **provision** (`source-forge`), so the builder resolves
the address and a credential the way it already resolves the image store? That is what would let
a forge move, be renamed, or be replaced without editing every module's stored URL, and it is the
unanswered half of [issue 085](../../04-ISSUES/085-the-packages-port-given-at-genesis-is-not-a-setting/00-report.md).
- The obstacle is known and written down: such a requirement **may go unanswered during genesis**,
and the mesh has no optional requirement. Deciding that is the real work.
- Should the mesh learn when a source moves? It records the fact and never discovers it: nothing
polls, and there is no receiver for a forge's push. `build --behind` answers a question only a
person can currently make true.
@@ -0,0 +1,130 @@
# The survey — seats, the forge, the registries, as the code has them
*2026-09-23. Read from the code and the records, not from memory. Every claim here was checked
against a file.*
## What a seat is
The glossary calls a seat *a named position at a scope with a capacity*, and a claim *a module
taking a spot on it*. **Capacity does not exist in the code.** There is no capacity field and no
bench; every claim is exclusive. The "shared seat" of the design is the word `provides` wearing
that name.
Claiming does two things and no others:
- **It refuses a second claimant when a node's modules are resolved** — within one node at any
scope, and against every other node for mesh and site scope.
- **It answers one question, once:** the controller finds where the broker is by looking for the
module claiming the broker's seat.
No host is ever told about a seat. It is a resolution-time predicate.
Two mechanical asymmetries matter. A mesh-scoped seat's exclusivity is defended only by nodes
that *resolve*: a node whose modules fail to resolve contributes nothing, so a broken node does
not hold its seat against a second claimant. And a site-scoped seat does nothing at all when
either node has no site.
## The seats that exist
Twelve claims, twelve names. Four are mesh-scoped: the controller, the store, the broker, the
catalogue. All four are things the mesh runs **for its own working**, and three of them are raised
by genesis and adopted in place — which is the argument of the record that named them.
The node-scoped ones split in two: a **role on this machine** (the build machine, the packet
filter, the intrusion prevention, the private network) and **a scarce machine resource** (the
resolver's configuration file, the DNS port). The second kind is a claim for the reason a port is:
there is one of it on the machine.
Two oddities worth stating:
- The image store claims its seat at **node** scope while providing its service at **mesh** scope.
The record that discussed this left open whether it should hold that claim at all.
- **Nothing claims the two public ports.** The proxy binds them and claims nothing, which is the
hole the route handover fell into.
## The forge, and what it serves
The forge serves three things: git over https, git over ssh on its own port, and an npm registry.
**Only the npm half is in the mesh's vocabulary** — the forge provides `package-registry` and says
where it answers. Git is served and declared nowhere: no provision, no `serves`, nothing that can
require it. The only trace is the public label in its route contribution, which the mesh is
explicitly not meant to interpret.
The image store is a separate module and a separate provision, plain HTTP, trusted because it is
reachable only over the private network.
A second module also provides `package-registry`. Two providers of one provision is not a refusal
but an ambiguity, and the mesh already has the command that settles it: pin one.
## Where the builder's three addresses come from
| What it needs | How it finds it |
|---|---|
| the broker | a sealed secret of its own |
| the image store | **a real provision binding**, resolved by the mesh |
| the package registry | **a file in its own manifest**, naming the module, `127.0.0.1` and the port |
| the source repository | nothing at all — a person types a clone URL per build |
The installer says this plainly in a comment: *the one binding nothing resolves: the builder dials
the forge by a number it carries.* Genesis can override the port of that literal, and nothing
else: not the forge's identity, not its address.
`package-registry` therefore has **zero consumers**. The builder does not require it. The
provision is fully built — two providers, grants, a served description including the registry's
path — and nothing asks for it.
## How a module's source is found
It is not found; it is supplied. The clone URL arrives as an argument, travels to the build machine
and reaches `git clone` unexamined. No credential is attached, so a private forge works only if the
build machine's own git configuration already authenticates. Each module remembers the URL it was
built from, so there are as many forge addresses as modules, each frozen at whatever was typed.
At genesis there are four independent source URLs, one per flag, unrelated to each other.
If the forge moves or is renamed: every stored URL goes stale independently, each failing at its
own clone, and there is no command to re-point them. Genesis keeps working, because it takes URLs
as flags — which is the existing record's position: the forge is reached *by a name outside the
mesh*, and failover is repointing that name.
## What genesis's forge is
A bare container run before any module exists, because the toolchain resolves the mesh's own
packages by version from an npm registry. It holds **no seat, no manifest, no provision**. The
mesh is told one thing about it — a port — and only when that port is not the default. **On an
ordinary genesis the mesh never learns the bootstrap forge exists.**
Its successor module differs from it in five ways: container name, network, data directory, the
address it is told to call itself, and — contradicting a comment that says they are pinned
identically — **the image digest**. So assigning the module raises a second forge beside the first.
A seat cannot fix that. A seat is checked between manifests; the bootstrap forge has no manifest,
so no seat can see it. The check that is actually wanted — same name, same data, same network — is
between an installer constant and a manifest.
## Would a git seat be coherent?
In form, yes: add the claim and a second forge is refused. In effect it buys one refusal nobody
has hit, and forbids an arrangement the mesh should allow — a second forge for something other
than the mesh's own source.
Measured against the four mesh seats, it does not fit: those are singular servers the mesh runs
for itself, and a forge is an application. Nothing would *use* it, because the only seat consumer
answers "where is the broker", and a forge address is not one fact: a module's source is a
repository **and a path and a ref**.
## What the code already supports and nobody uses
- **`package-registry` with no consumers.** The builder's literal file is a hand-rolled copy of
the binding the mesh would have written for it.
- **Pinning a provider** already expresses *this* forge, of several, without forbidding the second.
- **Grants** on the forge are the credential mechanism that would hand the builder an account.
Today the installer creates that account directly, because nothing requires the provision and so
the grant has no consumer to write for.
## The reading
The question *should there be a git seat* is the wrong shape for the problem under it. Every
symptom — the literal address, the port that would not follow a node, the takeover that is not a
takeover — is about **an address and a credential nobody resolves**. A seat resolves neither. The
provision mechanism does, and is already built.
@@ -0,0 +1,108 @@
---
status: graduated
became: 02-DECISIONS/0106-the-bus-is-nats.md
initiated: 2026-09-23
touches:
- 02-DECISIONS/0002-nodes-communicate-over-a-broker.md
- 02-DECISIONS/0033-the-substrate-is-a-store-and-a-broker.md
- 02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md
- 02-DECISIONS/0041-events-are-a-relationship.md
- 02-DECISIONS/0042-the-shape-of-an-event-on-the-wire.md
- 02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md
- 02-DECISIONS/0078-the-store-and-broker-are-modules.md
- 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
---
# 014 — The bus on NATS: replace the broker, and when
**The question, asked mid-migration:** the operator wants the mesh's bus — today an AMQP broker —
replaced by NATS, with everything the broker does today. Do we finish the migration on AMQP and
move to NATS after, or go to NATS directly?
**The short answer: decide NATS now, build it in the lab in parallel, and cut the mesh's bus over
in one rehearsed rollout — after the migration's core is done and never underneath it.** "First
everything on AMQP" is already the state and costs nothing more: the predecessor's broker was
merged into the mesh's tonight, and every module converted from here on targets the sdk's broker
contract, which names no protocol. The only code that speaks AMQP is the mesh's own, in three
places, and it is swapped once.
## What was measured
**Where AMQP is spoken** (non-test files): the controller's `internal/link` (5 files, ~2.4k lines
with tests) and the builder's main; the host's `internal/link` (~1.5k); the tool runtime's
`broker-amqp.ts` (~600 with its main). **The sdk speaks none** — its `Broker` is `request`,
`handle`, `publish`, `subscribe`, `close` ([ADR 0039](../../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md)
put the client in the runtime for exactly this). **Every module's tools and events go through the
sdk**; one module (`amqp-ping`, a probe) talks AMQP on purpose.
**What the bus carries:** a control queue (`control`, `.upgrades`, `.catchup`), one queue per node
(`node.<name>`), `builds`, the enrolment/report/alive/built flows, a topic exchange of events
(`mesh.events.<module>.<event>`, dead-lettered to `mesh.events.dead`), an RPC exchange (`mesh.rpc`)
and per-tool service queues (`serve.<module>.<tool>`), plus an MQTT exchange.
**What it relies on:** TLS for the bus a node enrols over; prefetch with reject/nack so the control
queue **holds messages unacknowledged while the store restarts** and retries
([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)); publisher confirms and
mandatory routing; a dead-letter exchange; **per-module accounts scoped by `emits`/`consumes`**
([ADR 0043](../../02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md)),
minted through the broker's management API; the broker as a module holding the `mesh-broker` seat
([ADR 0078](../../02-DECISIONS/0078-the-store-and-broker-are-modules.md)).
**The predecessor's world**, now on the same broker: ~50 AMQP connections per machine from four
machines, 595 queues, four vhosts, six users. All of it AMQP, all of it retiring module by module.
NATS speaks no AMQP: that world cannot move; it does not need to.
## How NATS answers each of those
| The mesh relies on | NATS | Note |
|---|---|---|
| topic routing keys | subjects with wildcards | same shape (`mesh.events.>`) |
| RPC and tool invocation over an exchange + reply queue | request/reply, native | simpler than today |
| competing consumers | queue groups | same |
| durable control queue, hold-unacked-and-retry, catch-up | JetStream: streams, durable consumers, ack/nak with delay, replay | the 0083 guarantee moves to JetStream; core NATS alone is at-most-once and would not do |
| dead-letter | max-deliver + advisories, or a stream fed from them | different mechanism, same effect |
| TLS bus | TLS | same |
| accounts scoped by emits/consumes, vhosts | accounts (isolation) with users and per-subject publish/subscribe permissions | stronger than today; the controller writes an auth config the host declares and the server reloads, instead of calling a management API |
| management API | `nats` CLI and an HTTP monitoring endpoint | no vhost concept — accounts instead |
| MQTT | built in | same |
| a module seat `mesh-broker` | unchanged — the seat is the server, the module changes | ADR 0079 |
| multi-node | clusters and leaf nodes | not needed now; a leaf per node is a later question |
Nothing the mesh needs is missing. The differences are in the shape of durability (JetStream must
be declared, streams and consumers are objects) and of accounts (configuration, not API calls).
## The cost, honestly
The bus is the mesh's nervous system. Moving it means: a new `nats` module in the catalogue taking
the `mesh-broker` seat; the controller's and the host's link packages rewritten; the tool runtime's
client swapped behind the unchanged contract; enrolment, reports, builds and the guard's ports
re-derived; the lab beds that prove the bus (store window, enrolment, upgrades) re-run on the new
one; and eight decisions amended or superseded. Weeks, not days — and none of it can be done
halfway on a live mesh: the controller, every host and every tool runtime move together.
## The sequencing question, answered
1. **NATS first, directly.** Stalls the migration for the length of the build; the mesh's bus
changes under a half-migrated node; the predecessor's clients still need AMQP, so a second
broker runs anyway, plus a bridge for whatever crosses. Rejected.
2. **Finish on AMQP, NATS after.** Wastes nothing — no module written from here on speaks AMQP —
but leaves the decision unmade while modules are written, and the runtimes idle. Adequate.
3. **Decide NATS now; build it in the lab in parallel; cut the mesh's bus over in one rehearsed
rollout after the core is migrated.** The predecessor's clients never notice: their broker is
the one the mesh adopted, kept as a compatibility module with an end date — the day the last
AMQP client is gone. **Recommended.**
## What this leaves open
- The **decision itself**, as a record: the bus is NATS; the AMQP broker becomes the predecessor's
compatibility broker and retires with the last AMQP client. Written when the operator says so.
- **JetStream's shape for the control plane**: one stream per concern (control, nodes, builds,
events) or one with subjects; retention; what the store window guarantee looks like as ack-wait
and nak-delay. Measured in the lab, not designed on paper.
- **Accounts as configuration**: the controller writes users and permissions into a file the host
declares, reloaded on change — which is the [ADR 0102](../../04-ISSUES/102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md)
discipline applied from the start — or the JWT/operator model. The first is simpler and matches
how the mesh already writes everything.
- **The MCP bridge** the operator asked for the same day is written against the sdk's contract, so
it moves with the bus and is not written twice.
- Whether a **leaf node per machine** replaces the hub-and-spoke bus later — out of scope here.
@@ -0,0 +1,142 @@
---
status: active
initiated: 2026-09-24
touches:
- 02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md
- 02-DECISIONS/0033-the-substrate-is-a-store-and-a-broker.md
- 02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md
- 02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md
- 02-DECISIONS/0078-the-store-and-broker-are-modules.md
- 02-DECISIONS/0084-which-provider-serves-a-consumer.md
- 03-DESIGN/00-as-is/03-provisioning.md
- 03-DESIGN/01-to-be/07-the-foundation.md
- 04-ISSUES/113-the-object-stores-images-were-withdrawn-upstream/00-report.md
---
# 015 — The object store after MinIO: which S3 implementation, and how the data moves
**The question.** The mesh's object store is MinIO. Its community edition is archived upstream,
its server and client images have been deleted from every public registry, and the pinned release
is four and a half years old and will never be patched
([issue 113](../../04-ISSUES/113-the-object-stores-images-were-withdrawn-upstream/00-report.md)).
Which S3-compatible implementation replaces it, and what is the migration track for the data and
the provisioning model that sit on top of it?
**Why now, and why not sooner.** Nothing is on fire: nodes that already hold the images keep
running, and issue 113 establishes that the deploy path tolerates an unfetchable-but-present
image by design. The forcing function is not an outage but a one-way door — **no node that does
not already hold the images can ever provision the module again**, so the mesh's ability to stand
a node up from its declarations is already broken for this module, and silently.
**The direction is not a departure from the design; it is the design.** The foundation document
already states the commitment:
> The dependency is on the **protocol**, not the product: AMQP for the bus, S3 for the object
> store, the OCI protocol for the registry. That is what keeps the naming safe rather than a
> commitment that cannot be revisited.
The object store is also **not** a foundation service — ADR 0028 removed it, and it is an
ordinary module required through the module graph by whatever wants one. (The "exception that is
not a swap" in that passage is the relational store, whose provisioning model borrows PostgreSQL's
own meaning of databases, roles and schemas. The object store carries no such coupling: a bucket
is a bucket.) So this effort is an instantiation of an existing principle, not a redesign — which
is the cheapest kind of decision to make and the strongest kind to cite.
## What the replacement has to carry, measured
Taken from the module's manifest, its composition, its tool surface, and a search for its
consumers across the catalogue — not from assumption.
| Requirement | Evidence in the module today |
|---|---|
| S3 API | The protocol every consumer speaks; already the design's stated dependency. |
| ~~OIDC login against the mesh's identity provider~~ | **Struck 2026-09-24. Not a requirement, and it never worked.** Six variables are wired and an entrypoint blocks on the provider, which reads as a live feature. The module's own hook comment records the end state as *"policy claim missing"* — a failing login. See [01](01-candidate-comparison.md). |
| **Per-application access keys, each scoped to a bucket** | The real requirement. A "user" of the store is normally an application; the mesh already mints a credential per provisioned bucket. |
| **One live consumer using it as opaque primary storage** | A file-sync application, since early 2023: objects named by internal id, metadata in its own database. Highest-risk consumer — a live copy drifts, and its bucket name must be preserved. |
| Erasure-coded multi-node topology | Four server nodes with two data directories each, behind a load balancer. |
| A single-node form | Declared as a flavour, for development and small nodes. |
| Buckets as a typed provision | The module declares a provision type of `bucket` on a named network; the mesh mints the credential and the provider creates it (ADRs 0048, 0084). |
| A tool surface | Bucket create/list/delete, object list/info/delete, presigned URL, and provisioning. |
| A console | Published on its own subdomain through the reverse proxy, with an unlimited request-body middleware for uploads. |
**Consumers, counted:** one application module, one capture module that takes a private bucket per
node, one workflow module's tools, and the delivery/rescue internals of the shared library. The
surface is small — the cost is concentrated in the provisioning handler, the tool handlers and the
OIDC story, not spread across the catalogue.
## Candidates
**Four candidates, not three.** The comparison was briefly narrowed to SeaweedFS on the strength of
console single sign-on; that axis turned out not to be a requirement, and the incumbent's own
maintained fork had been omitted altogether. Both errors, and why they happened, are recorded in
[01 — the candidates measured](01-candidate-comparison.md), which carries the evidence and the
requirement-by-requirement detail.
In short, and only in short:
- **The maintained fork of the incumbent** — the community edition was archived and its images
deleted, but a fork publishes, tracks CVEs, and preserves the on-disk format, S3 API and
environment surface. Costs **an image reference** where every other option costs a data
migration, two rewrites and a maintenance window. Does not end the dependence on an abandoned
codebase; buys time to choose deliberately.
- **Garage** — its permission model *is* the requirement (per access key, per bucket), its admin
API is the closest match to how the mesh provisions, and the highest-risk consumer is
first-party documented against it. Remaining cost: no object versioning, no server-side
encryption or object locking, partial lifecycle — **unmeasured against the ten buckets, and the
one thing that could still disqualify it**.
- **SeaweedFS** — longest field record and erasure coding. Its console sign-on is a paid feature,
which is now beside the point. What weighs against it is narrower: its S3 surface is a gateway
translating onto its own file-system API, with no first-party support for the opaque consumer.
- **RustFS** — closest in shape to the incumbent, so the least porting. But it reached general
availability eight days before this was written, and carries an open defect in the credential
path. Two earlier claims about it are corrected in 01: it is **not** a drop-in that retains
existing data.
- **Ceph RGW** — remains rejected as disproportionate where the object store is an ordinary
module rather than a platform.
**This is now two decisions, not one:** whether to repoint to the fork or migrate, and — if
migrating — to which. Repointing does not foreclose migrating, which is the argument for taking it
first. On the corrected requirement the migration ranking is Garage, then SeaweedFS, and not yet
RustFS. Two measurements gate any graduation: **which S3 endpoints the consumers actually call**
(Garage cannot be ranked fairly until counted), and **whether the fork can read the incumbent's
on-disk format in place** — tested on a copy, because the migration between them is one-way. Both
are in [01](01-candidate-comparison.md#what-is-still-unmeasured).
## The migration track, in outline
Data movement is the easy half, and deliberately reversible.
1. **Stand the replacement up beside the incumbent**, on its own ports, its own provision type and
**its own data directory**. Nothing removed. The data directory matters: reusing one the
incumbent already holds would put a fresh single-drive store on top of a live erasure set.
2. **Copy bucket by bucket with a neutral tool.** `rclone` rather than the incumbent's own client
— the client has been withdrawn upstream too, so building the migration on it would inherit
the same dependency this effort exists to remove.
3. **Verify per bucket** — object counts and checksums, not a transfer exit code.
4. **Repoint consumers through the connection the module already publishes.** Consumers read an
API URL from the module's declared connections rather than addressing the store directly, so
the cutover surface is that value plus the provisioning and tool handlers.
5. **Freeze writes, final incremental sync, flip**, and keep the incumbent read-only as the
rollback until confidence is earned. For the opaque consumer this is **not optional and not
instant**: it stores objects by internal id with metadata in its own database, so a copy taken
while it runs will drift. It needs a maintenance window for the final sync, and the window is
proportional to 82,496 objects rather than to 230 GiB.
6. **Retire**, and only then remove the module.
The genuinely new work is not the copy. It is the **provisioning handler** and the **tool
handlers**, both written against the incumbent's admin API. *The OIDC wiring was previously listed
here and is struck: it is not a requirement and it never worked.*
## Open questions
- ~~How much of the OIDC requirement survives, and in which build?~~ **Answered, and it was the
wrong question.** The console requirement does not exist, and the login it referred to never
worked. What replaced it: which S3 endpoints consumers actually call, and whether the fork reads
the incumbent's format in place.
- Does the mesh's bucket provision translate to the candidate's identity model without weakening
what ADR 0049 says about a consumer's identity fitting the tightest backend?
- Should this effort also answer issue 113's general question — mirroring third-party images into
the mesh's own registry — or is that a separate decision? Replacing one withdrawn product with
another unmirrored upstream leaves the same one-way door in place, just further from the hinge.
- Is the four-node erasure-coded topology still warranted, or was it inherited? Worth re-asking
while the product is being chosen, rather than reproducing a shape by default.
@@ -0,0 +1,176 @@
# 015 / 01 — The candidates measured
*Rewritten 2026-09-24. An earlier version of this document ranked the candidates on whether they
preserved single-sign-on to the object store's **console**. That was the wrong axis — it is not a
requirement — and a fourth candidate was missing entirely. Both errors are recorded at the end,
because how a comparison came to be ranked on the wrong thing is worth more than the ranking was.*
## The requirement, corrected
Taken from the operator and from the running system, not from the module's shape.
**A "user" of the object store is normally an application.** The requirement is therefore
**per-application access keys, each scoped to its own bucket** — not per-human single sign-on. The
mesh already works this way: it mints a credential for every provisioned bucket, and the consumer
reads an endpoint from the module's declared connection rather than addressing the store directly.
**The console is not a requirement.** It was the axis the previous version ranked on, and it should
not have been.
**The identity-provider login never worked.** The predecessor's module wires six OIDC variables and
blocks startup until the provider answers, which reads like a working feature. It is not: the
module's own hook comment records the end state as *"policy claim missing"* — a **failing** login,
written up as progress because it proved the provider had registered. The identity provider emits no
such claim, nothing in the module creates the mapper, and the configured scope alone would not carry
a custom one. Two days of logs show no genuine login attempts, only internet scanners failing on an
STS API version. **Nothing should be carried forward on the assumption this works**, and no
candidate should be credited or penalised for matching it.
**One consumer is live, opaque, and holds real user files.** A file-sync application has used the
store as its **primary storage** since early 2023: objects named by an internal id, with all
metadata in its own database. Three consequences — a copy taken while it runs will drift, its bucket
name must be preserved or its database references break, and it is the highest-risk consumer of the
lot.
## What is actually stored, measured
| | |
|---|---|
| Logical | **230 GiB, 82,496 objects, 10 buckets** |
| Raw on disk | **468 GiB** — eight drive directories at 59 GiB each |
| Implied scheme | 468 ÷ 230 = **2.03×**, confirming erasure coding at half parity |
| Headroom | ~1.3 TiB free on the filesystem holding it |
**All eight "drives" are directories on one filesystem on one machine.** The erasure coding is
therefore not buying independent-drive redundancy; the real failure domain is the array underneath,
which has its own. This single fact decides more of the comparison than any product feature: a
scheme's redundancy model is close to irrelevant here, and what remains is its storage overhead.
At 230 GiB with 1.3 TiB free, **storage overhead is not a deciding cost either.** Replication at
three copies would run ~690 GiB against the present 468 GiB — about **+222 GiB**, comfortably
absorbed. Erasure coding at a wider stripe would *save* roughly 146 GiB. Both are rounding errors
against the headroom, and neither should decide this.
## The candidates
Four, not three. The previous version omitted the first.
### The maintained fork of the incumbent
The community edition was archived upstream and its images deleted
([issue 113](../../04-ISSUES/113-the-object-stores-images-were-withdrawn-upstream/00-report.md)),
but **a fork is maintained and publishing** — `pgsty/minio`, from the Pigsty project. It restores
the console stripped from the community
build, rebuilt image and package distribution, tracks CVEs, and states that it preserves the on-disk
format, the S3 API and the environment-variable surface. Verified by pulling it: it reports a
current release, permissive-to-copyleft licensing unchanged from upstream, and identifies itself as
a community fork. Adoption is real — the server image has been pulled three quarters of a million
times.
**Why it reorders the comparison.** Every other candidate costs a data migration, a provisioning
handler rewritten against a different admin API, a tool surface ported, and a maintenance window for
the opaque consumer. The fork costs **an image reference**. It also closes the issue's one-way door:
a node holding nothing can provision the module again, and patches resume.
**What it does not do** is end the dependence on a codebase its original authors abandoned. It is
maintenance mode, largely one project's effort, with no new features intended. It buys time to
choose deliberately rather than under pressure — which is worth a great deal, and is not the same as
a decision.
### Garage
**The best fit for how the mesh provisions.** Its permission model is *per access key, per bucket,
read/write/owner* — which is the requirement above stated verbatim rather than approximated. Its
admin API is a first-class REST surface with tokens scopeable to exactly the two operations a bucket
provision performs. The opaque consumer is **first-party documented** against it, for primary
storage, including client-side encryption support.
Its previously-recorded penalties mostly dissolve under the corrected requirement: it has no console
and no identity-provider integration, neither of which is wanted; and it replaces AWS-style ACLs and
bucket policies with its own per-key-per-bucket model, which is the thing being asked for.
**What genuinely remains.** It replicates rather than erasure-codes — immaterial at this volume and
on a single array, as above. It does **not implement the full span of S3 endpoints**: object
versioning is absent, object locking and server-side encryption endpoints are absent, and lifecycle
is partial. **Whether any of the ten buckets depends on those is unmeasured, and it is the one thing
that could still disqualify it.**
### SeaweedFS
Longest field record of the group, permissive licence, erasure coding, and identity-provider
integration on the S3 API through token exchange. **Its console sign-on is a paid feature** — the
admin UI itself is open, its identity integration is not. That finding is what falsified the
previous version's narrowing, and it is now largely beside the point, since the console is not a
requirement.
What weighs against it here is narrower and more specific: its S3 surface is a **gateway
translating onto its own file-system API**, with acknowledged divergence from AWS behaviour at the
edges, and there is no first-party documentation for the opaque consumer. For a store already
holding real user files in an opaque layout, first-party support is worth more than a feature list.
It also carries more moving parts than a single-machine deployment needs.
### RustFS
Closest in shape to the incumbent — a similar admin API and client compatibility, so the existing
handlers would port with least effort — under a permissive licence, with erasure coding and a
console that does integrate an identity provider.
Two things were recorded about it earlier that were **wrong, and are corrected here**: it is *not* a
binary-level drop-in that retains existing data (API compatibility and on-disk compatibility are
separate paths, and the on-disk one is preview-scoped with documented encryption limits), and it
therefore offers no shortcut around the migration. It also carries **an open defect in the exact
area the mesh depends on** — an access key created by an identity-provider user reported denied on
all S3 operations.
Decisively for now: **it reached general availability eight days before this was written.** For a
component holding 230 GiB of real user files, field record is a feature, and it does not have one.
### Ceph RGW
Remains rejected, for the reason already recorded: disproportionate where the object store is an
ordinary module rather than a platform.
## Where this leaves it
**The decision is no longer "which product replaces the incumbent".** It is two decisions, and they
can be taken in either order but should not be confused:
1. **Repoint to the maintained fork, or migrate now?** Repointing is an image reference and it
closes the issue. Migrating now costs a data copy, two rewrites and a maintenance window, and
buys independence from an abandoned codebase sooner.
2. **If migrating, which?** On the corrected requirement the ranking is **Garage first** — its
permission model *is* the requirement, its provisioning API is the closest match, and the
highest-risk consumer is first-party supported. SeaweedFS second, on field record, with a
translation-layer caveat that matters more here than its feature list. RustFS not yet, on age.
Taking (1) does not foreclose (2), and that asymmetry is the argument for taking (1) first.
## What is still unmeasured
1. **Whether any of the ten buckets needs object versioning, server-side encryption or lifecycle.**
This gates Garage specifically and nothing else here answers it.
2. **Whether the fork's release can actually read the incumbent's on-disk format in place.** The
format is claimed compatible across a multi-year gap; the migration between them is one-way, so
this is tested on a copy or not at all.
3. **Whether the opaque consumer's maintenance window is acceptable**, and how long it actually is
at 82,496 objects.
4. **Whether the eight-drive erasure-coded shape is warranted at all.** The evidence above says it
is not buying what it appears to: eight directories, one array, one machine. It looks inherited.
**Nothing graduates to a decision before 1 and 2.**
## Two errors in the previous version of this document
Recorded because the shape of both survives anonymisation and neither is unique to this effort.
**It ranked on a requirement that did not exist.** Console single sign-on was treated as the axis
because the module's configuration showed it wired up, and a wired-up configuration was read as a
used feature. It was neither used nor working. *A configured feature is not an observed one*, and
the evidence needed was the operator's answer and the logs — both cheap, neither consulted before
the ranking was written.
**It omitted the incumbent's own fork.** The whole effort began because an upstream withdrew its
images; whether anyone had continued that upstream was the first question to ask and it was not
asked. The candidate list was assembled from a search for *alternatives*, which by construction
returns things that are not the incumbent. *When a dependency dies, "who took it over" precedes
"what replaces it".*
@@ -0,0 +1,57 @@
---
status: graduated
became:
- 02-DECISIONS/0114-a-shared-credential-rotates-over-two-credentials.md
- 03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md
initiated: 2026-09-26
touches:
- 02-DECISIONS/0113-the-vault-makes-every-secret.md
- 02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md
- 02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md
- 03-DESIGN/01-to-be/13-credentials-and-their-rotation.md
- 03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md
- 04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md
---
# 016 — How a credential can be rotated
**What.** Which rotation mechanisms the mesh's providers can actually support, measured against
their code rather than assumed. Every provider in the catalogue was read, found by listing every definition that provides something: how it names what it
makes for a consumer, what its remove destroys, whether it re-applies a password, whether its
backend can hold two secrets for one login or two logins on one resource, and how its own
administrative credential is set. The consumer side was read too: when a module reads a secret, and
what makes it read a new one.
**Why.** [ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md), as first drafted,
chose *overlap*: add a second login beside the first, move every reader, then remove the old one,
"through the adapter's existing create and remove", with "no consumer changes". A review showed that
claim false. In most providers the consumer's data is named after its login, and remove drops the data
with the login. Overlap as written would have deleted every consumer's database on its first
rotation. The mechanism has to be chosen on what the providers do.
**What it touches.** Rotation in 0113 and [to-be 27](../../03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md),
which [ADR 0114](../../02-DECISIONS/0114-a-shared-credential-rotates-over-two-credentials.md) decided on
these findings. The identity budget in
[ADR 0049](../../02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md), if a consumer
gets two logins. The rotation already implemented, which [to-be 13](../../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md)
describes.
**Documents.**
- [01 — The providers](01-the-providers.md): the survey, one row per provider, and what it shows.
- [02 — The readers](02-the-readers.md): how a secret reaches a running process, and what already
recreates it.
- [03 — The options](03-the-options.md): each rotation mechanism against those facts, and a
recommendation.
**Finding, in one paragraph.** All nine credential providers already re-apply a consumer's password
in place on every create, and the controller's `rotate` command relies on that. It is a working
rotation with a stated window. Eight of the nine name the consumer's resource after its login, and five
destroy the consumer's data when they remove the login. The harness, keyed by login, would do the same
on any change of login. Only one backend holds two passwords on one login, and two more hold several
tokens. Eight backends can grant two logins the same rights over one resource; the ninth can give one
login a second token. So every provider can hold **two credentials** over one resource, but only after
each adapter separates *the consumer's resource* from *the credential that reaches it*. In postgres
that also means the resource belongs to a role no login owns. Administrative credentials are a
different case. They have one party and a fixed name, and five backends take them only at first
initialisation, so changing one needs the old and the new value at once.
@@ -0,0 +1,103 @@
# 01 — The providers
Read from the catalogue's main branch: each provider's provisioner adapter (`create`, `remove`),
the client functions they call, and each definition's own credentials. The providers were found by
listing every definition that provides something and has a provisioner, not from memory. A first pass
of this survey worked from memory and missed one, mailu.
**The provisioner harness** in `mesh-sdk` calls `create` for a consumer when its contribution appears
or changes (its login, password or values), and after the provisioner restarts. It calls `remove` for
a login it applied earlier in the same process that is no longer contributed. Its record of what was
applied is kept in memory and keyed by login. Two things follow:
- a consumer whose derived login changes is removed under the old login and created under the new one,
in one pass;
- a contribution that disappears while the provisioner is down is never removed, and is left behind.
## The credential providers
`login` is the consumer's derived identity, which the adapter receives as `as`
([ADR 0049](../../02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md)).
| provider | the consumer's resource is named | remove destroys | create re-applies the password | two secrets on one login | two logins on one resource |
|---|---|---|---|---|---|
| postgres | a database named `login`, owned by the role `login` | the database and the role | yes, `ALTER ROLE … PASSWORD` when the role exists | no: a role has one password | yes, but only through a role that cannot log in owning the database, with each login working as it. Otherwise whatever one login creates is its own, and dropping that login means handing its objects over first. Not done today |
| mssql | a database named `login`, with the login mapped into it | the database and the login | yes, `ALTER LOGIN … WITH PASSWORD` | no: a login has one password | yes, two logins mapped to users in `db_owner`. A user owning a schema cannot be dropped, and a login with an open session cannot. Not done today |
| mongodb | a database named `login`, with a user holding `dbOwner` | the database and the user | yes, `updateUser` with the new password | no: a user has one credential | yes, two users with `dbOwner` on one database. Not done today |
| redis | the key prefix `login:` on an ACL user named `login` | the user, **not** its keys | yes: `ACL SETUSER … reset … >password` replaces all of them | **yes**: an ACL user holds several passwords, added with `>` and removed with `<`. Today's `reset` discards all but the new one | yes, two users on one key prefix, once the prefix is not the login |
| minio | a bucket derived from `login`, and a service account whose access key is `login` | the access key; the bucket **only if empty**. A bucket holding objects is left, and the failure logged | yes, by removing the access key and adding it again, which leaves a moment with no key | no, but an access key *is* the login: a second key is a second login | yes, two service accounts with one bucket policy. The access key is capped at 20 characters |
| lavinmq | a virtual host named `login`, and a user named `login` with permissions on it | the virtual host, with any queued messages, and the user | yes, the user is written again with the password | no: a user has one password | yes, permissions for two users on one virtual host |
| mosquitto | a client named `login`, with a role named for it on the topic prefix `login/#` | the client and its role | yes, the password is set when the client exists | no: a client has one password | yes, two clients holding one role, once the prefix is not the login. The MQTT client identifier is chosen by the consumer, not tied to the login; a duplicate one takes the older session over |
| mailu | a mailbox `login@domain`, unless the consumer contributes its own account name | the mailbox with its mail, for a login-named one; a contributed name is left for an operator | yes, the password is set when the user exists | no for the password; a user can hold several authentication tokens, per the backend's documentation | **no**: a mail user *is* its mailbox |
| gitea (npm) | a user named `login` on a team of an organisation that owns every package | the user; **packages survive**, because the organisation owns them | yes, the user's password is set on every run | no for the password; a user can hold several access tokens | yes, trivially: a second member of the same team |
## The other providers
| provider | answers with | credential |
|---|---|---|
| umami | a website, found by its public name | none. The site id it makes has no way back to the consumer today |
| cloudflare-dns | a public name derived from `login` | none handed to the consumer; its own API token is an operator value |
| showcase | a route | none |
| mesh-vault | custody: it records and withdraws sealed values in a ledger | it holds secrets; it makes none today |
verdaccio provides the npm registry too, and has no provisioner.
## Each provider's own administrative credential
| provider | identity | how the backend takes it |
|---|---|---|
| postgres | a fixed superuser | from a file **only at first initialisation** |
| mssql | `sa` | from the environment at first setup. The image documents no file form, and the definition records that as a declared exception |
| mongodb | a fixed `root` | from a file **only at first initialisation**, when the data directory is empty |
| mosquitto | a fixed admin client | seeded into the broker's dynamic-security file **once**; the seeding step skips when the file exists |
| lavinmq | a fixed admin name | per its own bootstrap code, **only on a first boot** with an empty data directory. No resource in the definition runs that bootstrap; what sets it on a running mesh is outside the catalogue |
| redis | the default user | from `requirepass` in a configuration the mesh renders, read when the server starts |
| minio | a fixed root user | from a file, read when the server starts |
**In five of seven, a new administrative value takes effect only through a command run with the old
one.** The credential file is mounted directly into both the server and the provisioner. So replacing
it recreates the provisioner, which then holds only the new value while the backend still expects the
old one, and the provisioner is locked out. That is worse than changing nothing.
Every provider module also has its own bus account, an own secret, read at start.
## What the tables show
1. **Every credential provider already rotates in place.** All nine re-apply the password on the
same login each time `create` runs. The controller's `rotate` command relies on that: it replaces
the credential in the inventory and sends both ends in one push. Its own comments state the window,
between the provider applying and the consumer restarting, in which the consumer cannot
authenticate.
2. **Eight of nine name the consumer's resource after its login.** Only gitea separates them,
because an organisation owns the packages. A second login therefore has no resource of its own to
reach, and cannot share the first one's without the adapter granting it.
3. **Five of nine destroy the consumer's data when they remove the login**: postgres, mssql and
mongodb drop the database, lavinmq drops the virtual host with its queued messages, and mailu
deletes the mailbox with its mail. minio drops only an empty bucket, and redis leaves the keys. In
those five, *retire a login* and *delete the consumer's data* are one call. With the harness keyed
by login, a changed login triggers it too.
4. **One backend holds two passwords on one login** (redis). Two hold several tokens beside one
password (gitea and mailu). A rotation built on two secrets per login would work for three
providers out of nine.
5. **Eight of nine can give two logins the same rights over one resource.** Group roles in postgres,
database roles in mssql and mongodb, permissions in lavinmq, a shared role in mosquitto, a shared
policy in minio, a shared key prefix in redis, a shared team in gitea. mailu cannot, because its
user is its mailbox, but it can give one user a second token. So every provider can hold **two
credentials** over one resource, though not every one as two logins. No adapter does either today.
6. **Ownership is a trap in two backends.** In postgres whatever a login creates is that login's, so a
second login cannot alter the first one's tables, and the first cannot be dropped while it owns
them. The one-step way out deletes them. In mssql, a login cannot be dropped with a session open,
nor its user while it owns a schema.
7. **The administrative credentials have one party and a fixed name**, and five backends take them
only at first initialisation. The provisioner needs the old and the new value at once to change
them. Today nothing can give it both.
8. **A consumer's identity is already the resource's name.** The login is derived from the
assignment, which is a module on a node, so the current login and "the consumer" are the same
string today. A second login would need a new name. The resource can keep the one it has.
## Seen on the way
The redis configuration names no ACL file, so a consumer's ACL user exists only in memory. A restart
of the redis server erases every consumer's user. The provisioner does not create them again until it
restarts itself, because its in-memory record says they are done. That is not a rotation finding, but
it is a live fault, and it is recorded here so it is not lost.
@@ -0,0 +1,46 @@
# 02 — The readers
How a secret reaches a running process, and what makes the process take a new one.
## No module watches a secret
A search of every module's code in the catalogue found no file watching of any kind, and no
re-reading of a secret while running. **Every reader reads a secret when it starts.** There is no
consumer that takes a new value live, so every rotation that changes what a consumer presents ends in
the consumer restarting.
## The host already recreates what read a changed file
The node host records, for every long-running container, the digest of each file it read when it
was created: its env-files, and every file bind-mounted into it directly. When a digest changes, the
host recreates the container, even though its spec is otherwise unchanged. This is the fix for
[issue 103](../../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md).
It is on the host's main branch, while the issue is still recorded as located, not fixed.
Two cases are deliberately left out and need `restart-on` in the definition:
- a file read out of a **mounted directory**, because the host cannot know whether the service reads
it once or watches it (a route proxy re-reads its routes live; a provisioner polls what it receives);
- a **process** rather than a container.
## What the catalogue does with it
22 definitions declare a secret they receive. In 16 of them it reaches the service through a
rendered file, usually an env-file. That case the host already covers. 14 declare `restart-on` for
something. Whether each of the 22 is fully covered depends on how its secret travels: through an
env-file or a direct mount, which the host covers, or through a directory or into a process, which
needs `restart-on`. **That was not classified module by module.** It is the check to run before a
rotation mechanism relies on it.
The count covers only the `secrets` field. The 49 modules with their own bus account, and 54 with any
own secret, are readers too, and their bus accounts are rotated like any credential two parties hold.
Their files are mounted directly, which the host covers, but the classification has to name them.
## What this means for rotation
- The *read at start* half of 0113's recipient model is already true, and mostly already handled by
the host. The restart is derived from the files a container reads, not declared per secret.
- Any mechanism, in place or overlapping, ends with the reader being recreated. What differs is
whether the credential it held until then still works.
- For a single-party secret, a module's own, the reader is also the only holder. There is nobody to
overlap with, and delivering the new file recreates the reader.
@@ -0,0 +1,79 @@
# 03 — The options
Three mechanisms, weighed against [01](01-the-providers.md) and [02](02-the-readers.md).
## A. In place, as today
The vault makes a new value. Every applier re-applies it on the same login, which all nine
providers already do. Every reader is recreated by the host.
- **Works with:** every provider, unchanged. It is what `rotate` does now.
- **Costs:** a window per consumer, from the provider applying to the consumer being recreated. They
are on different machines, and nothing orders them. A reader whose machine is unreachable from the
mesh but still reaches its provider stays locked out until the mesh reaches it again.
- **Admin credentials:** the natural form. The provider module is the only party, and it has to
apply the new value with the old one anyway (finding 6).
## B. Two secrets on one login
The applier adds the new password beside the old one, readers move, and the old one is removed.
- **Works with:** redis natively, and gitea and mailu through tokens. **Not** with the other six, whose
backends hold one password per login (finding 4).
- **Verdict:** not a mechanism, a special case. Using it where it exists and something else
elsewhere is the "this way or that way" the design is trying to remove.
## C. Two credentials per consumer, over one resource
The consumer has two credentials and uses one at a time. The applier ensures the other with the new
value and gives it the same rights over the consumer's resource. Readers move to it, and then the old
credential is retired, which removes the credential only, never the resource. **What a credential is,
is the adapter's**: a second login for eight providers (finding 5), a second token on the same login
for mailu. The mesh sees one mechanism.
- **Works with:** every provider, **after** each adapter changes:
- the resource is named after the consumer, not the login. Today the two are the same string (finding
8), so existing resources keep their names, and the current login stays one of the two;
- the resource is owned by the resource, not by a login. In postgres that is a role no one logs in
as, which each login works as, and ownership of an existing database moves to it once (finding 6);
- both credentials get the same rights, over data and structure;
- *retire a credential* and *remove the consumer* become two operations. Today they are one call, and
in five providers that call destroys data (finding 3). The harness must key by consumer, so that a
changed login is not a removal. This is the whole of the danger, and it has to be split, whatever
else is chosen.
- **Costs:**
- every credential adapter changes;
- the harness learns the alternation and a confirmation per step, and rotation state has to live
somewhere that survives a restart, which the harness's memory does not;
- the second login's name must fit the tightest backend. That is 20 characters for a minio access
key ([ADR 0049](../../02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md)), and
a suffix spends part of it;
- retiring an mssql login has to end its sessions first.
- **Gains:** no window. A reader that cannot be reached keeps a working credential until it can.
- **Does not apply to** single-party secrets: admin credentials and a module's own secrets. There is
no second party to overlap with.
## Independent of the choice
- **Split remove, and key the harness by consumer.** Retiring a credential, or a login changing, must
never be able to destroy a consumer's data. That holds under A too, because A's remove is the same
call.
- **Admin credentials are applied by their own provider**, using the old value, with the new one staged
beside it. Five backends take the value only at first initialisation. Replacing the file first locks
the provisioner out (finding 7).
- **Classify the readers** ([02](02-the-readers.md)), bus-account readers included, before relying on
derived restarts.
## Recommendation
- **Two-party credentials, consumer credentials and bus accounts: C**, because it is the only
mechanism every provider supports, and it closes the window instead of shortening it. Its prerequisite, separating the resource from the login and
retiring a login from removing a consumer, is worth doing on its own, because it removes a
data-loss path that exists today.
- **Single-party secrets (admin credentials, a module's own): A, staged.** In place, applied by the
provider that holds them, with the new value beside the old until it has taken.
- **Until the adapters are changed, A stays** as `rotate` implements it, with its window stated. It is
not replaced by a mechanism the providers cannot yet carry.
This is two mechanisms, split by a property of the secret rather than by provider: whether it has one
party or two. Every provider is treated the same way for the same kind of secret.
@@ -0,0 +1,47 @@
---
status: active
initiated: 2026-09-26
touches:
- 00-META/mission.md
- 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0010-delivery.md
- 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
- 03-DESIGN/01-to-be/06-the-controller.md
- 03-DESIGN/01-to-be/09-the-node-lifecycle.md
- 03-DESIGN/00-as-is/09-interfaces-and-observability.md
---
# 017 — A mesh that heals itself
**What.** The behaviour the operator wants: a mesh that runs itself. It notices what is wrong,
repairs what it can, and hands what it cannot repair to someone who can, with the reason. This effort
writes that wish down as intended behaviour, designed for the bus the mesh is moving to
([ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md): NATS). It also records what can be done
pragmatically before that move.
**Why.** The mission is *a mesh that controls itself* ([mission](../../00-META/mission.md)). The
mesh can tell whether it is up. It cannot tell whether it is right. The as-is page on observability says so
([as-is 09](../../03-DESIGN/00-as-is/09-interfaces-and-observability.md)). To-be 06 names an
`observability` context in the controller and leaves its store undecided. Nothing routes a condition
the mesh cannot fix to anyone. The cost is measurable: **46 of the 116 issue reports in this
repository describe a failure that was silent.** A mesh that heals itself is, first, a mesh that stops
failing silently.
**What it touches.** The controller's observability context, the node lifecycle's liveness, delivery
([ADR 0010](../../02-DECISIONS/0010-delivery.md), [ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)),
the provisioner harness, and rotation, which is proposed alongside to-be 27 as ADR 0114.
**Documents.**
- [01 — The intended behaviour](01-the-intended-behaviour.md): the wish, as principles and as how
the mesh behaves once the bus is NATS.
- [02 — Now, pragmatically](02-now-pragmatically.md): what is done before NATS, why it does not
build anything the move would throw away, and what has been done already.
**Next.** Two measurements this effort owes before it can graduate:
1. **Every loop in the mesh**: what it converges, and whether it compares against observed state or
against its own memory. Issue 120 found the provisioner harness trusting memory. The same pattern is
expected elsewhere.
2. **The 46 silent failures, classified**: a missing observation, a loop trusting memory, or a missing
escalation. That shows which mechanism removes the most of them.
@@ -0,0 +1,85 @@
# 01 — The intended behaviour
The operator's wish, written as behaviour: what a person or an agent sees the mesh doing. This is a
target to design toward, not a design. Every part of it is to be decided through a record before it
is built.
## Principles
**1. Every loop compares what should be with what is, never with what it did.** Desired state is the
mesh's: assignments, requirements, seats. Observed state is read from the thing itself: the container,
the backend, the node. A loop that compares against its own memory of what it applied is blind to
anything that changed behind its back. That is issue 120, and it is the pattern this whole effort is
written against.
**2. Healing is the ordinary path run again, never a second path.** Repairing a lost login is
provisioning it. Repairing a dead container is converging the node. Repairing a stale declaration is
delivering it. A repair that needs its own code is a second way of doing something, which is exactly
what the mesh is removing everywhere else.
**3. A repair never destroys.** Healing may recreate, re-provision, re-deliver and restart. It may
never delete a consumer's data, retire a credential someone still uses, or pick a winner between two
contradictory states. Where the only repair is destructive, it is escalated.
**4. Nothing fails silently.** Every condition the mesh cannot repair within its budget becomes
visible. It is named, it says since when, why, and who can resolve it. It is visible until it is
resolved, and resolved by observation, not by someone clicking it away.
**5. What the mesh cannot fix goes to an agent.** Per the mission, an agent may be human or not. A
condition that needs judgement is handed to one, as work, with what the mesh knows. It is not handed
over as a notification that someone may or may not read.
**6. Correctness, not only liveness.** A running process that authenticates with a dead credential,
serves an old version, or routes nowhere is not healthy. What a provision's contract promises is what
is checked: the credential authenticates, the route answers, the version is the declared one.
## The loop, everywhere
Every part of the mesh that owns something runs the same loop:
1. **know** what should be true: from assignments, requirements and seats;
2. **observe** what is true: from the thing itself, on its own cadence;
3. **repair** the difference by running the ordinary path again, within a budget of attempts and
time;
4. **raise** a *condition* when the budget is spent or the only repair is destructive;
5. **clear** the condition when observation shows it resolved.
A **condition** is a durable fact about something the mesh owns, such as a node, an assignment, a
provision, a seat or a rotation: what is wrong, since when, the evidence, what was tried, and who can
resolve it. Conditions are the one thing a person or an agent looks at to know whether the mesh is
right. `status` is the list of open conditions. When it is empty, the mesh is right, not just up.
## On NATS
[ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md) moves the bus to NATS, and NATS makes most of
this cheaper, because observation becomes something every component publishes rather than something
a central process polls.
| the wish needs | on NATS |
|---|---|
| every component says it is alive | a heartbeat on a subject per node and assignment; silence past its interval is a condition, and nobody polls |
| every component says what it observed | observations published on subjects (`mesh.observed.<node>.<assignment>`, for instance), consumed by whoever owns the comparison |
| the last known state survives restarts | a JetStream key-value bucket of observed state per owner; the provisioner's "what I applied" and a rotation's step live there, not in memory |
| conditions are durable and watchable | conditions as entries in a key-value bucket, watched by anyone who cares: a surface, an agent, the controller |
| the bus itself is observed | the server's advisories (a consumer exceeding its deliveries, a slow consumer, a client disconnecting) and its monitoring endpoint become observations like any other |
| a repair is retried, not lost | JetStream redelivery with delay, which is the same mechanism [ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)'s guarantee moves to |
| work handed to an agent | a condition that needs judgement published as a task on a subject an agent's queue group consumes |
**Who compares.** Each owner compares its own: the host for its node's containers and files, a
provisioner for its backend, the vault for rotations, the controller for delivery and seats. The
controller's observability context does not repair anything. It holds conditions, their history,
and the view across the mesh. It notices what no owner can see about itself: an owner gone silent.
## What stays human
Some repairs need the operator's key, and the mesh says so rather than pretending otherwise:
re-raising the vault or the broker, and recovering a node's identity. These are conditions too, with
the procedure named, and they are the only ones that can never clear themselves.
## Open
- The budgets: how many attempts, over how long, per kind of repair.
- How a condition that needs judgement reaches an agent, and how the agent's action is recorded.
- Where the observability context stores history (to-be 06 left it open; volume argues against the
relational store).
- Which correctness probe each provision's contract offers, and how often it runs.
@@ -0,0 +1,46 @@
# 02 — Now, pragmatically
The intended behaviour lands on NATS. The bus moves after the migration's core
([ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md)). Until then, work toward it is chosen by one
test:
**Does it survive the move?** A change to what a loop compares against, or to what an adapter can
tell about its backend, survives, because it is independent of the bus. A new AMQP queue for health
reports, a poller written against the broker's management API, or a condition store built on the
current broker does not survive, and is not built.
## Done
**The provisioner asks the backend, not memory** (issue 120). The SDK's harness gained an optional
`holds` on the adapter, asked for every applied consumer every minute. A consumer the backend no
longer holds is provisioned again. Being unable to ask is not treated as loss. The cache module
implements it first, because its server keeps its users in memory and forgets them all on a restart.
That was verified against a real server: a restart erases every consumer's user, and `holds` answers
correctly for absent, present, wrong-password, disabled and deleted.
Changes: mesh-sdk PR #7 (0.1.1) and mesh-catalog PR #84.
This is principle 1 applied to one loop. It survives the move unchanged. On NATS, the harness's
record of what it applied moves from memory into a key-value bucket, and `holds` stays as it is.
## Next, in order of silent failures removed
1. **`holds` for the other credential providers.** Each backend can answer whether a login exists
with the mesh's password without changing anything. Where a backend cannot check a password without
logging in, logging in is the check.
2. **The harness's other blind spot.** A consumer that goes away while its provisioner is down is never
removed. The fix is the same principle in reverse: list what the backend holds, and compare it with
what the mesh asks for. Removal stays subject to ADR 0114's rule that it never follows from a login
changing.
3. **`status` reports what owners already know.** The host knows which containers it recreated and
why. A rotation knows who it waits on. Delivery knows what is outstanding. Surfacing those as
conditions in the existing `status` needs no new transport. It is the shape the NATS condition store
will hold.
4. **The loop inventory and the classification** in [00](00-overview.md). They decide what comes after
these three.
## Not now
- Heartbeats, observation subjects, key-value state, advisories: all NATS, all after the move.
- Handing conditions to agents: designed with NATS, where a task on a subject is native.
- Choosing the observability store: decided when there is something to store, which is after the
move.
@@ -1,6 +1,6 @@
--- ---
topic: what runs on it topic: what runs on it
status: proposed status: accepted
date: 2026-08-30 date: 2026-08-30
deciders: jochen deciders: jochen
reconstructed: false reconstructed: false
@@ -0,0 +1,105 @@
---
topic: how we work
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0019-how-this-repository-works.md
---
# 25. The design record is read where it is written, never copied to be found
## Context
**These documents cannot be found by searching the mesh's memory, and never could.** Checked on
2026-08-23 and again on 2026-08-31, against both the symptom-indexed store and the structured
archive, using a decision record's full title and a distinctive phrase from a design document: no
result, no partial match, no stale copy.
That matters because of what was promised. The objection to giving this material its own
repository was that the mesh already has a knowledge store, and a second one repeats the mistake
that store was created to fix. **The answer offered was indexing rather than location** — that
these documents would be returned beside everything else in a search, so where they were authored
became a separate question. The indexing was never built.
**The claim has since stopped being load-bearing**, which is why this is a decision rather than an
incident. [`README.md`](../README.md) names the gap in the place the claim used to sit, and
[ADR 0019](0019-how-this-repository-works.md)'s reasoning rests on cadence, reviewers and scope —
none of which depend on being searchable from elsewhere. What remained was an unbuilt capability
and an open question, recorded as
[`04-ISSUES/006`](../04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md).
**A signpost was added on 2026-08-31 and measured.** One entry in the mesh's memory naming what
lives here and when to come looking. A search for *design records, decisions, repository* returns
it; a search phrased the way somebody actually asks — *why is the mesh built this way* — returns
nothing, because the store matches terms and not meaning. **Reachable is not the same as
surfacing**, and the measurement is what established which one a signpost buys.
## Considered Options
1. **A one-way sync into the mesh's memory.** A job reads this repository on a schedule and writes
the documents into the searchable store. It works with what exists today and needs nothing
built first. **Rejected**, because it creates a second copy of every document, and the failure
mode of a derived copy is the one this repository is least able to tolerate: *the copy that is
searched quietly stops matching the copy that is edited*, and the enforced one wins. A design
record that has silently diverged from the reasoning it claims to carry is worse than one that
cannot be found — the first misleads, the second merely fails.
2. **Leave the signpost and close nothing.** Honest, free, and it keeps the gap visible.
**Rejected as an end state**, though it is what stands until the option below exists. It
answers only for a reader who already suspects these documents exist, which is precisely not
the reader the mesh's memory is designed for.
3. **An agent reads this repository directly, and the search consults it.** Nothing is copied.
**Adopted.**
## Decision
**The design record is read where it is written.** Retrieval is an agent reading this repository,
not a copy living in a second store — and a search of the mesh's memory consults that agent, so
what it knows appears beside ordinary results rather than only when it is asked.
Both halves are the decision. The first alone is merely a reader, and would leave this repository
reachable but not surfacing — the state measured above. **The second half is what discharges the
promise** that these documents are returned beside everything else.
**There is no copy, and that is the point.** No sync, no schedule, no reconciliation, and nothing
that can drift, because there is only ever one of each document. It is also always current,
including for work that is not yet committed.
**The direction of reading is one-way and stays that way.** The agent reads this repository and
answers from it. Nothing flows back: this repository is public, the mesh is not, and a return path
would be how installation-specific detail arrives into documents that must not carry it
([`README.md`](../README.md)).
## Consequences
**This repository stops being a fourth knowledge system, properly.** The original objection was
about adding a knowledge *system*. An agent with read access adds no store at all — which answers
the objection more completely than the indexing that was promised, rather than merely as well.
**ADR 0019's promise is amended, not satisfied.** It said these documents would be *indexed*. They
will not be. They will be *read*, and the search will ask. The commitment that survives is the one
that mattered — that a searcher finds them without already suspecting they exist — and the
mechanism behind it is different from the one named.
**It is gated on an agent that does not exist yet.** Until it does, the signpost is what stands,
and this repository is reachable rather than surfacing. That is a known and stated gap, not a
silent one — and the gap is now a build task with a decided shape rather than an open question.
**The search must degrade honestly.** When the agent cannot be reached, a search has to say that
this material was not consulted. A result set that silently omits it looks identical to one where
nothing matched, and *silence and success must never look alike*
([ADR 0004](0004-a-node-and-how-it-joins.md)) — the rule this repository has now paid for twice.
**A rule states how it is checked, and this one is checkable.** The check is the measurement that
produced this record: search the mesh's memory for a phrase that appears only in a design document
here, and require it back. That check fails today, deliberately, and passing it is what closes
`04-ISSUES/006`.
## References
- [`04-ISSUES/006`](../04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md) —
the gap, the two measurements, and why closing it early was refused
- [ADR 0019](0019-how-this-repository-works.md) — the promise this amends
- [`README.md`](../README.md) — the objection, and the gap named where the claim used to sit
@@ -0,0 +1,130 @@
---
topic: what runs on it
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0004-a-node-and-how-it-joins.md
---
# 26. The mesh has a session of its own, and it is the node session's mechanism
## Context
[ADR 0004](0004-a-node-and-how-it-joins.md) gives every node a session: one per node, permanent,
remembering across callers, its system prompt the node's engram, reachable over the broker like
everything else. **Any node can message any node**, and that is called the one part of the system
that is genuinely a mesh — symmetric, with no centre.
**There is no way to address the mesh itself.** A question that spans machines — *what is running
across all of this*, *which nodes are behind*, *why is it built this way* — has to be put to some
node, which then asks the others. That works, and it makes a mesh-wide question **nobody's
question**: every node answers it as a foreigner, from a position where the whole is not in view.
**Three things independently arrived at the same missing piece.**
[ADR 0025](0025-the-design-record-is-read-not-copied.md), taken hours before this one, commits to
an agent that reads the design repository directly and answers into search. That agent has to
exist, run somewhere, and be askable — and nothing in the record says what it is or where it
lives.
[`14-model-access.md`](../03-DESIGN/01-to-be/14-model-access.md) records, as a gap deliberately
not half-built: *this worker uses that licence is a binding to an agent, not to a node* — and the
provisions model has no consumer identity other than a node. A session that must be assigned a
licence is exactly that consumer, and node sessions are already one.
**And ADR 0004 never said how a session is set up.** It describes behaviour and stops: nothing
states how a session starts, where its context lives, how the engram reaches it, or how a message
off the broker becomes a prompt. There is no design document for it. That gap was invisible until
something had to be built *like* a node session, because describing a second instance of a
mechanism requires the mechanism to have been described once.
## Considered Options
1. **No mesh session; keep relaying through a node.** Costs nothing and works today. **Rejected.**
It leaves mesh-wide questions belonging to nobody, and it does not survive contact with
ADR 0025 — that agent still needs a home, so the thing gets built anyway, unnamed, as an
attachment to whichever node happened to host it.
2. **A new kind of agent, built separately.** Purpose-built for the whole mesh. **Rejected.** It
would hold a session, a memory, a licence and broker plumbing — every one of which the node
session already has. Two implementations of one mechanism drift, and the vocabulary collision
that [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) exists to undo began exactly this
way: two things that were nearly the same, built twice, until neither word meant one thing.
3. **The same mechanism, started in a different context.** **Adopted.**
## Decision
**The mesh has one session, addressed as the mesh, and it is a node session in every respect but
three.**
| | |
|---|---|
| **the context it starts in** | the mesh's, not a machine's — this is the whole of what makes it different |
| **its engram** | the mesh's system prompt, as a node's engram is that node's |
| **its licence binding** | assigned in its own right, not inherited from the machine it runs on |
Everything else is unchanged and deliberately so: it is permanent, it remembers, it is reachable
over the broker, it holds its own tools, and switched off it still answers *I am switched off*
rather than falling silent.
**It runs on the node that holds the control plane** — not for convenience, but because that node
is already the one place excepted from *compromise of a node is compromise of that node*
(ADR 0004). An agent able to reach everything, placed anywhere else, creates a **second** such
place. Putting it where the authority already sits concentrates nothing new.
**It is an addition to per-node messaging and never a replacement.** Every node remains directly
addressable. This is not a preference: ADR 0001 holds that losing the control plane costs *change,
not operation*, and a mesh whose only conversational surface lives on that node would lose the
ability to ask anything while every machine kept running perfectly. **The front door may not be
the single point.**
**It is not an employee** ([ADR 0003](0003-agents-are-persistent-employees.md)). Nobody hires it,
it holds no task queue, it is never drained or reassigned. What it does with work that belongs
somewhere else is **dispatch it** — to node sessions, or to workers — which is what a node session
already does when asked something it does not have.
**It is ADR 0025's reader.** The agent that reads the design repository and answers into search is
this session, not a second one. One agent, one memory, one place to reach; two would both need
that repository and would eventually disagree about what it says.
**"One per node" is about address, not about process count.** ADR 0004's rule — *two and nothing
decides which replies* — forbids ambiguity in who answers when a **node** is addressed. The mesh
session answers when the **mesh** is addressed. The control-plane node therefore hosts two
sessions and no ambiguity, and stating this here is what stops it reading as a contradiction
later.
## Consequences
**The node session's setup must now be designed, and it never was.** This decision is expressed as
*the same as a node session, elsewhere*, which is only meaningful once that mechanism is written
down. The design document covering both is the immediate consequence of this record, not a
follow-up to it.
**A consumer that is not a machine stops being deferrable.** The licence binding above is the gap
`14-model-access.md` names, and it now has two consumers rather than a hypothetical one. Until it
exists, a session's model access can only be expressed as *this module on this machine*, which
cannot say *this node's session uses the personal licence and the mesh's uses the company one* —
the thing the binding is for.
**Symmetry is preserved, and it is worth being precise about why.** ADR 0004's claim is about what
a node can reach, and it is untouched: node-to-node messaging is unchanged, nothing is routed
through the mesh session, and it is a participant rather than a hop. What arrives is a
participant that happens to be the one a person usually addresses.
**Availability degrades to inconvenience rather than to silence** — but only because of the
addition rule above. If that rule is ever relaxed, this consequence inverts, and it inverts
quietly: everything keeps working and nobody can ask about it.
**The surface a person uses is not decided here.** That a board is a good place to talk to it is
likely and is not this record's business; the session is reachable over the broker like everything
else, and what puts a text box in front of it is a separate choice.
## References
- [ADR 0004](0004-a-node-and-how-it-joins.md) — the node session this extends
- [ADR 0025](0025-the-design-record-is-read-not-copied.md) — the reader this session is
- [ADR 0003](0003-agents-are-persistent-employees.md) — the vocabulary this is not
- [`03-DESIGN/01-to-be/14-model-access.md`](../03-DESIGN/01-to-be/14-model-access.md) — *a
consumer that is not a machine*, the gap this makes concrete
@@ -0,0 +1,109 @@
---
topic: what runs on it
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0009-modules-and-the-graph.md
---
# 27. A provision names what the consumer is coupled to, not the role it plays
## Context
Provisions are named after roles. The catalogue and every test fixture built so far say:
```
provides: database
requires: database
```
**Nothing distinguishes one engine from another.** A module requiring `database` is satisfied by
any module providing `database`, so a module written against PostgreSQL can be matched to a
provider of Microsoft SQL Server, resolve as satisfied, deploy, and fail on its first query.
**The mesh runs several engines** — PostgreSQL, Microsoft SQL Server, MariaDB, and others behind
products that expose their own. This is not a hypothetical collision.
**The failure is in the direction that hides.** Resolution *succeeds*. Nothing is refused, nothing
is logged, and the breakage surfaces later as an error inside an application, on a machine, with
nothing connecting it back to a match made elsewhere by something that thought it had done its
job. **A wrong answer delivered confidently costs more than a refusal**, and the whole point of
refusing on ambiguity ([ADR 0009](0009-modules-and-the-graph.md)) was to not do this.
**How it got in:** every test written for the resolver had exactly one provider of each name, so no
mismatch was expressible and none was caught. The fixtures agreed with the design. That is the same
fault as [`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)'s
imagined output and [`019`](../04-ISSUES/019-a-comment-asserting-a-fact-about-a-machine/00-report.md)'s
unchecked comment, at the level of a name rather than a line.
## Considered Options
1. **Keep role names; let the operator assign correctly.** The mesh would refuse ambiguity when two
providers exist, so a person picks. **Rejected.** It makes correctness depend on somebody
knowing that the module they are assigning speaks a particular dialect — which is exactly the
knowledge the provisioning model exists to remove. And with one provider of each name, nothing
is ambiguous and nothing is asked.
2. **A role name plus a `flavour:` or `engine:` qualifier**, matched as a second field.
**Rejected.** Two fields that must agree is a constraint the resolver has to enforce and a
manifest author has to remember, to express something one field already can. The name is the
contract; splitting it invites a requirement that names a role and forgets the qualifier, which
then matches everything again.
3. **The name says what the consumer is coupled to.** **Adopted.**
## Decision
**A provision is named for the thing a consumer's code is written against.**
```
provides: postgres-database
requires: postgres-database
```
**The test is whether the consumer can tell the difference.** If swapping the provider would break
the consumer, the name must say which provider — because a match that breaks the consumer is not a
match. If the consumer genuinely cannot tell, a role name is correct and better.
| provision | | why |
|---|---|---|
| `postgres-database`, `mssql-database` | **specific** | applications are written against a dialect; a swap breaks them |
| `route` | **role** | the consumer wants its name reachable and does not care what proxies it |
| `resolver` | **role** | the consumer wants names to resolve |
| `artifact-store` | **role** | the consumer fetches by digest over a protocol, and nothing else |
**`database` is not a provision and may not be provided.** There is no context in which an
application talks to a generic database: it talks to PostgreSQL or it talks to SQL Server. A name
that cannot be true of any real consumer should not be expressible.
**This is about coupling, not about products.** Two providers of `postgres-database` — a container
on this node and a managed instance elsewhere — are interchangeable and *should* both match. What
may not be interchangeable is what the consumer's queries are written in.
## Consequences
**Every manifest that names a database changes.** Doing this now costs a rename across a handful of
examples. Doing it after modules are migrated costs it across all of them, plus every deployment
that resolved against the old name.
**Wrong requirements now fail loudly, and at the right moment.** A module requiring
`postgres-database` where only `mssql-database` is provided is unsatisfiable, so it is **refused at
resolution** with both names visible — rather than deployed and broken later. This is the property
that was lost, restored.
**Generic role names are still right, and the rule says when.** This does not push specificity
everywhere; it puts it exactly where a consumer is coupled. Naming `route` after a particular proxy
would be the same error in the other direction, and would prevent a swap that genuinely changes
nothing.
**It is checked, not merely stated** ([`00-META/how-we-build.md`](../00-META/how-we-build.md) §5).
A manifest providing a name known to be engine-generic is refused, naming what to say instead.
Without that, this record is a convention, and a convention is what the previous naming was.
## References
- [ADR 0009](0009-modules-and-the-graph.md) — provisions, and refusing on ambiguity
- [`03-DESIGN/01-to-be/07-the-foundation.md`](../03-DESIGN/01-to-be/07-the-foundation.md) — *the
provisioning model uses databases, roles and schemas as PostgreSQL means them*, which is this
record's point made about the substrate before it was made about modules
@@ -0,0 +1,118 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0006-the-substrate-and-the-control-plane.md
---
# 28. The substrate supplies the control plane and nothing else
*Corrects one row of [ADR 0006](0006-the-substrate-and-the-control-plane.md) and makes explicit
something it left unsaid. The rest of that record stands.*
## Context
ADR 0006 defines the substrate by a circularity: **what the control plane needs in order to run,
and cannot ask itself for, because it is not running yet.** Two questions, and both must be
answered *yes* for something to be substrate.
Its membership table admits the object store on this line:
| role | product | |
|---|---|---|
| object store | **MinIO** | it cannot grant itself a bucket |
**That answers the second question and assumes the first.** It is true that a control plane cannot
grant itself a bucket. Nothing establishes that it needs one.
**It does not.** Verified 2026-08-31 against `mesh-control`: no S3 client, no bucket, no object
storage of any kind outside comments. Artifacts reach nodes as content-addressed blobs in the OCI
registry, and the code records the decision and its reasoning:
> One store, and it is the registry the bootstrap already pulls from. An OCI registry is a
> content-addressed blob store that happens to also understand images… The alternative considered
> was a second store beside it — S3-shaped, buckets, signed URLs. It is the right answer for
> objects that are *mutable*, or need per-reader access, or are not build output. None of that
> describes a digest-pinned archive, and standing up a second service to hold one kind of
> immutable blob means two things to run, two things to back up and two ways for an artifact to be
> missing.
**The row is inherited from the system being replaced**, where an object store distributed module
tarballs. Here nothing does, and the row was never re-tested against the definition it sits under.
**A second thing ADR 0006 never says:** whether a substrate service and a module of the same
product are the same instance. It says the substrate is *not the control plane* and *not a place
for logic*, and stops. The question is not idle — an application wanting a database, on a mesh
whose substrate is already running PostgreSQL, has an obvious wrong answer available.
## Considered Options
1. **Leave the object store as substrate, unused.** Harmless-looking. **Rejected.** A membership
list that includes something nothing needs is a list that has stopped being derived from its
test, and the next member is admitted by precedent instead of argument. It also mandates that
every mesh run a service no mesh uses.
2. **Applications share the substrate's instances.** One PostgreSQL, one of everything.
**Rejected**, below.
3. **The substrate is exactly what the control plane consumes; everything else is a module.**
**Adopted.**
## Decision
**The object store is not substrate.** It fails the first half of the test: the control plane does
not need one. An object store is an ordinary module, required through the module graph like
anything else, and a module wanting one depends on a module providing one.
**The substrate has four members, not five**: a relational store, a message bus, an image registry,
and conditionally an identity provider. The registry stays — the control plane genuinely cannot
deliver an artifact without somewhere to put it.
**A substrate service and a module of the same product are different instances, and are not
shared.** The mesh's own PostgreSQL and a PostgreSQL a workload was given are two servers, two
containers, two lifecycles.
Three reasons, and the first is the one that matters:
**The substrate is not in the module graph.** It is raised from the pinned bundle the host carries,
before any mesh exists to declare it. A workload depending on it would depend on something the
graph cannot see, cannot rotate a credential for, and cannot move — which is every property the
provisioning model exists to provide.
**It would put workload data in the control plane's own store.** The mesh's contexts own their
stores exclusively ([ADR 0008](0008-a-context-owns-its-store.md)). An application sharing that
server can exhaust it, lock it, or fill its disk, and the failure is the control plane going down
— which is the one failure that makes every other one harder to fix.
**They are bounded differently.** The substrate is sized, backed up and upgraded as part of
bootstrapping a mesh. A workload's database follows the workload — moved with it, destroyed with
it, restored with it.
## Consequences
**Migrating an object store is ordinary module work**, not substrate work. It was previously going
to be done as part of completing the substrate, which would have been the wrong shape and would
have coupled every mesh to a service the mesh does not use.
**A mesh with no workload needing one runs no object store at all.** That is the correct outcome
and was not previously available.
**Two PostgreSQL containers on a node that hosts both is expected**, not duplication to be
optimised away. Anyone tidying them together should find this record first.
**"Substrate by role and ordinary by delivery" loses one of its two members.** ADR 0006 uses that
phrase of the object store and the registry — things that are substrate but provisioned once a
control plane exists. It now describes the registry alone.
**The definition is applied, not just stated.** Both halves of the circularity test are asked of
each member, and *cannot grant itself one* is not sufficient on its own — it is true of almost any
service, which is what made it possible to admit a member on that half alone.
## References
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — the definition, and the table this
corrects one row of
- [ADR 0008](0008-a-context-owns-its-store.md) — a context owns its store exclusively
- `mesh-control internal/builder/registry.go` — where artifacts go, and why not S3
@@ -0,0 +1,112 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0005-the-node-host.md
---
# 29. A network is a shape, because an action cannot be undone
## Context
**A module of several containers has no way to let them reach each other by name.** A container
declaration carries a `network` field, and it only ever *joins* one that already exists — it was
added so the control plane could reach the store and the broker on the machine it was raised on.
Nothing in the vocabulary **creates** a network.
Without one, containers on a machine share the runtime's default bridge, which gives addresses and
no name resolution between them. So a module that is several containers can only be written by
publishing ports onto the machine and pointing its own parts at the host — which puts a module's
private wiring on the machine's own address space, where anything else on the machine can reach it
and any other module can collide with it.
**This is the gap a mail system meets and nothing else so far does**
([`00-work-breakdown.md`](../03-DESIGN/01-to-be/00-work-breakdown.md) 3.3). It is being taken now
rather than then, because 3.3 is the task most likely to send work back into the declaration
language and the least useful place to discover it.
**Adding a shape is not a small change, and the host says so** — the vocabulary is asserted
against a stated number, with the reason written into the failure: *every addition widens what a
compromised control plane can express, so a change here is a decision.* The host applies what it
is told; the only bound on a hostile control plane is what the language can say
([ADR 0004](0004-a-node-and-how-it-joins.md)).
## Considered Options
1. **An `action` that creates the network.** The vocabulary already has one, the bundle already
uses seven of them, and `docker network create` with a `verify` is exactly the shape an action
takes. Nothing would need adding. **Rejected**, on removal:
> An action has no footprint the host can undo — it ran, and whatever it did belongs to
> whatever it acted on.
A network made this way **leaks when the module is unassigned**, and the mesh cannot tell: the
record says an action ran, and there is nothing to reverse. Unassigning a module would leave a
network behind on every machine it was ever on, and the only way to find them would be to go
and look. *A resource the mesh can create and never clean up is one it should not create.*
There is a second reason, and it is the one that generalises: an action is opaque. **The mesh
cannot tell what an action did**, so a network created by one is not a thing the mesh knows
about — it cannot be reported, counted, or reasoned about, and a module could not require one.
2. **Publish ports on the machine instead.** No new shape, and it works today. **Rejected.** It
makes a module's internal wiring part of the machine's address space: two modules that each
want a database on a fixed port collide, and anything else on the machine can reach what was
meant to be private. It also makes the module's manifest depend on what else is installed,
which is the thing provisioning exists to remove.
3. **`network` as a ninth shape.** **Adopted.**
## Decision
**`network` joins the vocabulary, and the vocabulary is nine shapes.**
```
{"id": "internal", "type": "network", "name": "mail"}
```
**A name and nothing else.** Not a driver, a subnet, an address range or a gateway: every one of
those is a thing a module would have to know about the machine it lands on, and a module that
names a subnet is a module that collides with whatever else chose the same one. The runtime picks;
the mesh names.
**It is created if absent and removed when no longer declared** — an ordinary shape, with the same
lifecycle as a directory. That is the whole reason it is a shape.
**Declared before the containers that join it.** Resources are applied in the order the module
wrote them, and orphans are removed in **reverse** — so a network written first is created first
and removed last, after the containers attached to it are gone. This is not a new rule; it is the
existing one, and it happens to be exactly right here. A network written *after* its containers
would fail to remove while they still hold it, and that failure is reported rather than silent.
**What it does not do:** it does not reach across machines. A network is one machine's, like
everything else the host applies. Modules on different machines reach each other over the private
network the mesh already provides ([ADR 0007](0007-connectivity.md)), and a shape that tried to
span machines would be a second overlay with a worse contract.
## Consequences
**The vocabulary is nine, and the count moves with a record.** The test that asserts it names this
one, so the next person to change it finds the argument rather than a number to edit.
**A compromised control plane can now create and destroy networks on a machine.** Stated plainly
because that is the cost, and the bound is the point: it can create a named network and remove
one, and it can do neither to anything it did not declare. It cannot inspect, attach to, or
reroute what is already there — those would be different shapes, and are not being added.
**A multi-container module becomes expressible**, which unblocks 3.3 and, less obviously, makes
several smaller modules simpler: anything that is a service plus a sidecar currently has to
publish a port to talk to itself.
**Nothing is required to use it.** A module of one container declares no network and joins none,
exactly as now. The substrate keeps using `host`, which is a runtime-provided network and not one
the mesh creates.
## References
- [ADR 0005](0005-the-node-host.md) — the host's vocabulary, and why each shape is a decision
- [ADR 0004](0004-a-node-and-how-it-joins.md) — what may be pushed is bounded by form, not by trust
- [`03-DESIGN/01-to-be/00-work-breakdown.md`](../03-DESIGN/01-to-be/00-work-breakdown.md) — 1.3,
and the mail system at 3.3 that this is for
@@ -0,0 +1,96 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0005-the-node-host.md
---
# 30. Data outlives the mesh that declared it
## Context
**The conversion runs on live services holding real data**, and starts on the node that holds all
of it. Identity, mail, everything. The requirement stated plainly: a data directory may be
*moved*, and may never be *lost*.
**The host deleted them.** A directory that stopped being declared was an orphan, and an orphan
directory was removed with `os.RemoveAll` — everything under it — while the report said
`removed`. A module unassigned took its database's files with it, and nothing anywhere said what
had been in there.
Reproduced before it was fixed: assign a module, let a service write into its directory, unassign
the module, and the file is gone.
**A directory stops being declared for ordinary reasons**, which is what makes this sharp rather
than theoretical. A module unassigned from a node. A manifest edited to move a data folder — the
exact operation the conversion needs. A resource renamed. A typo. **Every one of those is a normal
day's work, and every one of them was destructive.**
**The removal order was already right, and that is what makes a fix possible.** Everything the
mesh puts inside a directory is itself a declared resource, and orphans are removed in reverse
declaration order — so by the time a directory is reached, what the mesh wrote there is already
gone. Anything still present was put there by something else.
## Considered Options
1. **A `keep` flag on the directory.** A module declares which of its directories hold data, and
the host leaves those. **Rejected.** It is safe only when somebody remembered, and the failure
of forgetting is total and silent. A rule that protects data only when it was asked to is not
a rule about data, it is a rule about attentiveness — and this is the one place in the system
where being wrong does not recover.
2. **Never remove a directory.** Simple and unarguably safe. **Rejected**, narrowly: every module
ever assigned would leave its directories behind for ever, and a machine that accumulates
things nobody can account for is one where nobody can tell what is still in use. The clean-up
that is genuinely the mesh's is worth keeping.
3. **Remove a directory only when it is empty.** **Adopted.**
## Decision
**A directory that still holds anything is kept, and the mesh says so.** An empty one is removed.
**This is the host's existing line applied to the one shape where getting it wrong is
unrecoverable** — *it removes what it made and leaves what it merely configured*
([ADR 0005](0005-the-node-host.md)). An empty directory is what the host made. A full one is not.
**No flag, no declaration, nothing to remember.** Emptiness is the test, and it is derived from
the removal order rather than asserted: the mesh's own contents are gone by then, so what remains
is by definition something nobody declared.
**It is reported, not silent.** The outcome is `kept`, naming how many items are inside and saying
they are for a person to deal with. A directory quietly left behind is how a machine accumulates
things nobody can account for — which is the objection to option 2, and it is answered by saying
so rather than by deleting.
**Files are unchanged.** A declared file is the mesh's own — it wrote it, it owns it, and losing a
configuration file is not the failure this is about. The distinction is deliberate: **directories
hold what other things produced; files are what the mesh itself put there.**
## Consequences
**Moving a data directory is now safe by default.** The manifest changes, the old path stops being
declared, and the data stays where it is until somebody has looked at it. That was the operation
most likely to destroy something during the conversion, and it is now the operation that does the
least.
**Unassigning a module leaves its data.** Correct, and it means unassignment is no longer a way to
clean up — removing data is a person's act, done knowingly. Given what unassignment did before,
that is the trade being made and it is the right way round.
**A machine can accumulate directories nobody removed.** Accepted, and mitigated by saying so
every time rather than by a periodic sweep. A sweep would be the deletion this record exists to
prevent, on a timer, with nobody watching.
**It is not a backup, and must not be mistaken for one.** This stops the mesh destroying data. It
does nothing about a disk, a mistaken `rm`, or a service corrupting its own store. The conversion
still needs backups taken and **restored** before anything is moved — a backup nobody has restored
is a belief, not a copy.
## References
- [ADR 0005](0005-the-node-host.md) — the host removes what it made
- [`03-DESIGN/01-to-be/00-work-breakdown.md`](../03-DESIGN/01-to-be/00-work-breakdown.md) — the
conversion this was found by planning
@@ -0,0 +1,72 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0006-the-substrate-and-the-control-plane.md
---
# 31. The control plane authenticates nobody, so identity is a module
## Context
[ADR 0006](0006-the-substrate-and-the-control-plane.md) left one member of the substrate
conditional, and said exactly why:
| role | product | |
|---|---|---|
| identity provider | — | **conditional**: substrate only if the control plane delegates authentication, which is undecided |
[`07-the-foundation.md`](../03-DESIGN/01-to-be/07-the-foundation.md) carried it as an open question —
*whether identity is the fifth* — noting it followed from a decision nobody had taken.
**The decision is taken: the control plane does not delegate authentication.** There is no mesh
identity provider.
**Nothing in the mesh's own machinery ever needed one.** A node proves itself with a keypair it
generated, over a broker account issued at enrolment
([ADR 0004](0004-a-node-and-how-it-joins.md)). Declarations are verified by signature. None of
that touches an identity provider, and the conditional was never about machines — it was only ever
about whether a *person* signing in to a mesh surface would be authenticated by something else.
## Decision
**Identity is a module**, like the mail system and the forge. It runs *on* the mesh, not *of* it
([ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)) — a provider other modules require,
which is the ordinary shape and needs nothing new to express.
**So the substrate is three, and no longer conditional**: a relational store, a message bus, and
an image registry. Together with
[ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md), which removed the
object store, the list is settled and every member is there for the same reason — the control
plane needs it and cannot ask itself for it.
**A mesh that wants no identity provider runs none.** That is now expressible, and was not while
it sat in the substrate as a maybe.
## Consequences
**The last open question about substrate membership is closed.** Both halves of ADR 0006's test
now have an answer for every candidate, and the answer for identity is *the control plane does not
need it*.
**It does not settle how a person signs in to a mesh surface**, and that is deliberately left
open. What is settled is that whatever answers it is not part of what must exist before the mesh
does — so it can be decided late, changed, or replaced, which is precisely what being substrate
would have prevented.
**It becomes a real test of the module graph.** An identity provider is a module that *other
modules require* — the object store already consumes it — so it exercises the provider chain more
seriously than anything ported so far, where the provider was written alongside its consumer.
**Ordering follows from it rather than from preference.** Anything requiring identity has to move
after it, which is a dependency the graph can state rather than something a person has to
remember.
## References
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — the conditional this closes
- [ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md) — the other member
removed, and the test applied properly
- [ADR 0004](0004-a-node-and-how-it-joins.md) — how a node proves itself, which needs none of this
@@ -0,0 +1,79 @@
---
topic: how we work
status: superseded
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0031-the-control-plane-authenticates-nobody.md
superseded-by: 02-DECISIONS/0034-the-local-account-owns-the-mesh.md
---
# 32. The local account owns the mesh; a surface delegates to a module
## Context
[ADR 0031](0031-the-control-plane-authenticates-nobody.md) settled that the control plane
authenticates nobody, and deliberately left one thing open: **how a person signing in to a mesh
surface is authenticated.** This answers it, and answers a question 0031 did not ask — *who owns
the mesh at all.*
**There was no answer, and the absence was invisible** because every operation so far has been run
by the person sitting at the machine. Nothing had to say whether that was the design or the
circumstance.
## Decision
**The account that installed the host owns the mesh on that node.** Authority is a local login,
and there is nothing else to hold.
**No mesh user model.** No accounts, no roles, no grants, nothing to administer. A person with a
shell on a node can do anything the mesh can do there, because that is already true and pretending
otherwise would be a boundary that does not exist.
**This follows from what was already decided rather than adding to it.**
[ADR 0004](0004-a-node-and-how-it-joins.md) says there is no authorisation between nodes — every
node is the operator's own, so a message from one is a message from them, and *the mesh boundary
is therefore the security boundary*. A user model inside that boundary would guard nothing: anyone
who could be stopped by it could equally read the node's key off the disk.
**The board is different, and the difference is the network.** A surface reachable by a browser
has to know who is asking, because the people reaching it are not, by construction, people with a
shell on the machine. **So the board delegates to an OAuth provider** — which is a module.
## What this does not change
**The identity provider is still not substrate** (ADR 0031). A *surface* delegating
authentication is not *the control plane* delegating it. The control plane runs, applies
declarations and reaches nodes with no identity provider in existence; only the board needs one,
and only to decide whose browser it is talking to.
The test is unchanged and still answers no: *does the control plane need it in order to run?*
## Consequences
**The board depends on a module, and says so.** An ordinary edge in the graph, which means the
board cannot come up before the provider it authenticates against — stated as a dependency rather
than discovered as an outage.
**Moving the identity provider takes the board with it.** During that module's own conversion the
board is unavailable, and that is acceptable: it is a surface, nothing depends on it, and a brief
interruption is the trade already accepted everywhere else. Nothing that keeps a service serving
goes through it.
**Anyone with a shell on a node has full authority there.** Written down rather than left implied,
because it is the sentence that decides who gets an account on a machine. The protection is the
machine's own login, and the overlay that keeps the machine unreachable from outside
([ADR 0007](0007-connectivity.md)).
**A node cannot be operated by somebody without a login on it.** Deliberate, and the cost of
having no user model: there is no way to give a person authority over one node without giving them
a shell there. If that is ever wanted, it is a new decision and not a gap in this one.
## References
- [ADR 0031](0031-the-control-plane-authenticates-nobody.md) — the control plane authenticates
nobody; this answers what it left open
- [ADR 0004](0004-a-node-and-how-it-joins.md) — no authorisation between nodes, and why the mesh
boundary is the security boundary
- [`03-DESIGN/01-to-be/11-a-board.md`](../03-DESIGN/01-to-be/11-a-board.md) — the surface this is
about
@@ -0,0 +1,91 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md
---
# 33. The substrate is a store and a broker
## Context
Third correction to one table in one day, all found the same way: by asking whether **both** halves
of the substrate test were actually answered for a given member, or only the second.
The test ([ADR 0006](0006-the-substrate-and-the-control-plane.md)) is *what the control plane needs
in order to run, and cannot ask itself for, because it is not running yet.* ADR 0006 admits the
image registry on this line:
| role | product | |
|---|---|---|
| image registry | **an OCI registry** | it cannot grant itself a repository |
**That is the second half again.** It is true that a control plane cannot grant itself a
repository. Nothing establishes that it needs one *in order to run*.
**Counted rather than argued.** `substrate-first-node.lock` — the only bundle there is, and what a
first node actually becomes — raises twelve resources, and no registry is among them:
```
container runtime · the store · one database per context · the schemas
· the broker's certificate · the broker · the control plane
```
The registry arrives afterwards, as an ordinary module the mesh assigns. That is what the lab
asserts, in those words: *the mesh runs its own artifact store.*
**ADR 0006 half-said this already**, calling the registry *substrate by role and ordinary by
delivery, provisioned once there is a control plane to do it.* A member that is provisioned by the
thing it supposedly precedes is not a member; the phrase was carrying a contradiction rather than
resolving one.
**The registry is a closer call than the object store, and the difference is worth keeping.** The
control plane never touches an object store at all — no client, no bucket, ever
([ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md)). It genuinely
*uses* the registry: the builder pushes to it, hosts pull from it, and nothing reaches a machine
without it. **So the registry is a real dependency of the mesh operating, and not of the control
plane starting** — and it is the second that the word substrate means.
## Decision
**The substrate is two things: a relational store and a message bus.** Both are in the bundle,
both must exist before the control plane's first instruction, and neither can be asked for.
**The registry is an ordinary module.** The mesh cannot deliver anything without one, and it
installs one the way it installs everything else. The first node's chicken-and-egg is already
solved and needs nothing from this list: it fetches upstream images directly, then runs a registry
of the mesh's own.
**The test is applied to both columns, every time.** *Cannot grant itself one* is true of almost
any service and settles nothing on its own. It is what admitted the object store, and then the
registry, and both were removed by asking the other question.
## Consequences
**The substrate is now exactly what the bundle raises**, which is the strongest form this list can
take: it can be checked by counting rather than by reading an argument. A member that is not in
the bundle is not substrate, and the two statements cannot drift apart.
**A mesh that builds nothing still needs a registry** — to receive anything at all — but it needs
it as a module, on its own schedule, replaceable. That was already true and was obscured by the
list.
**The word may now be doing too little work.** "Substrate" for *a database and a broker* is a term
of art for two things everybody can name. Renaming is not taken here and is worth considering
separately; what this record fixes is the membership, not the vocabulary.
**Three removals from one table in one day is itself the finding.** Each member was admitted on the
half of the test that is easy to answer, and the design read plausibly throughout. The rule that
comes out of it is not about substrates: **a test with two conditions is a test only when both are
asked.**
## References
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — the definition, and the table this
corrects a second row of
- [ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md) — the object
store, removed for the same reason
- [ADR 0031](0031-the-control-plane-authenticates-nobody.md) — identity, which was conditional and
is now a module
@@ -0,0 +1,81 @@
---
topic: how we work
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
supersedes: 02-DECISIONS/0032-the-local-account-owns-the-mesh.md
---
# 34. The local account owns the mesh, and a web application's login is not that
*Supersedes [ADR 0032](0032-the-local-account-owns-the-mesh.md), which decided the right thing and
described it wrongly. The decision below is unchanged; what it said about the board was an
invention.*
## Context
ADR 0032 answered *who owns the mesh* — the account that installed the host — and then framed the
board as **a surface that delegates authentication**, a category it made up for the occasion. It
does not need one.
**The board is a web application.** It has a login, provided by the identity module, in the way
every web application has a login. That is a fact about an application, not a property of the
mesh, and giving it a name in the mesh's vocabulary implied a relationship that is not there.
The cost of the invented category was not cosmetic. It made the identity module look like part of
the mesh's own authority — something the mesh *depends on* to know who anybody is — when the truth
is that the mesh knows nothing about people at all, and one of the applications running on it has
a login.
## Decision
**The account that installed the host owns the mesh on that node.** Authority is a local login.
There is nothing else to hold, no user model, no roles, and nothing to administer.
**This follows from what was already decided.**
[ADR 0004](0004-a-node-and-how-it-joins.md) says there is no authorisation between nodes — every
node is the operator's own, so a message from one is a message from them, and *the mesh boundary
is the security boundary.* A user model inside that boundary would guard nothing: anyone it could
stop could read the node's key off the disk.
**A web application's login is its own business.** The board authenticates its users through the
identity module. So might anything else the mesh runs. **None of that is mesh authority**, and the
mesh does not learn who anybody is from it.
## The line this draws, which is the reason to write it down
**Signing in to an application must not, on its own, become authority over the mesh.**
Today it cannot: the board reads and does not act
([`11-a-board.md`](../03-DESIGN/01-to-be/11-a-board.md) — *not the way to change things*). Looking
at a page tells you what is true and changes nothing.
**The moment the board can assign a module, whoever it lets in has mesh authority** — and it would
arrive as a feature rather than as a decision. That is the failure this record exists to make
visible, because it is the kind that is only obvious afterwards.
So: **a surface that can change the mesh is a change to who owns the mesh**, and is taken as one.
Not forbidden — wanting to manage nodes from a browser is reasonable — but not something that
turns up in a pull request titled *add assign button*.
## Consequences
**The identity module is not special.** Not substrate ([ADR 0031](0031-the-control-plane-authenticates-nobody.md)),
not part of the mesh's authority, and nothing about the mesh stops working when it is down. Some
applications cannot be logged into, which is what it means for an application's login provider to
be unavailable.
**Anyone with a shell on a node has full authority there.** Unchanged from ADR 0032, and still the
sentence that decides who gets an account on a machine. The protection is the machine's own login
and the overlay that keeps it unreachable from outside ([ADR 0007](0007-connectivity.md)).
**A node cannot be operated by somebody without a login on it.** The cost of having no user model.
If that is ever wanted, the paragraph above says what it costs.
## References
- [ADR 0032](0032-the-local-account-owns-the-mesh.md) — superseded; same decision, invented category
- [ADR 0004](0004-a-node-and-how-it-joins.md) — the mesh boundary is the security boundary
- [ADR 0031](0031-the-control-plane-authenticates-nobody.md) — the control plane authenticates
nobody
@@ -0,0 +1,115 @@
---
topic: what runs on it
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0034-the-local-account-owns-the-mesh.md
---
# 35. One implementation, several surfaces, and what that costs
## Context
The mesh is operated from a command line today. It needs to be operable from a browser and from a
model's tools as well, and the three must not be three different systems.
**The pattern is already in the code and unnamed.** `board` serves HTTP by calling the same
functions the CLI calls; it holds nothing and decides nothing. What follows makes that the rule
rather than a property of one command.
**The board is a presentation layer over the control plane.** Not an application beside it holding
a database credential — the thing that shows what the control plane knows, and asks it to do what
a person asked for.
## Decision
**The logic lives once, in the context that owns it. A surface is an adapter with no decisions in
it.**
| surface | for |
|---|---|
| **command line** | a person on a machine, and the recovery path below |
| **HTTP** | the board, and anything else that speaks to the mesh over a network |
| **model tools** | an agent asking the mesh to do something |
**Every surface refuses identically, because the refusal is not in the surface.** An assignment
that cannot be satisfied is refused by the same resolution whichever way it arrived. The moment a
surface can accept something another would reject, the mesh has two answers to one question and
people learn which to trust.
**Reading and doing are both exposed.** The HTTP surface is not read-only: managing the mesh from
a browser is the point. This takes the decision
[ADR 0034](0034-the-local-account-owns-the-mesh.md) said had to be taken deliberately —
**a browser login now carries authority over the mesh** — and takes it knowingly rather than
letting it arrive with a feature.
**The networked surfaces authenticate through an OAuth2 identity provider.** Named by protocol
rather than by product, like every other dependency the mesh takes — AMQP for the bus, S3 for an
object store, OCI for the registry
([ADR 0006](0006-the-substrate-and-the-control-plane.md)). What fills the role today is a module
running Keycloak; what the control plane knows is that it validates a token against a provider
speaking OAuth2, and replacing that provider is a migration rather than a redesign.
**The command line does not authenticate at all**: it is already behind the machine's own login,
which is what owns the mesh (ADR 0034).
## What this is not: a kernel every module imports
**The shared library is the failure this project was started over**, and the difference has to be
stated or it will be rebuilt. The old one is 155 files and 34,636 lines *containing code from
every context* — work-domain logic sitting in the kernel every module imports, each piece landing
there to avoid a cycle between two modules that both needed it.
**Shared surfaces are not a shared library.** What is shared here is that three adapters call the
same functions. Those functions stay in the context that owns them — provisioning's logic in
provisioning, identity's in identity — and no module imports another's. A surface may call many
contexts; a context still may not reach into another's store
([ADR 0008](0008-a-context-owns-its-store.md)).
The test, when something is about to be put "somewhere shared": *does this belong to a context, or
does it only belong to the surface?* If it belongs to a context it goes there, even if two
surfaces want it.
## The loop this creates, and the way out
**The control plane's networked surfaces will depend on a module the control plane assigns.**
An identity provider is an ordinary module ([ADR 0031](0031-the-control-plane-authenticates-nobody.md)). When
it is down, or being migrated, or misconfigured, the HTTP and tool surfaces cannot authenticate
anybody — including the person trying to fix it.
**The command line is the way out, and it is why local ownership matters more rather than less.**
It authenticates through nothing, needs no network, and is available on the machine to the account
that owns the mesh. **A mesh must always be operable by somebody standing at it.**
So the rule: **no capability exists only behind an authenticated surface.** Anything the board can
do, the command line can do. That is not a courtesy to CLI users; it is the recovery path, and a
capability that exists only over HTTP is one that disappears exactly when identity does.
## Consequences
**Identity is still not substrate**, and the test still answers no: the control plane runs, applies
declarations and reaches nodes with no identity provider in existence
([ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md)). What is unavailable without it is two
surfaces, not the mesh.
**Whoever the identity provider admits has authority over the mesh.** That is now a real perimeter
with real consequences, where before it guarded a page that only read. Who may log in, and to
which realm, becomes a decision about the mesh rather than about an application.
**A surface must not grow an opinion.** The likely erosion is a validation added to the board
because it was quicker there — and then the CLI accepts something the board rejects, or worse the
reverse. Adapters hold no decisions.
**Three surfaces over one implementation is a cost paid three times if it is not one
implementation.** The reason to write this down now is that the second surface is the cheapest
moment to get it right, and the third is where the drift usually starts.
## References
- [ADR 0034](0034-the-local-account-owns-the-mesh.md) — the local account owns the mesh, and the
line this record deliberately crosses
- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — the shared library this must not
become
- [ADR 0008](0008-a-context-owns-its-store.md) — a context owns its store, which a surface does
not change
@@ -0,0 +1,99 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0035-one-implementation-several-surfaces.md
---
# 36. Bootstrap ends at a usable mesh, and the first credential comes from a person
## Context
Bootstrap currently ends when the control plane starts
([`07-the-foundation.md`](../03-DESIGN/01-to-be/07-the-foundation.md)). That is a mesh that runs and
cannot yet be used by anybody who is not standing at the machine: the networked surfaces need an
OAuth2 identity provider ([ADR 0035](0035-one-implementation-several-surfaces.md)), the provider is
a module, and no module has been assigned.
**So bootstrap should go further** — through the identity provider and the first login — and stop
at a mesh somebody can actually use.
**One thing in the way, and it is not incidental.** The mesh has never held a readable secret. The
sealing code says what it does and why:
> Make generates a secret and seals it to both ends, **keeping no readable copy.**
An initial administrator's credential is the first value a **person must read**. Everything else
the mesh generates is something no human ever sees, and everything a human provides is something
the mesh immediately stops being able to read.
## Considered Options
1. **The mesh generates it and prints it once**, to the terminal of whoever ran the bootstrap.
Convenient, and needs no prompt. **Rejected.** It would give the control plane a plaintext
secret for the first time — briefly, and only to one terminal, but the capability would then
exist. *An exception made for one case does not stay one*: the next credential that is awkward
to supply gets printed too, and the property that a copy of the mesh's database is a copy of
nothing stops being checkable by reading the code.
2. **No password: a one-time link that lets the operator set their own.** The nicest to use.
**Rejected for now** — it needs a mechanism that does not exist, and the thing it improves is
one prompt, once, on a new mesh.
3. **The operator supplies it.** **Adopted.**
## Decision
**Bootstrap runs to a usable mesh**: the substrate, the control plane, the identity provider as an
ordinary module, its realm and client provisioned, an administrator able to log in, and the
networked surfaces available.
**The administrator's credential is supplied by the person doing the bootstrap**, on standard
input and not echoed — the path that already exists for a model-access key. The mesh seals it and
cannot read it afterwards.
**What is created is an account in the identity provider, not a user of the mesh.** The mesh still
has no user model and gains none here ([ADR 0034](0034-the-local-account-owns-the-mesh.md)). What
this produces is the first login for the applications that have one.
**The provisioning is ordinary.** A realm, a client and a first account are what an identity
module's provisioner makes from what the mesh granted it — the same shape as a database and a
bucket, which are built and proven.
**The surfaces arrive when their dependency does.** The command API is not started with the
control plane and then broken until identity exists; it becomes available once it can authenticate,
the way anything else waits for a provider.
## Consequences
**The mesh still never holds a readable secret**, and that sentence needs no exception clause.
That is the whole reason for the prompt.
**An unattended bootstrap is still possible, and the value still comes from outside.** Automation
supplying the credential is the operator supplying it. What is refused is the *mesh inventing*
one — so an unattended bootstrap with no credential provided produces a mesh with no
administrator, which is correct rather than broken.
**Bootstrap gains an interactive step**, and it is the only one. Worth stating because a bootstrap
that cannot run without a person is a real constraint on how a node is stood up, and this is
deliberate rather than an oversight.
**The identity provider is still not substrate.** It is assigned by the control plane, so it comes
after it, and a thing that comes after cannot be a thing that must exist before
([ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md)). Bootstrap running through it does not
move it: bootstrap is a sequence, the substrate is a dependency.
**And the recovery path is unchanged.** When the identity provider is broken later — which is the
failure that matters, not the one at first start — the command line still works, because it
authenticates through nothing (ADR 0035).
## References
- [ADR 0035](0035-one-implementation-several-surfaces.md) — the surfaces, and why the command line
must keep working
- [ADR 0034](0034-the-local-account-owns-the-mesh.md) — the local account owns the mesh; this adds
no user model
- [ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md) — what must exist before the control
plane, which this does not change
+103
View File
@@ -0,0 +1,103 @@
---
topic: building it
status: proposed
date: 2026-09-01
deciders: jochen
reconstructed: false
rests-on: 02-DECISIONS/0009-modules-and-the-graph.md
---
# 37. Where a module lives
## The question
The mesh's own module descriptions currently sit in `examples/` inside the control plane, beside
the small programs that hand out logins. That was fine while there were three of them. It is
wrong now, and the name is doing active harm: everything in `examples/` reads as a sketch, and one
of them shipped naming a container image that nothing in the repository builds. A directory called
*the catalogue* would have made *does this actually work* the obvious question to ask of it.
So: **one repository holding the modules we ship?** And if so, where does everything that is not
ours go?
## What a module actually is, counted
The system being replaced has **126 modules** on its main branch. The shape of them is the whole
argument, so it is measured rather than assumed:
| | count | what it is |
|---|---|---|
| **the module is software** | 47 | its own source tree lives inside the module — a daemon, a service, a library |
| **helper scripts only** | 44 | no application of its own; scripts it runs at install time or offers to an agent |
| **a description and nothing else** | 35 | a package to install and some files to write |
**Two thirds of modules contain code.** The largest is a shared library of 182 source files. A
speech-capture module carries a complete daemon — audio capture, mixing, transcription, a model
runner. Treating a module as *a description of something else* is true of barely a quarter of them.
That kills the simplest answer. A catalogue cannot be "a folder of manifests" when most modules
are programs.
## The four kinds, which want different homes
**1. What the mesh is made of.** The control plane, the host, the shared library, the board.
These are not modules that happen to be ours; they are the mesh, expressed as modules so it can
install itself. They belong in the repositories that build them, which already exist.
**2. Something the world made, that we describe.** A forge, a mail system, an identity provider,
a media server. Nobody upstream ships a description; somebody has to write one, and it is the same
description for everybody who runs it. **This is what a catalogue is for.** It is also where the
small programs that create accounts belong, because such a program is part of describing that
service, not part of the mesh.
**3. Something we wrote, that runs somewhere.** An application, a site, a side project. The
description belongs **with the code, at the root of its own repository**, because the two change in
the same commit. A repository that gains an environment variable and a description that gains it
elsewhere will drift, and there is no mechanism that could stop it. This is already how it works
and it should stay that way.
**4. A package and some files.** A tool, a font, a shell. Thirty-five of these, and each is a few
lines. The catalogue.
## The proposal
**A `mesh-catalog` repository** holding kinds 2 and 4: descriptions of software we did not write,
and the programs that provision it. Not kind 1, which is the mesh itself. Not kind 3, which lives
with its own code.
**The mesh's list of modules is not this repository.** It is a table in the control plane, filled
by adding a description to a running mesh. The catalogue is a *source* to add from — one of
several, and the mesh already records which: every module carries where it came from, the branch
followed there, and the commit its description was read at. **Nothing needs inventing to support
modules from anywhere**; a repository of our own is simply the source we curate.
**A description is checked by the tool, not by a test that imports the tool.** Today a test in the
control plane parses the example manifests by reaching into the control plane's internals, and
another reads the control plane's own build file to check every image a module names can be built.
Two jobs tangled. A `module check` command on the control plane's binary would let the catalogue
hold data validated from outside, and would give the same check to somebody describing their own
application in their own repository — which is the case that matters most and currently has no
check at all.
## What this costs, and the argument against
**It is early.** Ten modules exist, four of them ours. Moving ten files is a morning; moving a
hundred is a week — but the hundred is not here yet, and splitting now adds a second repository to
release across before there is anything to release.
The counter is that the tangle is already producing faults rather than merely threatening to. A
manifest naming an unbuildable image, and a test reading a build file two directories up, are both
symptoms of one repository doing two jobs. And the moment the first module is adopted on a real
machine, the descriptions stop being examples and become the thing deployments come from. **That
is the moment this becomes urgent, and it is close.**
## What it does not settle
**Where a provisioning program's image is published**, and how a description pins it. A description
names an image by digest; the image is built from the catalogue; the catalogue must therefore both
produce an image and refer to it, which is the same knot the bootstrap has and solves by writing
the digest down after building.
**Whether kind 4 deserves a module at all.** Thirty-five descriptions that say *install this and
write these files* may be better as one module with settings than as thirty-five modules. Left
open deliberately; it is a question about the shape of the catalogue, not about whether to have one.
@@ -0,0 +1,82 @@
---
topic: what runs on it
status: accepted
date: 2026-09-01
deciders: jochen
reconstructed: false
rests-on: 02-DECISIONS/0009-modules-and-the-graph.md
---
# 38. The mesh assigns the port, and a module does not care
## The problem, as met
A database module cannot start on a machine that runs the control plane. The mesh keeps its own
store there and holds 5432; the module publishes 5432. Nothing notices until a container runtime
three layers down says `port is already allocated`
([`028`](../04-ISSUES/028-two-things-want-one-port-and-nothing-says-so/00-report.md)).
A module cannot fix this by choosing better, because **a module cannot know what else is on the
machine.** It is written once and assigned anywhere. Any number it picks is a guess about a
machine it has never seen, and two modules guessing the same number is not a mistake either of
them made.
## The number is written three times, and nothing makes them agree
Every module says its port in three places:
| where | for | example |
|---|---|---|
| `listens` | the rule set that lets traffic in | `{port: 5432, from: mesh}` |
| `serves` | what a consumer must know to connect | `{port: 5432}` |
| a container's `ports` | what the runtime publishes | `"5432:5432"` |
They agree today because one person wrote all three. Nothing checks it. A module whose `serves`
said 5432 and whose container published 5433 would resolve, compose, apply, and hand every
consumer a port that answers nothing.
## The decision
**The mesh assigns the machine-side port, and the module says only what it needs.** A module
declares that a container port must be reachable and what it is for. Which number the machine uses
is the mesh's to choose, because the mesh is the only thing that knows what else is there.
**One source, and the other two are derived.** `serves` carries the assigned port so a consumer is
told where to connect without the module having written it down; the rule set is computed from the
same assignment. Three copies become one fact.
**An assignment is made once and kept**, exactly as a credential is. A port that moved on every
push would restart both ends each time and would hand consumers a number that was true when it was
read.
## Some ports cannot move, and that is a claim
Mail is 25, submission is 587, IMAP over TLS is 993. A mail system on a strange port is not a mail
system. So a module may say a port is **fixed by the protocol** rather than assigned.
**A fixed port is exactly a claim** — the thing the mesh already has for what is singular on a
machine: one seat, one display server, one artifact store. Two modules wanting 25 on one machine is
the same shape as two wanting the seat, and gets the same answer: the second is refused, by name,
when it is assigned rather than when it is applied.
That is why this does not need a new mechanism so much as it needs the existing one pointed at
ports.
## What follows
- **A module becomes portable in a way it was not.** Two databases on one machine stop being a
collision and become two assignments.
- **The substrate has to be visible.** The mesh cannot assign around its own store while it has
never heard of it. What the bundle holds must be written down somewhere the assignment can read
— which the bundle does not say today.
- **A refusal can be useful.** *25 is held by the mail system on this machine* is a sentence a
person can act on. `port is already allocated` is not.
- **`serves` stops being written by hand**, which is a small vocabulary change with a large
consequence: what a consumer is told is now derived from what actually happened.
## What this does not settle
**Whether a module should publish to the machine at all.** Assignment makes publishing safe; it
does not make it necessary. Consumers could instead reach a provider on the module's own network by
name, with nothing published — which would make the question moot for anything inside the mesh, and
would still leave it for anything reached from outside.
@@ -0,0 +1,106 @@
---
topic: building it
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
---
# 39. What the SDK holds, and what it refuses
_Reconciliation note (2026-09-05): supersedes the earlier "repository structure" decision, which the consolidation folded; no standalone record remains to point at, so body references to it now point at the nearest surviving record, [ADR 0015](0015-applications-live-in-their-own-repository.md)._
## Context
The earlier "repository structure" decision (folded in consolidation; see the reconciliation note
above, and [ADR 0015](0015-applications-live-in-their-own-repository.md) as the nearest survivor)
named `mesh-sdk` "contracts shared across tiers:
types, not behaviour." That line is superseded here, because it draws the boundary in the wrong
place. The boundary that matters is not *types versus behaviour* — it is **how often the thing
changes**.
The current SDK is the cautionary tale, and its failure is precise. `hal/sdk` holds all the
code, including a per-module API client for every service (`clients/plex.ts`, `clients/gitea.ts`,
…) and a per-module tool implementation for each (`tools/plex.ts`, …). Every module depends on
the SDK, so **every edit to any of that per-module code rebuilds every module** — the cascade.
The SDK is under constant maintenance precisely because it became the place all the volatile
per-module logic accumulated.
The root cause is worth stating exactly, because the fix follows from it: the pressure was never
to share a client *between* modules. It was to share a client between one module's *own features*
— plex's tools, its health check and its hooks all wanted the same `PlexClient` — and the only
place to share code across a module's features was the global SDK. So **intra-module sharing
leaked out as inter-module coupling.**
## Decision
The SDK holds the **stable spine** that modules build against, and earns its place by rarely
changing. The test for membership is change-frequency, not kind.
### What it holds
- The **tool-serving harness** — the worker and registration mechanism, and the tool-definition
type. *How* a tool is declared and served is settled; it does not change when an individual
tool does.
- The **messaging and event framework** — the broker client, the event consumer, the envelope.
- The **contracts** — the manifest, declaration, provision and link shapes.
- **Core primitives** — sealing and crypto, semver, the shared resolution helpers.
These change rarely and deliberately. When one of them does change, a rebuild of everything is
the *correct* outcome, because the contract every module shares has genuinely changed.
### What it must not hold — the more important half
- **A module's API client.** A Plex client, a Gitea client, a MinIO client belong in their
module. They change when that service's API or the module's use of it changes, which is often,
and which has nothing to do with any other module.
- **A module's tool implementations.** Same reason, same place: in the module.
- **Anything volatile** — anything that changes when one service's features change.
The rule, stated so it can be applied without re-deriving it:
> If editing a thing recompiles unrelated modules **and** it changes often, it does not belong
> in the SDK.
Both conditions are load-bearing. A rare change that cascades is fine — that is a contract, and
the cascade is correct. A frequent change that stays local is fine — that is a module minding its
own business. Only **frequent *and* cascading** is the disease, and per-module clients and tools
are its carriers.
### Where per-module shared code lives instead
Code shared among a module's *own* features lives **in the module**. The default is the plainest
thing that works: an ordinary shared file the features import — `plex/client.ts`, imported by
`plex/tools/`. Within one module, features are files importing sibling files; no package
boundary, no ceremony.
A **module-local SDK** (a sub-package with its own version) is warranted only for the few modules
whose shared surface is large enough to version on its own. It is the exception, not the shape.
Either form gives the property the global SDK could not: editing a module's shared code rebuilds
**that module and nothing else**.
## Consequences
- The cascade becomes **structurally impossible for module logic**. There is no longer an edge
from one module's internals to another, so the only thing that can rebuild everything is a real
change to a shared contract in the SDK — which is rare, and when it happens, is right.
- The SDK is small and stable **by construction**, not by discipline. Its size is no longer a
thing anyone has to police.
- **Converting a module from the current system is partly a de-coupling, not just a move.** Its
client and its tools are pulled *out* of the shared SDK and *into* the module. A conversion
that copied `clients/plex.ts` into the SDK's replacement would rebuild the exact mistake.
- The host still does not import the SDK. It depends on nothing
([ADR 0005](0005-the-node-host.md)) and **mirrors** the contracts rather than
importing them, exactly as its apply-shapes table already does deliberately. The SDK is shared
by the tiers that *can* share code; the host is not one of them.
## References
- The earlier "repository structure" decision — named the repositories; its `mesh-sdk`
description ("types, not behaviour") is superseded by this record (folded in consolidation;
nearest survivor [ADR 0015](0015-applications-live-in-their-own-repository.md)).
- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — the decomposition this serves: code
belongs to the boundary that owns it.
- [ADR 0005](0005-the-node-host.md) — why the host mirrors the contracts instead of
importing the SDK.
+101
View File
@@ -0,0 +1,101 @@
---
topic: what runs on it
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
extends: 0009-modules-and-the-graph.md
---
# 40. What a module is
_Reconciliation note (2026-09-05): supersedes the earlier "grouped by domain" decision, which the consolidation folded into how-we-build.md; no standalone record remains to point at._
## Context
[ADR 0009](0009-modules-and-the-graph.md) settled that everything is a module, but never said what a
module *is* beyond "a directory the mesh processes." That gap let the catalogue's breadth read as a
smell: a module can carry a container, a built image, tools, a provisioner, migrations, health,
config, seat claims, requires and provides — so much that the unit seemed ill-defined.
The earlier "grouped by domain" decision (folded in consolidation; see the reconciliation note
above) tried to organise modules by domain, which is the wrong axis. This record states what a module is, drawn from the cases that
stress-tested it: the shell, i3-vs-sway, umami, and "database."
## Decision
**A module is one self-contained piece of software the mesh installs and manages** — everything
needed to make that one thing real and integrable: what runs, the seats it claims, what it provides
to other modules, what it requires from them, and what operates it.
The **software is the module's identity.** Capabilities, seats and provisioned resources are the
**relationships *between* modules**, not what a module is — and that is what binds a module into one
thing. umami is bound by *being umami*: its container runs umami, its provisioner creates umami sites,
its tools query umami, its `requires` gets umami a database. Every feature serves the one software.
### The three relationships
1. **Shared seat** — several modules fulfil a capability and coexist; one may be default. bash, zsh
and fish all join `shell`.
2. **Exclusive seat** — modules contend for a single slot; one holds it. i3 (needs x11) and sway
(needs wayland) contend for `display-session`.
3. **Provide / require** — a provider ships the **provisioner** that creates instances of the
resource it offers and returns sealed credentials; a consumer requires it and the mesh wires the
credential in. Symmetric: umami requires a database *and* provides analytics.
### Interfaces are mesh-owned; providers adapt to them
The mesh **defines the interface** for a capability — the provider-neutral contract of what a
consumer receives and how it integrates. Both sides conform: a provider's provisioner **adapts** its
software's real API to the mesh contract; a consumer depends on the **interface**, never on a
provider. Swap one provider for another and the consumer does not change.
### The naming rule — draw the interface at the consumer's real coupling
Name a `provides`/`requires` at the **widest boundary across which the consumer genuinely does not
care which implementation serves it**:
- Where the consumer's coupling is thin — an analytics embed snippet and dashboard, opaque to it —
the mesh defines a neutral interface (`analytics`) and providers (umami, amumi) adapt. Swappable
across vendors.
- Where the consumer **speaks a protocol** — a database's wire protocol and query dialect — the
interface *is* the protocol: `postgres-database`, `mssql-database`, `mongodb-database`. Swappable
only among protocol-compatible implementations, **never across**, because the application cannot
cross it either. "database" is not a capability; the protocol is.
- **Never false genericity.** A name must not promise a swap the contract cannot deliver
([research 005](../01-RESEARCH/005-domain-grouping/analysis.md)).
This is [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md)'s rule made general — "names the protocol, not
the product; a database names the engine because the app targets it" — with the reason stated: the
contract sits where the coupling is.
### What is not a module
- A **library** (built against, never deployed — [ADR 0039](0039-what-the-sdk-holds-and-refuses.md)).
- A **control-plane context** (the mesh itself — [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)).
A **swappable machine mechanism** (a firewall — ufw, nftables) *is* a module implementing a
capability. The host hardcodes no firewall, supervisor, package manager or runtime; it owns only the
generic apply primitives and platform detection, so it runs where none of those exist — an Android
phone has no ufw, systemd, pacman or Docker.
## Consequences
- **Supersedes the earlier "grouped by domain" decision** (folded in consolidation; see the
reconciliation note above). Modules are
organised by their relationships (seats, provisions), not grouped into domain folders.
- **Refines [ADR 0009](0009-modules-and-the-graph.md).** Everything the mesh runs and integrates is
a module — but a module is defined by the *software it delivers*, not by being a bucket of features.
- The target is a **self-fulfilling mesh**: declared wants bound to swappable modules, provisioners
wiring credentials, nothing hardcoded. The control plane's whole job is the binding.
- Converting a module from the old system includes pulling its per-module code out of the shared SDK
([ADR 0039](0039-what-the-sdk-holds-and-refuses.md)) and shipping its provisioner as an adapter to a
mesh interface — a de-coupling, not just a move.
## References
- [ADR 0009](0009-modules-and-the-graph.md) — everything is a module; this says what one is.
- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — contexts are the mesh, not modules.
- The earlier "grouped by domain" decision — superseded (folded in consolidation; see the note above).
- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — protocol-not-product, generalised here.
- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — per-module code lives in the module.
- [research 011](../01-RESEARCH/011-the-module-graph/00-overview.md) — the graph of these relationships.
@@ -0,0 +1,101 @@
---
topic: what runs on it
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
extends: 0040-what-a-module-is.md
---
# 41. Events are a relationship, the lighter sibling of provisioning
## Context
[ADR 0040](0040-what-a-module-is.md) names two relationships between modules — seats and
provide/require (provisioning). A third is latent in the mesh and worth making first-class: the
broker every node already runs ([ADR 0002](0002-nodes-communicate-over-a-broker.md)) can carry a
module's activity as **events**, which any other module reacts to. A logger that writes an audit
trail, a module that acts when another module acts, observability — all of it is one mechanism, and
today it is ambient rather than declared.
## Decision
**A module emits events and consumes events, and both are declared** — parallel to `provides` /
`requires`, so the mesh knows the event graph the same way it knows the provisioning graph.
### Events are provisioning's lighter sibling
| | provisioning | events |
|---|---|---|
| shape | **1:1**, a provider creates a resource *for* one consumer | **1:many**, a module emits, any number listen |
| credential | yes — sealed, per consumer | none — it is broadcast |
| machinery | a provisioner (the reconcile adapter) | nothing but the broker's topic routing |
| declared as | `provides` / `requires` | `emits` / `consumes` |
Because an event is broadcast and credential-free, there is no provisioner and no per-consumer
setup — only a subscription. That is why it is the *lighter* relationship, and why most
inter-module reaction should be an event, not a provision.
### An event carries what an audit needs
Every event carries its **type** (a dotted topic key, so listeners match by prefix), its **source**
module, the **node** it came from, and the **time**. A body follows. The metadata is not optional:
a reaction may only need the body, but an audit trail needs to know who did what, where and when,
and an event that cannot answer that is not auditable.
### The audit logger is just a consumer of everything
A logger that records the whole mesh's activity is **not a privileged component** — it is an
ordinary module that consumes `#` (every event) and writes them down. It holds no special access;
it only listens widely. That it falls out of the model with no new machinery is the check that the
model is right.
### `consumes` is validated like `requires`
A `consumes` for an event that **nothing** `emits` is a dangling edge, and the mesh refuses it
before deploy — the same rule that catches a `requires` for a resource nothing provides
([research 011](../01-RESEARCH/011-the-module-graph/00-overview.md)). A listener waiting for an
event that can never arrive is a silent failure, and this repository's whole discipline is against
silent failure.
### One runtime serves all three
The per-node module runtime that serves a module's tools also wires its `consumes` (subscribe,
dispatch to the handler) and lets its code `emit`. Tools are *invoked* (request/reply), resources
are *provisioned* (1:1, credentialed), events are *emitted and consumed* (1:many, broadcast) —
three relationships, one broker, one runtime, all declared on the manifest.
## Consequences
- The mesh gains a declared **event graph** alongside the provisioning graph — visible, validated,
reasoned over.
- **Reaction becomes the default coordination**: a module acts on another's event without either
knowing the other, and without a credentialed link. Coupling drops.
- An **audit trail** is a module, not a platform feature — and can be swapped, extended or run more
than once (a file logger and a queryable one) with no change to anything that emits.
- The runtime must dispatch a module's event handlers as well as its tools; that generalisation is
small (both arrive by importing the module's entrypoint) but it is real work.
## Progressive insight
> **Progressive insight — 2026-09-26.** *"No provisioner and no per-consumer setup" was a fact
> about the transport, and the transport changed.* This record's table says an event's machinery is
> "nothing but the broker's topic routing", and the text that an event needs "no per-consumer setup
> — only a subscription". That was true of a topic exchange, where a binding cost nothing and the
> broker fanned out. On NATS
> ([ADR 0106](0106-the-bus-is-nats.md)) a subscription is a **durable consumer**: a real object
> with a name, an ack policy, a delivery limit and its own ack subject, created when a module is
> assigned and removed when it is not. Per-consumer setup exists, and the controller does it.
>
> The decision is untouched — events are declared on both sides, 1:many, credential-free, and
> still provisioning's lighter sibling; the lightness is now relative rather than absolute.
> [ADR 0126](0126-a-module-declares-its-own-seats.md) adds the relationship this record's two
> columns had no room for: work addressed to a role, where exactly one holder must act.
## References
- [ADR 0002](0002-nodes-communicate-over-a-broker.md) — the broker events ride.
- [ADR 0040](0040-what-a-module-is.md) — the relationships this extends.
- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — `emit`/`on` are stable sdk surface; the
broker binding and the runtime are not.
- [research 011](../01-RESEARCH/011-the-module-graph/00-overview.md) — the graph these edges join.
@@ -0,0 +1,116 @@
---
topic: what runs on it
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
extends: 0041-events-are-a-relationship.md
---
# 42. The shape of an event on the wire
## Context
[ADR 0041](0041-events-are-a-relationship.md) made events a relationship — `emits`/`consumes`, the
graph, the audit logger. It did not say what an event *is* on the broker: the exchanges, the
routing keys, the headers, the queues and their configuration. That shape is a contract every
emitter and consumer conforms to, exactly as [ADR 0010](0010-delivery.md)
is for declarations — and it was being decided ad-hoc in code. This settles it, so the sdk and the
runtime implement one contract and a module never reinvents it.
## Decision
### Two exchanges, kept apart
- **`mesh.events`** — a durable topic exchange. Every event rides it: module, mesh and node.
- **`mesh.rpc`** — a durable topic exchange. Tool invocations (request/reply) ride it.
Kept separate because RPC is not an event: a `#` subscription on `mesh.events` is then a complete
audit of what happened, with none of the invocation traffic.
### The routing key is the event type, namespaced by origin
Dotted and hierarchical — `<origin>.<name>.<event…>` — with three reserved origins:
- `module.<module>.<event>` — `module.umami.site.created`
- `mesh.<context>.<event>` — `mesh.delivery.deployed`, `mesh.provisioning.granted`
- `node.<node>.<event>` — `node.anchor.joined`, `node.anchor.unreachable`
Topic matching gives a consumer `node.*.joined`, `module.umami.#`, or `#`. The origin roots are
reserved; everything after is the emitter's own namespace.
### Metadata in headers, payload in the body
An event's identity and provenance are AMQP **headers**, so a consumer — or the broker, or an
audit tool — reads who/when/what without parsing the body, and the body is only the domain payload.
**Required headers**
| header | meaning |
|---|---|
| `x-event-id` | a unique id — for dedup and audit (delivery is at-least-once, below) |
| `x-source` | the emitter: the module, context or node name |
| `x-node` | the node it was emitted from |
| `x-time` | emit time, RFC-3339 |
| `content-type` | `application/json` |
**Optional headers**
| header | meaning |
|---|---|
| `x-causation-id` | the event or command that caused this one — tracing |
| `x-schema` | a version of the body's shape, so a body evolves without silent misreads |
The routing key already carries the type; it is not duplicated as a header. An **unknown `x-`
header is ignored, not refused** — unlike a declaration, an event is observed by parties that need
not all understand every header, and refusing would couple every consumer to every emitter's
additions.
### Messages are persistent
Events are published persistent (delivery-mode 2). An audit trail that loses events on a broker
restart is not one, and the cost is disk the broker already spends on everything durable.
### Queues: one per consumer, durable, dead-lettered
- **A consumer's queue** is `<node>.<module>.events`, durable, bound to that module's consumed
patterns. Durable so a restart does not drop what arrived while it was down. **Manual ack** after
the handler succeeds — at-least-once.
- **Prefetch** bounds in-flight work (default 32) so one slow consumer does not pull the whole
backlog into memory.
- **A dead-letter exchange** `mesh.events.dead` receives a message rejected past a redelivery limit,
so a poison event is set aside for inspection rather than looping forever or vanishing silently.
- **The audit logger's queue** `<node>.audit-logger.events`, bound to `#`, is the same shape —
durable, persistent, dead-lettered — because completeness is its whole job.
- **RPC reply queues** are exclusive, auto-delete and server-named; **RPC serve queues**
`serve.<key>` are durable and shared, so several runtimes serving one tool key compete rather than
each answer.
### At-least-once, and consumers are idempotent
A handler may see an event twice — a redelivery after a crash between doing the work and acking.
Consumers must be idempotent, and `x-event-id` is what makes dedup possible. **Exactly-once is not
offered**: it is a promise no broker keeps honestly, and saying so is better than pretending.
## Consequences
- The event shape is a versioned, enforced contract, not conventions each module reinvents. The
sdk's `emit`/`on` and the runtime's AMQP binding implement it; a module never sees an exchange or
queue name.
- Metadata-in-headers means the body is exactly the domain payload, and a consumer that only wants
provenance never parses it.
- Adding a header or an origin root widens the contract and is reviewed as one — the discipline
[ADR 0010](0010-delivery.md) applies to the
declaration vocabulary.
- The sdk's first cut carried source/node/time in the *body*; this supersedes that — they move to
headers. That is code to align, in `mesh-sdk` (`emit`/`on`) and `mesh-tools` (the binding, queue
config, dead-letter).
## References
- [ADR 0041](0041-events-are-a-relationship.md) — events as a relationship; this is their wire shape.
- [ADR 0010](0010-delivery.md) — the precedent: a wire
contract, versioned, additions reviewed as security.
- [ADR 0002](0002-nodes-communicate-over-a-broker.md) — the broker.
- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — `emit`/`on` are stable sdk surface; the
binding, queue config and dead-letter are the runtime's, not the sdk's.
@@ -0,0 +1,101 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0041-events-are-a-relationship.md
---
# 43. A module's broker account is scoped by what it emits and consumes
## Context
[ADR 0041](0041-events-are-a-relationship.md) made events a relationship — `emits` and `consumes`
on the manifest. [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) gave them a wire shape — the
`mesh.events` exchange, the durable per-consumer queue, the reserved routing-key origins. Neither
said how a module *reaches* the broker: what account it holds, and what that account is allowed to
do.
As the code stands, there is no answer. The mesh can provision a **node** account (at enrolment)
and a **builder** account (scoped to the build queue), and it can *deliver* any module a sealed
own-secret at a declared path — but it has no way to provision a broker **account** for a general
module. A module that declares `own-secrets: {broker: …}` and nothing more receives thirty-two
random bytes, not a credential. So on the broker, `emits` and `consumes` are enforced by nothing: a
running module could bind any queue, consume any pattern, and publish under any origin, and the
manifest that says otherwise would be describing a boundary no code draws — the exact shape of fault
[04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) records, a scope
declared in manifests and read by nothing.
This settles it, so a module's place on the bus is a thing the broker enforces rather than a thing
the manifest merely claims.
## Decision
### A module gets a broker account when it is assigned, and its permissions are the manifest
When the mesh assigns a module to a node it provisions a broker account for that module on that node,
sealed to the node ([ADR 0004](0004-a-node-and-how-it-joins.md)) and delivered as the
module's `own-secrets` broker — `amqps://` with the mesh's fingerprint, the shape
[ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) already carries. The account's permissions are
derived from the manifest, and are exactly these:
- **What it consumes.** Read on `mesh.events`, and configure-and-read on its own queue
`<node>.<module>.events` bound to the patterns in `consumes`. It cannot bind or read another
module's queue. A module that consumes nothing gets no read on the events exchange at all.
- **What it emits.** Write to `mesh.events`, restricted to routing keys under its own origin,
`module.<name>.*`. It cannot publish as another module, and cannot publish under the reserved
`mesh.*` or `node.*` origins — those belong to the mesh and the host (ADR 0042). A module that
emits nothing gets no write.
- **Nothing else.** The events account reaches `mesh.events` and that module's own queue, and no
more. Tool serving and calling over `mesh.rpc` is a separate grant on the same principle — a
module serves the tool keys it declares and calls the ones it is bound to — and is scoped the same
way rather than folded in here.
### Consuming everything is a privilege, granted deliberately
`consumes: ["#"]` — the audit logger — is read across the whole bus: every module's events, the
mesh's, every node's. That is not a pattern like any other; it is the power to see everything, and
the account is where it becomes visible. The grant that lets one module read the entire bus is one
the mesh issues on purpose and can be audited — the answer to *who can read everything* is a row, not
a guess — rather than a breadth any manifest acquires by typing a single character. A `#` consume is
a reviewed grant, not a default one.
### The account is how the declaration is enforced
Because the account can do only what `emits` and `consumes` name, the broker itself refuses a module
that tries to consume a queue it did not declare or emit under an origin it does not own. That is what
makes an event relationship a rule and not a comment — the discipline that a stated rule says how it
is checked. A manifest that over-declares grants more than the module uses, which is visible and
reviewable; one that under-declares makes the module fail closed at the broker, which is the safe
direction to be wrong in.
## Consequences
- The control plane gains a **generic module broker-account**, derived from the manifest. The
builder stops being a special case: its access to the build queue becomes an ordinary expression of
what it consumes and serves, not a bespoke account method. One rule, and the builder is an instance
of it.
- The runtime reads its credential from a file (the broker own-secret), `amqps://` verified against
the mesh's fingerprint. The `guest` account is for raising the substrate, never for a module — a
module documented as holding its own credential and handed the broker's administrative one is worse
than one with no credential story at all.
- `emits` and `consumes` stop being advisory. They are the module's authority on the bus, so the
manifest is now a security boundary and is reviewed as one, the discipline
[ADR 0010](0010-delivery.md) applies to the declaration
vocabulary.
- *Who can read the whole bus* becomes an answerable question, because `#` is a grant and not an
accident.
## References
- [ADR 0041](0041-events-are-a-relationship.md) — events are a relationship; this scopes the account
by that relationship.
- [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) — the wire this account secures: the queue,
the origins, the `amqps` credential shape.
- [ADR 0004](0004-a-node-and-how-it-joins.md) — the link is the security boundary; a
module's account is sealed to its node the same way a node's is.
- [ADR 0010](0010-delivery.md) — a declaration is owned
and its additions reviewed; a module's broker permissions are that discipline applied to the bus.
- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — a scope declared
in manifests and enforced by no code: the fault this decision closes for events.
@@ -0,0 +1,94 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0027-a-provision-names-what-the-consumer-is-coupled-to.md
---
# 44. A public name is provisioned, not registered by hand
## Context
The mesh names and resolves its own machines internally: the overlay generates
`<service>.<node>.<suffix>` wildcards, dnsmasq answers them (`wildcard-resolution`), and the mesh
issues a certificate for each internal name. A service reachable at a *public* domain —
`plex.example.com`, not `plex.anchor.internal` — needs three things that machinery does not give it:
- a **public DNS record** at a registrar or DNS provider, so the name resolves on the internet;
- a **publicly-trusted certificate** for it, because the mesh's own authority is trusted by nobody
outside the mesh;
- and routing from that name to the module — which the reverse proxy already does: a module
`requires` the `route` capability and the proxy provides it, routing by the host it was asked for.
The routing exists. The public DNS record does not: the mesh has no way to make a name resolve on
the public internet, so today that is a step someone does by hand at a DNS provider, outside the
mesh, remembered nowhere. A public name is therefore the one part of reaching a service that the
declaration graph cannot grant or withdraw — which means it is created once and outlives whatever it
was for, the shape of drift this project exists to remove.
## Decision
### A public name is a capability, requested like any other
A module reachable at a public host declares `requires: ["public-dns"]` and contributes the hostname
it wants — beside `requires: ["route"]`, which exposes it through the proxy. The name is then
provisioned on declaration ([ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md)): created
when the module is assigned, removed when it is withdrawn, reconciled like every provision.
### The interface is neutral; the providers are the registrars
`public-dns` is drawn at the consumer's coupling: the consumer wants *a public name that resolves to
me*, and does not care whether Cloudflare, Route 53 or a registrar's own API puts the record there.
So the interface is neutral and the providers are provider-scoped — `cloudflare-dns`,
`route53-dns`, `porkbun-dns` — each implementing the one `public-dns` contract, the same way a
neutral database coupling is answered by `postgres-database` and `mssql-database`. A module names
`public-dns`; it never names a registrar.
### The record points at the mesh's public ingress, not at the node
What the name resolves to is the address the reverse proxy answers on, not the consuming machine's.
A public service is reachable only *through* the proxy — the proxy holds the `route` grant and routes
by host to the module — so the public name must resolve to the proxy. `public-dns` and `route` are
the two halves of one public exposure: the name, and what the name reaches.
### The record is a fact, not a secret
A DNS record is public by definition, so the grant returns the fully-qualified name and its TTL and
nothing sealed. The only secret is the provider's own API credential, which is the provider module's
own-secret and never leaves it — the module that wanted the name never sees it.
### Events
The provider emits `module.<provider>.record.created` and `module.<provider>.record.removed`
([ADR 0041](0041-events-are-a-relationship.md)), so *which names the mesh publishes, and where* is a
question answered from the event trail and the grants, not from a folder of records edited at a
provider.
### The public certificate is the proxy's, and is named here only to pair it
A public name without a publicly-trusted certificate is reachable and not trusted — the same pairing
the internal name and the mesh-issued certificate already have. Obtaining that certificate (ACME
against the now-resolving public name) is the reverse proxy's to do, and its mechanism is its own
decision; it is named here so the pairing is not forgotten, not resolved here.
## Consequences
- A public name is created and torn down with the module, so it cannot outlive it, and the mesh can
say which public names it publishes without anyone reading a registrar's dashboard.
- Adding a registrar is adding a provider that answers `public-dns`; the modules that want names do
not change.
- Public exposure of a service is a trio of separate, declared, enforced relationships: the firewall
opens the proxy's public port ([ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)),
`route` routes the host to the module, and `public-dns` makes the host resolve.
## References
- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — a capability is provisioned on
declaration; a public name is one.
- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the firewall, the other
half of the reachability question this was asked with.
- [ADR 0041](0041-events-are-a-relationship.md) — the provider's record events.
- [ADR 0040](0040-what-a-module-is.md) — a provider and its interface; the neutral-interface,
scoped-provider naming this follows.
@@ -0,0 +1,93 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0005-the-node-host.md
---
# 45. A machine's firewall is the sum of what its modules listen on
## Context
The reverse proxy is a *provider*: a module `requires` the `route` capability and a running proxy
provides it, routing traffic by name and reaching back to the consumer. A fair question follows —
is the firewall the same shape? Should a module *register* a port with a firewall provider the way
it requests a route?
It should not, and the difference is the point. A reverse proxy is a service another component
performs; a firewall is a property of the machine — a packet filter the host applies to itself.
Modelling it as a provider would invent a credential and a reach-back for something that has neither.
And the mesh already has the registration: a module declares `listens: [{ port, from }]` — the port
it accepts connections on, and from where. That *is* how a service says it wants a port open. What is
missing is not a model but enforcement. [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md)
records that a `scope:` key five manifests carry is read by no code: a manifest can appear to
restrict a port and restrict nothing — the exact fault
[how-we-build.md](../00-META/how-we-build.md) names, *an unenforced rule is indistinguishable from a
wrong one*, made worse because the declaration reads as a restriction.
## Decision
### The firewall is derived and host-applied, not a provider
A machine's firewall is the sum of what the modules assigned to it declare they listen on, computed
by the host and applied as one of its owned resources ([ADR 0005](0005-the-node-host.md):
the host applies, it does not decide; [ADR 0010](0010-delivery.md):
the declaration is owned resources). It is not a capability, not a per-consumer grant — opening a
port is a declarative fact about a machine, so it is computed and applied, not requested and
credentialed.
### `from` is the whole of public-versus-internal
The distinction the question is really about lives in `from`:
- `listens: [{ port: 5432, from: mesh }]` — open to the private overlay only.
- `listens: [{ port: 443, from: anywhere }]` — open to the public internet.
A module registers a port on the firewall by listening on it and saying from where. There is no
separate firewall capability, because the firewall is not a thing that reaches back or holds a
secret; it is the machine's own filter over the ports its modules named.
### The host enforces it both ways, and unknown keys are refused
A port a module listens on is opened to exactly the scope it named; a port nothing declares is
closed. And a key the firewall does not read — the `scope:` of issue 003 — is refused at the
manifest, not accepted and ignored, so a declaration that reads as a restriction is one. This is the
discipline [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) applied to the
broker account, applied here to the packet filter: the declaration is the enforcement, or it is a
comment.
### A public service is exposed through the proxy, not by opening its own port
Reaching the public internet is normally not `from: anywhere` on the service's own port. The service
listens `from: mesh` — only the proxy reaches it — and `requires: route`, so the sole machine with a
public opening is the one running the reverse proxy, and the service is exposed by name through it.
`from: anywhere` is the deliberate direct-exposure case, for a service that is its own front door.
## Consequences
- Issue 003 is closed: the firewall is computed from `listens` and enforced, so a declared scope is
real and an undeclared port is shut. Rejecting unknown manifest keys is the general fix, of which
the `scope:` key was one instance.
- The firewall and the reverse proxy stop being confused for one model: the firewall is the machine's
filter (host-derived from `listens.from`); `route` is a name-router (a provider); the public DNS
name is a third thing ([ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md)). A
public service uses all three.
- The modelling question is answered: a module registers a port by declaring `listens`, and reaches
the public internet by name through `route` + `public-dns` — never by the firewall being a
provider.
## References
- [ADR 0005](0005-the-node-host.md) — the host applies; the firewall is one of
the things it applies.
- [ADR 0010](0010-delivery.md) — the firewall is a derived
owned resource, not a grant.
- [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) — the same discipline:
a declaration is enforced, or it is a comment.
- [ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md) — the public name, the other
half of the reachability question this was asked with.
- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — the unenforced
`scope:` this closes.
@@ -0,0 +1,88 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0027-a-provision-names-what-the-consumer-is-coupled-to.md
---
# 46. A module's configuration is its assignment's, not its manifest's
## Context
A module is assigned to a node — `assign <node> <module>`, always to a machine; there is no
assignment to the mesh. "Mesh" is a *scope*, not a place: a `provides` or a `claim` scoped `mesh`
reaches the whole mesh, but the module still runs on a node. So the two kinds of thing a module can
carry are the manifest (what the module *is*) and, separately, what it should do *here* — which
differs by deployment and by node.
The mesh already has the second: **settings**. `settings set <module> [--node <node>]` — with a node
it is that machine's, without it the whole mesh's — layered over what the module declares and applied
at resolution, changeable without editing the module and without a rebuild. That is the surface a
meshboard would edit.
But settings today reach only a module's **config-file content** (a mergeable file the module owns).
Configuration that is not a file has been landing in the manifest instead, statically — a registrar's
zone and domain, the address public names point at, and, most sharply, `listens.from`. That last one
is the tell: whether a port is open to the private overlay or to the public internet is a
*per-node deployment choice* — the same database internal on one machine and public on another — and
a value fixed in the manifest is one value for every machine, so it cannot be. Static configuration in
the manifest is configuration in the wrong place: it cannot vary per node, and it cannot change
without a new module version.
## Decision
### The manifest is identity and defaults; the assignment's settings are the configuration
A module's manifest declares what it is — what it provides, requires and claims, the shape of its
resources — and, for anything configurable, a **default**. The values that make a running instance
*this* instance are settings, carried by the assignment: per-node, or mesh-wide when no node is named,
applied over the defaults at resolution. Change one and the next reconcile carries it; nothing is
edited on a machine and nothing is rebuilt.
### Settings drive the configurable fields the manifest marks, not only file content
Settings extend beyond a config file's content to the manifest fields a module declares settable —
foremost:
- **`listens.from`**: a module declares its safe default (`from: mesh`), and a per-node setting
raises or lowers it. postgres declares `listens: [{ port: 5432, from: mesh }]`; on the machine that
should expose it, a setting makes that port `from: anywhere`. Same module, different exposure, and
the firewall ([ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)) is computed
from the effective value, so the packet filter follows the setting.
- **A provider's own configuration**: a registrar's zone, domain and the ingress its names point at
([ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md)) are mesh-wide settings, not
manifest constants — one mesh's Cloudflare zone is not another's, and the module description is the
same for both.
### Unset is the default, and an unknown setting is refused
A field with no setting keeps the manifest's default, so a module runs correctly configured by nobody.
A setting that matches no settable field — like a config value that reaches no file today — is named,
not silently dropped, so a misspelled setting is found rather than believed (the discipline of
`UnusedSettings`, and of [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md):
a declaration is enforced or it is a comment).
## Consequences
- The postgres case works: one module, `from: mesh` by default, `from: anywhere` where a setting says
so — internal on ace, public on novox, changeable live.
- Provider modules stop carrying a mesh's specifics: `cloudflare-dns` describes *a Cloudflare
registrar*, and *which* zone and ingress is a setting, so the same module serves every mesh.
- Configuration becomes a thing a meshboard manages — set per node or mesh-wide, applied on the next
reconcile — rather than a manifest edit and a rebuild ([ADR 0011](0011-managed-files-are-generated-never-edited.md):
the way you change a managed thing is not by editing it).
- What a manifest may not do is grow a value that differs per machine; if it differs per machine it is
a setting, and the manifest holds only the default.
## References
- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — what is provisioned on
declaration; its per-instance values are the assignment's.
- [ADR 0011](0011-managed-files-are-generated-never-edited.md) — a managed thing is changed through
the mesh, not by editing it; settings are that, for configuration.
- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the firewall follows the
effective `listens.from`, so making `from` a setting makes exposure a setting.
- [ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md) — the provider whose zone and
ingress are settings, not manifest constants.
@@ -0,0 +1,87 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md
---
# 47. A module runs its code as its own process, with its own account
## Context
A module is one self-contained thing ([ADR 0040](0040-what-a-module-is.md)), and it gets a broker
account scoped to what it emits and consumes ([ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md)).
The catalogue now gives modules **tools** and **events** — real code, in the module ([ADR 0039](0039-what-the-sdk-holds-and-refuses.md)) —
but nothing has said what *runs* that code. The audit-logger showed one shape and was treated as an
exception: a container running the tool runtime carrying the module's compiled code, holding the
module's own scoped account. Every module with tools or events needs the same, and the tempting
alternative does not work.
**A node-wide runtime that loaded every assigned module's code cannot hold a per-module account.** It
would run under one account with the union of every module's permissions — able to emit as any of
them and read any of their queues — which is exactly the isolation ADR 0043 exists to draw. So the
runtime is per-module, not per-node, and treating the audit-logger as special left the other
modules' code with nothing to run it: the conversion produced tools and events that, as it stands,
never execute.
## Decision
### A module with tools or events runs a process of its own
A module that has tools or events runs a **runtime process** — a container, the tool runtime carrying
that module's compiled code — assigned and started like the module it is, holding the single broker
account the mesh scoped to it (ADR 0043). One module, one process, one account.
### It serves its tools, each on its own key
A tool is served on its own key (`serve.<tool>`), and a caller invokes a named tool. Only the module
that serves it answers, and the module's account is scoped to exactly its tool keys — so one module
cannot answer another's calls, the isolation ADR 0043 gives events extended to tools. This supersedes
a single `tools.invoke` endpoint that dispatched by name: that shape assumed one runtime for the
whole node, and per-module runtimes competing on one key would each be handed calls for tools they do
not have.
### It runs its events in the same process, under the same account
Emitting under the module's own origin and consuming its own queue ([ADR 0042](0042-the-shape-of-an-event-on-the-wire.md))
happen in that same process, with that same account — not a second one to scope and seal. A module's
tool code, its event code and, for a provider, its provisioner are the one module's code and run as
the one module's process.
### The runtime image is the tool runtime plus the module's code
Built from the module's source like any module image — the audit-logger's shape, made the rule, not
the exception. The module declares a `container` for it carrying `MESH_BROKER_FILE` (its sealed
credential, ADR 0043) and its compiled code. A module with **neither** tools nor events runs no such
process: a plain service module — the plex *server*, dnsmasq the resolver — is its service and files
and nothing more. A module that is both a service and code declares both containers: the service, and
the runtime beside it.
## Consequences
- The catalogue's tools and events become runnable: each tools-or-events module gains a runtime
container with its scoped credential, and the audit-logger stops being special. Until this, the
converted modules held code with nothing to execute it.
- A process, and a small image, per tools-or-events module. That is the cost of ADR 0043's isolation:
one account per module means one process per module. It is paid deliberately — a shared runtime is
cheaper and cannot be scoped, and a mesh where any module can emit as any other is not one worth the
saving.
- `serve.<tool>` per key replaces the single `tools.invoke` dispatch. The sdk's serving and a module's
account scope both come to name tools individually.
- **A provider's provisioner is a runtime process too.** It already runs as its own container; its
events (`bucket.created`, `database.provisioned`) belong to *that* process and need the same
credential. So a provisioner that emits carries `MESH_BROKER_FILE` and its scoped account like any
runtime — or it does not emit. (This is the fix for provisioners that emit today with no broker
bound: the emit is a runtime's, and the provisioner is a runtime.)
## References
- [ADR 0040](0040-what-a-module-is.md) — a module is one self-contained thing; its code runs as one
process.
- [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) — the scoped account
this process holds, and the isolation that makes it per-module.
- [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) — the events this process runs, and the
`serve.<key>` queue tools now use.
- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — the code lives in the module; this runs it.
@@ -0,0 +1,131 @@
---
topic: what runs on it
status: accepted
date: 2026-09-05
deciders: jochen
reconstructed: false
---
# 48. A provider creates the credential the mesh minted, and seals nothing
## Context
A provider module stands up a per-consumer resource — a database, a cache bucket, an object
store user — and the consumer must end up holding a credential that authenticates against it.
Building the module runtime (ADR 0047: a module runs its own code as its own process under its
own account), the provider's provisioner was run for the first time as a delivered thing, and
it did not work. It reads a seal key from the environment that nothing sets, and it seals every
credential it produces to that key with a symmetric passphrase.
Tracing the credential's path turned up something larger than a missing key. **The provisioner
harness the whole catalogue is built on describes a credential flow the mesh does not have, and
duplicates — incorrectly — one it does.**
What the sdk's `runProvisioner` does today:
- reads request files named `*.grant.json` — which nothing in the mesh writes;
- calls an adapter whose `create` **generates its own password** and returns it;
- seals that password with a symmetric key (`$MESH_SEAL_KEY`) and writes a `*.credential`
file — which nothing in the mesh reads, and no consumer ever unseals.
What the mesh already does, and has wired end to end:
- The control plane mints one password per (consumer, provider) pair (`Inventory.SecretFor` →
`secrets.Make`) and seals it to **both** node keys asymmetrically — a copy the consumer's
host can open and a copy the provider's host can open. No shared symmetric key exists
anywhere, on purpose: a key both ends hold is a key the mesh would have to distribute, which
is the same problem one level down, and the control plane's own code refuses it.
- The provider is handed, at the path its `receives` names, one contribution per consumer:
the **login to create** (`As`, derived by the mesh so the two ends agree by construction),
the consumer's address and requested values, and a **`Secret` file** holding that consumer's
password sealed to the provider and unsealed onto the machine by its host.
- The consumer is handed the *same* password, as plaintext its own host wrote by unsealing its
copy and substituting it into a config file. The consumer never unseals anything itself and
holds no key.
So the password a provider's provisioner invents is not even the password the consumer was
given: a consumer authenticating with the mesh's password against a resource the provisioner
created with its own would simply fail. The symmetric seal is not an incomplete feature to
finish delivering a key for. It is a second, contradictory credential model bolted beside the
real one, and it cannot be made to work without building the very thing the mesh was designed
not to have.
This is a decision and not a patch because the harness is the **provider contract**. Every
provider — the four that exist and the many a real mesh grows — is built on
`runProvisioner(resource, adapter)`. Whatever it says a provider is, they all inherit; and
changing it later is one migration per provider. It is cheaper and more honest to settle what
a provider is now.
## Decision
**A provider is handed the credential; it does not make one, does not seal one, and does not
hand one back.** The provisioner's only job is to make the mesh's grants true in its own
software.
Concretely, for the sdk harness and the adapter contract:
- The harness reconciles the **contributions the mesh delivers** to the provider's `receives`
path — the list of consumers, each with its login name (`As`), address, requested values,
and the path to its unsealed password (`Secret`). It does not read `*.grant.json` and it
does not write `*.credential`.
- For each consumer present, the harness reads the password from that consumer's `Secret` file
and calls the adapter to bring the resource into being under the given login. For each
consumer no longer present — the mesh drops it from the contributions file when its consumer
goes away — the harness calls the adapter to withdraw it.
- The adapter shrinks to the per-software half and nothing else. It is given the login, the
password, and the values, and it makes the resource exist or removes it. It generates no
password, derives no name, seals nothing, and returns no credential:
roughly `create({ as, password, values })` and `remove({ as })`, both returning nothing.
- `$MESH_SEAL_KEY`, the symmetric `seal()`/`writeSealedCredential` path, and the `*.grant.json`
/ `*.credential` files are removed from the provisioning path entirely. The credential
reaches the consumer through the mesh's own asymmetric channel, which already crosses node
boundaries and holds no shared secret.
Identity stays the mesh's to say. The login the provider creates is the name the mesh derived
and gave the consumer to present; the provider never invents a name, because a name the
consumer cannot learn is a name it cannot authenticate with.
## Consequences
- A provider module becomes smaller and unable to be wrong in this way: with no password to
generate and no key to seal to, the class of bug where the two ends hold different secrets
cannot be written. A provider added after this inherits the corrected contract and has no
seal to reintroduce.
- The four current providers (redis, postgres, minio, umami) each lose their `generatePassword`
+ seal code and gain a `create` that takes the password it is given. Their teardown becomes
"withdraw the login named `As`".
- The symmetric `seal()`/`unseal()` primitive loses its only caller and leaves — checked, not
assumed: nothing else in the sdk or the catalogue called it, so it is removed with the
provisioner it belonged to.
- **How this is verified:** redis is assigned as a provider in the lab, the contributions and
the unsealed password the mesh would deliver are put in its `receives` path, and a client
authenticates as that consumer with the mesh's password and gets PONG — where a provider that
invented its own password answers WRONGPASS — with `$MESH_SEAL_KEY` set nowhere and no
`.credential` file written. Proven: `provider-uses-mesh-credential` is green.
**What this does not cover — credential provisions, not data provisions.** This decision is about a
provision whose credential is a *secret the mesh mints* — a login and password (redis, postgres,
minio). A provider that instead *generates* the thing the consumer needs, and that thing is not a
secret — umami's `analytics`, where the consumer wants back a `siteId` umami assigned — does not fit,
because a contract that returns nothing has no way to hand that data back. The seal-key fault was
never umami's (it sealed no password; it returned a public id), so removing the seal does not break
it further, and it still reconciles its sites off the mesh's contributions. But delivering
provider-generated data back to a consumer is a *return path* the mesh does not have and this
decision does not build — a separate shape, left to a separate decision.
- Teardown beyond "remove the login" — data an object store leaves behind when a consumer
leaves — is named by each provider's adapter, not by the harness, and is out of scope here
except to say the contract must leave room for it.
## References
- [04-ISSUES/032](../04-ISSUES/032-provider-runtime-has-no-seal-key/00-report.md) — the
observation and the cross-repo trace this decision rests on.
- ADR 0047 (the module runtime) — what first ran a provider's provisioner as a delivered
process and exposed this; link to be filled when 0047 lands on the trunk.
- Control-plane mechanisms this relies on already existing: `mesh-control` —
`internal/inventory/secrets.go` (`SecretFor`, `SecretsFrom`), `internal/secrets/seal.go`
(`Make`, the two-blob asymmetric sealing), `cmd/mesh-control/plan.go` (`grantsFor`, the
`Grant.Sealed = ForProvider` delivery), `internal/catalogue/declaration.go` (the `receives`
contribution: `As`, `At`, `Values`, `Secret`).
- The path being removed: `mesh-sdk` — `src/provisioner/index.ts` (`runProvisioner`, `sealKey`,
`writeSealedCredential`) and the symmetric `src/primitives/index.ts` `seal()`/`unseal()`.
@@ -0,0 +1,130 @@
---
topic: what runs on it
status: accepted
date: 2026-09-05
deciders: jochen
reconstructed: false
---
# 49. A consumer's identity is bounded by the tightest backend that must accept it
## Context
The mesh says who a consumer is, once, and hands the same name to the provider (to create) and the
consumer (to present), so the two ends agree by construction rather than by two conventions (the
principle behind `ConsumerIdentity`, 04-ISSUES/023). The name is `mesh_<node>_<module>`, cleaned to
lower-case letters, digits and underscore.
Proving the provider contract per backend (ADR 0048) turned up 04-ISSUES/034: redis and postgres
create that name verbatim, but **minio refuses it** — an S3 access key is capped at 20 characters,
and `mesh_anchor_bucketuser` is 22. The provisioner then retries for ever, per consumer, and the
consumer holding that same too-long name could never present it either.
Two things about the existing derivation decide most of this:
- **The charset is already right.** `[^a-z0-9_]` is deliberately conservative, and its own comment
says it reaches "a PostgreSQL role, a MinIO access key, an LDAP uid and a Keycloak client without
quoting." That much is true.
- **The length is wrong.** `CheckIdentity` refuses names over `identityLimit = 63`, commented as
"the shortest identifier limit among the systems these names reach: PostgreSQL's". It is not the
shortest — S3's 20 is shorter — so the guard that was meant to catch exactly this lets it through,
and the failure lands at provision time as a silent retry instead of at assignment as a refusal.
So this is a small wrong constant with a real cost attached: whatever bound we set, `mesh_` (5) plus
a node name plus `_` plus a module name has to fit inside it.
## The options
**A — Bound the identity by the true minimum, and refuse early.** Lower `identityLimit` to the real
shortest (20, S3's), so `CheckIdentity` refuses an over-long name *at assignment* with a clear
message, the way it already refuses over-63 names. The derivation does not change; long names are
simply rejected before anything is provisioned.
- *For:* smallest change; keeps "the mesh says the identity once, verbatim" intact; the failure
moves from a per-consumer provision-time retry to an up-front, legible refusal — which is what
`CheckIdentity` exists to do.
- *Against:* a hard budget. `mesh_` + node + `_` + module ≤ 20 means node + module ≤ 14 characters.
`anchor` + `bucketuser` (16) is already over. It pushes the constraint onto how machines and
modules are named, which is a real limitation on legible names.
**B — Keep the readable name when it fits, compact it when it does not.** Below the bound, the name
is `mesh_<node>_<module>` as today; over it, the mesh substitutes a deterministic short form (e.g.
`mesh_` + a truncated hash of node+module) — still one derivation, so both ends still agree.
- *For:* no naming constraint; short backends always satisfied; the common case stays legible.
- *Against:* some identities become opaque, and a provisioner tracing "whose login is this" loses
the answer for exactly the consumers that overflowed. The mesh now owns a fallback format and its
collision properties (a truncated hash is not free of collisions at 15 characters).
**C — Let each interface declare its identifier bounds, and derive within the tightest a consumer
reaches.** `s3-bucket` states `identifier: { max: 20 }`; `postgres-database` states 63; the mesh
derives a name that fits the **minimum** bound across the providers a given consumer is granted.
- *For:* the most precise — each provision gets exactly the room it has, and a database consumer
keeps long legible names while an S3 consumer gets a short one; the constraint lives where the
fact does (on the interface).
- *Against:* the most work, and a consumer of two interfaces with different bounds must satisfy the
smaller — so its name shortens for both, reintroducing B's opacity in a narrower case. It also
means one consumer can hold **different** identities per provision, which the "said once" model
currently forbids.
**D — Let the provider generate a backend-valid identity and hand it back (rejected).** minio mints
its own access key and returns it to the consumer. This is the data-provision return path this era
keeps meeting — but it directly contradicts 023 and ADR 0048: the identity would no longer be the
mesh's single derivation the two ends share, it would be a value one side invents and the other must
be told. Listed for completeness; not recommended.
**E — A module (and a node) may declare a short slug; the identity is built from it.** The identity
becomes `mesh_<node-slug|node-name>_<module-slug|module-name>`: where a slug is declared it is used,
otherwise the cleaned name. A slug is a deliberately short, operator-chosen identifier — `kc` for
keycloak, `wkstn` for a workstation. It is optional: short names (`anchor`, `redis`) need none.
- *For:* this is the escape hatch B wanted to be, without the opacity. The name stays legible — a
provisioner can read `mesh_wkstn_kc` and know who is asking — because a person chose it, not a
hash function. And it makes an early refusal *palatable*: if even the slug-built identity overflows,
the refusal points at the slug, a field made for exactly this, rather than at the machine's name.
Both ends still derive it from one declared thing, so they agree by construction.
- *Against:* a new optional manifest field, and someone must pick the slug — but only for names that
would otherwise overflow, and picking a short legible identifier is a better job than being handed
a hash.
## What implementing A revealed
A was tried first. At `identityLimit = 20`, the readable budget is `mesh_` (5) + node + `_` + module
≤ 20, i.e. **node + module ≤ 14 characters** — far tighter than it looked. The catalogue's own
existing tests use `workstation`+`keycloak` (25), which compacts to `mesh_dbbc02f8dde34d3`; common
mesh names (`home-server`, `the-build-node`, `workstation`) blow the budget with any module. So B's
compact fallback would fire for the *common* case, not the rare overflow — which inverts A+B: most
identities would be opaque hashes. A alone (hard refusal at 20) would refuse most realistic names.
This is what moved the recommendation to E: the problem is not the limit, it is that the *readable
name* is the wrong source when it is long, and a slug is a better source than either a hash or a ban.
## Recommendation
**E, over a per-consumer bound (start with the global minimum, 20).** Build the identity from an
optional slug, keep it when it fits, and refuse at assignment with "declare or shorten `<module>`'s
slug" when it does not — no hash, no lost legibility, and the fix is a first-class field. Set the
bound to the true minimum (20) now; it needs no per-interface machinery to unblock S3, and a module
that consumes S3 simply declares a short slug. Graduate to **C** (per-interface bounds) later if it
turns out that non-S3 consumers are paying for S3's limit often enough to mind — E and C compose:
slugs are the mechanism, per-interface bounds refine where the ceiling sits. **B is dropped**: a
declared slug is a strictly better escape hatch than an opaque hash. **D stays rejected.**
## Consequences (of E)
- A module manifest gains an optional `slug`; a node may carry one too. `ConsumerIdentity` prefers
the slug over the cleaned name for each half. `identityLimit` becomes 20 (the true minimum), and
`CheckIdentity` refuses at `module add` / assignment — now with a message naming the slug to set.
- The common case stays legible; only names that overflow the budget need a slug, and what they get
is a name a person chose, not a hash.
- Existing modules/nodes whose names overflow declare a slug once — a migration cost paid as a clear
refusal with an obvious remedy, not a silent hash or a silent truncation.
- minio (04-ISSUES/034) is unblocked: an S3 consumer declares a short slug and its access key fits.
- **How it is checked:** the minio grant e2e — a consumer whose (slugged) identity fits reaches its
bucket with the credential the mesh delivered — plus unit tests that a slug is preferred, that an
un-sluggable over-long identity is refused (naming the slug), and that two consumers never collide.
## References
- [04-ISSUES/034](../04-ISSUES/034-mesh-login-exceeds-s3-access-key-limit/00-report.md) — the
observation.
- ADR 0048 — a provider creates the credential the mesh minted; the identity it creates it under is
the one this decision bounds.
- `mesh-control` `internal/catalogue/identity.go` — `ConsumerIdentity`, `identityUnusable`,
`identityLimit`, `CheckIdentity` — where the constant and the check live.
@@ -0,0 +1,206 @@
---
topic: what runs on it
status: accepted
date: 2026-09-05
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0024-model-access-is-a-provision.md
---
# 50. Model access is vendor-agnostic, and a vendor is an adapter
## Context
[ADR 0024](0024-model-access-is-a-provision.md) settled that model access is a provision and that
a licence is a named thing an operator uses. What shipped, and runs, is a single vendor: the mesh's
"claude" feature. A read-only trace of that feature (2026-09-05, in the code workspace) was made to
answer whether the model-access provision is Anthropic-shaped or genuinely general. The finding is
that **the vendor-agnostic layer already largely exists**, and the Anthropic specifics are a thin
band around it that a per-vendor adapter can hold.
**What is already general, with evidence.** `mesh-control internal/licences` models
`licence(name, vendor, serves)` and `licence_holder(licence, node, module, sealed)`, and each
holder's credential is sealed per-holder through `internal/secrets`. `serves` carries the
non-secret facts (a base URL, a model) and is not vendor-specific. The `accept` verb
([ADR 0024](0024-model-access-is-a-provision.md), and
[`14-model-access`](../03-DESIGN/01-to-be/14-model-access.md)) already takes an operator-supplied
value, seals it to each holder and discards the plaintext. A model the mesh runs itself answers
`model-access` at node scope with no licence at all. None of that mentions Anthropic.
**What is Anthropic-specific.** The credential is not a static key: it is a subscription OAuth grant
— an hourly access token plus a refresh token. That shape drags four things behind it that a static
key does not need: **central rotation** (one manager node refreshes under a lease and publishes the
new token), **delivery that strips the refresh token** so a consuming node holds only an access token,
an **identity guard** that reads the credential to catch a mis-binding, and a **usage** reading with
Anthropic's own `utilization%` semantics. Most vendors are a single static key, which the sealed-key
model already handles and which needs none of these four.
**The tension at the centre of this.** A refreshable credential cannot be both *sealed so the mesh
cannot read it* and *rotated centrally*. Central rotation means some node in the mesh holds the
refresh token in readable form, because that is what refreshing requires. Per-holder sealing means no
node but the holder can read the credential. For a static key the two never meet — there is nothing to
rotate. For a refreshable grant they collide directly, and this record exists to say which gives way,
and by how much.
## Considered Options
1. **Keep Anthropic special-cased in the core.** Leave the three binding columns and the `claude_*`
schema, and add other vendors beside them the same way. **Rejected.** It is exactly what
[ADR 0024](0024-model-access-is-a-provision.md) ruled against: a module that names a vendor cannot
be moved onto another model without editing it, and moving it is the point. It also grows the core
by one band per vendor, when the bands are the same shape.
2. **One provision, and refuse to hold any refresh token — re-seal only.** Make every credential
purely sealed per-holder, including refreshable ones; let each holder refresh its own grant.
**Rejected.** It throws away the hard half [ADR 0024](0024-model-access-is-a-provision.md) says
already works — the lease, the single-refresher, the switch-on-exhaustion — and replaces it with N
nodes each holding a refresh token, which is the very thing today's delivery strips on the stated
ground that *a node never holds a refresh token*. A refresh token is the long-lived secret; spraying
it across every holder is strictly worse than keeping one copy on one node.
3. **One provision, and abandon central rotation entirely** for refreshable vendors — treat the grant
as opaque and let it expire. **Rejected.** For a subscription-seat vendor an expired access token is
a dead licence; without rotation the feature that works today stops working. This is option 2's cost
without option 2's autonomy.
4. **One vendor-blind provision, plus a per-vendor adapter, with a bounded carve-out for the
refreshable case.** **Adopted**, below.
## Decision
**`model-access` is one consumer-facing, vendor-blind provision.** A consumer declares
`requires: model-access`, and is coupled to *reaching a model* — a base URL, a model name, a key —
and not to which vendor answers. That is the coupling the name is drawn at
([ADR 0040](0040-what-a-module-is.md)'s rule: name the interface at the widest boundary across which
the consumer does not care which implementation serves it). Where a consumer were genuinely coupled to
a specific wire API it could not swap across, the same rule would split the name — but the consumers
that exist reach their model through a CLI or SDK that hides the vendor, so `model-access` is the true
coupling and stays one name. This extends [ADR 0024](0024-model-access-is-a-provision.md) and
[ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) without changing them.
**A vendor is an adapter, keyed by the licence's `vendor` field.** The lifecycle a licence needs is
vendor-specific and lives in a per-vendor adapter selected by `licence.vendor`, exactly as
`public-dns` is one neutral interface answered by registrar-scoped providers —
`cloudflare-dns`, `route53-dns` ([ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md)).
A consumer names `model-access` and never a vendor, the same way a module names `public-dns` and never
a registrar.
**The field is named `vendor`, not `provider`.** The inventory already uses "provider" for the
provider-pin — *which node answers a brokered provision*. Reusing it for *which company sells this
licence* would collide two unrelated facts on one word. `vendor` is the licence's, and is separate.
### The adapter's capabilities, all but one optional
An adapter declares:
- **`shape`** — `static-key` or `refreshable-grant`. This is the switch the carve-out below turns on.
- **`accept(value) → sealed`** — take an operator-supplied credential and seal it to the holders, the
`accept` verb [ADR 0024](0024-model-access-is-a-provision.md) already defines.
- **`refresh(licence)`** — refreshable-grant only: the lease / rotate / publish machinery.
- **`identity(credential) → account-id`** — the mis-binding guard, for a vendor whose credential
carries an identity worth checking.
- **`usage(licence) → normalised rows`** — the vendor's usage reading, mapped to the common shape below.
- **`deliver`** — the credential *value* only; the destination path is the consumer's, not the
adapter's.
**A static-key vendor implements almost nothing** — `shape: static-key`, `accept` is the generic
seal, `deliver` is the value, and `refresh`, `identity` and `usage` are absent or trivial. The
abstraction earns its keep by making the common vendor small, not the rare one clever.
### The carve-out — the one place the guarantee is relaxed, said plainly
The mesh's standing principle is that it cannot read what it stores: `accept` seals to the holders and
discards the plaintext ([ADR 0024](0024-model-access-is-a-provision.md)), and a provider seals nothing
because the credential travels the mesh's own asymmetric channel
([ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)). A `refreshable-grant`
credential cannot honour that principle and be centrally rotated at the same time, and central rotation
is the working half [ADR 0024](0024-model-access-is-a-provision.md) is explicit about keeping.
**So, for `refreshable-grant` vendors only:**
- the **manager node holds the refresh token encrypted at rest** — readable by that node, because
rotation requires it. This is the bounded exception.
- **access tokens are still sealed per-holder**, as every credential is; a holder reads its own and no
other node reads it.
- the **refresh token is stripped on delivery** — it never reaches a consuming node. *A node never
holds a refresh token* stays true for every node but the one manager.
**Static-key vendors keep the full guarantee.** There is no token to rotate, so there is nothing to
hold readably, so `accept` discards the plaintext and the carve-out never fires. The majority of
vendors are static-key, and the majority therefore lose nothing.
The exception is stated rather than hidden because a relaxed guarantee that is not written down is
indistinguishable from a broken one. It is bounded on three axes at once: **refreshable-grant vendors
only, the refresh token only, the manager node only.**
### The settled details this record also fixes
- **Usage is normalised to `(licence, consumer, period, metric, value)` plus the raw response as
jsonb.** The metric is vendor-defined — Anthropic's `utilization%` is one metric, a token count is
another — and no common unit is forced across vendors. The raw response is kept so a reading can be
re-derived if the normalisation is later found wrong.
- **Binding is explicit per consumer, and an unchosen consumer is refused — no implicit fallback.**
This is the direction the resolver already takes, and it is the safe one: a mesh with several ways
to reach a model refuses a consumer that has not said which, naming the candidates and the command,
rather than silently choosing one ([ADR 0024](0024-model-access-is-a-provision.md), and
[`14-model-access`](../03-DESIGN/01-to-be/14-model-access.md)).
- **Subscription-seat authentication lives entirely inside the adapter**, never in the generic core.
So does an interactive `/login` — an adapter-specific "adopt" origin for a credential a person must
produce in a browser; the generic `licence key <name>` covers the static-key case.
- **Anthropic is the first `refreshable-grant` adapter**, carrying the OAuth refresh, the usage
reading, the identity guard and access-token-only delivery. **`anthropic-api-key` is a
`static-key` adapter for the same vendor's plain API keys**, and is the early second case that
proves the abstraction is not a single vendor wearing a coat: it exercises the whole path with the
carve-out switched off.
## Consequences
- **The carve-out is the mesh's one deliberate relaxation of "it cannot read what it stores."** It is
bounded to refreshable-grant vendors, to the refresh token, and to the manager node; the static-key
majority keep the full guarantee unchanged. This is the open risk the analysis carried here, and it
is recorded as an exception rather than pretended away.
- **Anthropic collapses from special case to adapter.** The three binding columns
(`nodes.node_license`, `nodes.hal_claude_account`, `agents.claude_account`) become three ordinary
consumers of `model-access`; the `claude_*` schema becomes the generic licence tables plus one
adapter. What was hardcoded becomes data keyed by `vendor`.
- **Adding a vendor is adding an adapter, and a static-key vendor is nearly free.** The modules that
want a model do not change when a vendor is added — they named `model-access`, not a vendor.
- **The refresh-token concentration is now a stated property to defend, not an accident.** The manager
node is a place a long-lived secret lives readably, and losing it or compromising it is a bounded,
named blast radius rather than a surprise.
### How each claim here is checked
- **Vendor-blind provision, static-key path.** A lab scenario binds an `anthropic-api-key` licence to
a consumer; the consumer resolves, receives a key sealed to its node, and reaches a model — and the
key is **nowhere in the control plane's database** nor in anything that crossed the broker. This is
the `licence_holder` sealed-per-node check that [`14-model-access`](../03-DESIGN/01-to-be/14-model-access.md)
already runs, now asserted for a second vendor.
- **The carve-out is exactly as narrow as stated.** For a `refreshable-grant` licence, a test asserts
the refresh token exists (encrypted) **only on the manager node**, is **absent from every holder's
delivery**, and that the delivered credential is access-token-only — and that for a `static-key`
licence no refresh token is stored anywhere.
- **Adapter selection is keyed by `vendor`.** A scenario with two vendors on two licences verifies each
licence's lifecycle runs its own adapter, and that a consumer naming `model-access` never names a
vendor to get one.
- **Refuse-if-unchosen.** Already checked in [`14-model-access`](../03-DESIGN/01-to-be/14-model-access.md):
a consumer with more than one candidate is refused with the candidates and the command named.
## References
- [ADR 0024](0024-model-access-is-a-provision.md) — model access is a provision, a licence is a named
thing, and `accept`; this record generalises its single vendor and keeps its working central
rotation.
- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — a provision names the
coupling; `model-access` is drawn at the consumer's.
- [ADR 0040](0040-what-a-module-is.md) — the naming rule and the neutral-interface / scoped-provider
shape a vendor adapter follows.
- [ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md) — registrar-scoped `public-dns`
providers, the precedent a `vendor`-scoped adapter mirrors.
- [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md) — the mesh seals credentials
and holds no readable copy; the carve-out here is the bounded, named exception to that for a
refreshable grant.
- [`03-DESIGN/01-to-be/14-model-access.md`](../03-DESIGN/01-to-be/14-model-access.md) — the design this
record extends, amended to describe the adapter generalisation.
- The read-only vendor-agnostic analysis, 2026-09-05 (code workspace) — the inventory and the decisions
taken on the open questions this record encodes.
@@ -0,0 +1,166 @@
---
topic: what runs on it
status: accepted
date: 2026-09-05
deciders: jochen
reconstructed: false
extends: 0030-data-outlives-the-mesh-that-declared-it.md
---
# 51. Shared data is the operator's, and a module is granted access to it
## Context
**Eight modules declared one filesystem as eight private ones.** The media stack — a library
server, the acquisition managers for films, series, music and books, a subtitle fetcher and two
download clients — shares directories on one machine: the download clients write into
`/services/media/downloads` and the managers read it; the managers write into the libraries and
the library server reads them. That sharing is the entire point of the stack. Yet each module
declared every shared directory it touched as its own `directory` resource, with an owner and a
mode. `/services/media/downloads` was written seven times, as seven private directories that
happen to be the same path.
**The resolver refuses exactly that, and is right to.** Two modules declaring one path on one node
are refused by name, with no exemption for identical content and no merge — because two owners of
one path is the class of fault this repository keeps recording ([04-ISSUES/036](../04-ISSUES/036-six-modules-own-what-they-must-share/00-report.md)).
So the stack as written refuses its own only sensible assignment: all of it on one machine,
sharing one filesystem. It passed today only because no test co-resolves any two of the eight. The
first machine assigned two of them together is where the refusal would have surfaced.
**The vocabulary had one word for two intentions, and this was already seen.**
[04-ISSUES/026](../04-ISSUES/026-the-data-directories-are-not-declared/00-report.md) found the
same gap from the other side and named it precisely: two kinds of mount are spelled identically —
*the directory my data lives in*, which the mesh creates and owns, and *a facility I was granted*,
which already exists and the mesh only reaches. That issue deferred inventing a field to tell them
apart, because doing so is a design decision and it declined to make one to get a check green. This
is that decision.
**A `directory` resource is owned, on every axis.** The host creates it, sets its owner and mode,
and removes it when it is empty and no longer declared — it *removes what it made and leaves what
it merely configured* ([ADR 0005](0005-the-node-host.md), [ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md)).
A shared media library is none of that. It existed before the mesh, several modules read and write
it at once, and losing it is the one failure that does not recover. It is not any module's
resource; it is the operator's, and a module only needs to be let at it.
## Considered Options
1. **Make the stack one module with several containers**, the way the mail system already is. The
issue raises it directly: is a set of modules that must share a filesystem really one module?
**Rejected.** The eight are independently assignable and independently useful — a person may run
the download client without the library server, or the film manager without the music one — and
folding them into a single module to express a shared directory would make *what a module is*
turn on an incidental filesystem contract. It also does not generalise: the next pipeline of
modules handing files to each other on one machine (an ingest folder, a spool, a drop directory)
would face the same wall and the same wrong remedy.
2. **One module owns the directories and the rest `require` them.** **Rejected**, and this is the
heart of the decision. Nobody owns shared operator data. The library predates the mesh and
outlives any one module, so making the library server or a manager its owner means unassigning
that module orphans everyone else's access — and the owner would set the owner and mode of a
tree it did not create. Ownership is the wrong relationship to model, because the true owner is
not a module at all.
3. **A flag on a directory resource** — `external: true`, or an owner of `operator`. **Rejected.**
It overloads one shape with a boolean that inverts every one of its semantics: created becomes
*must already exist*, owned becomes *touch nothing*, removed-when-empty becomes *never removed*.
That is the *two-kinds-spelled-identically* trap of 04-ISSUES/026 reintroduced with a single
quiet field — a reviewer reading `type: directory` would have to check one boolean elsewhere to
know whether the mesh owns the thing at all.
4. **A distinct `accesses` declaration, separate from resources.** **Adopted.**
## Decision
**Shared, pre-existing data is operator-owned and external. The mesh does not create it, does not
set its owner or mode, does not reconcile it and does not remove it.** A media library, a download
spool, an ingest directory is the operator's, and the mesh is a guest in it.
**A module declares that it needs *access* to such a path, not that it owns a resource there.** The
manifest field is `accesses`: a list of `{path, mode}`, where mode is `read` or `read-write` and
absent narrows to `read` — the safe default, because the danger with an access is being given more
than was meant, not less. A module's own configuration and state directories stay owned
`directory` resources; only the shared, pre-existing paths become accesses.
**The host mounts an accessed path and owns nothing about it.** It reaches the machine as a new
declaration shape, `access`, distinct from `directory`. The host confirms the path is present and
does nothing else — no create, no chown, no mode, no removal.
**An accessed path absent at apply time is refused, clearly, not created.** The mesh does not own
it, so conjuring it would be a lie the host then acts on — and specifically the lie 04-ISSUES/026
records, where a bind mount whose source does not exist is made by the container runtime as root
with the wrong ownership. The host says the operator must provide the path instead.
**Several modules accessing one path is normal, and never refused.** The duplicate-path refusal is
about *ownership*, not *use*: it applies to resources a module owns and to those alone. An access
is not a resource and never enters the check, so the eight-module stack co-resolves. What stays
refused is genuine rivalry — two modules owning one path — and the new contradiction it exposes: a
path one module owns while another merely accesses it, because that asserts both that the mesh owns
the directory and that the operator does.
This is a decision and not a patch because it settles *what a module may say about a path it did
not make*, which every co-located file-handoff in the catalogue now and later depends on — and
because it draws the ownership line [ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md)
started: the host owns what it made and keeps what it merely configured, and this adds the third
case it did not have a word for — what it neither made nor configured, and must not touch.
### How each claim is checked
- **The stack co-resolves.** A control-plane unit test assigns two modules that declare access to
one path on one node and asserts no refusal — the exact case the resolver refuses when the same
path is owned. The mirror test, two modules *owning* one path, still refuses, so the sharing
vocabulary does not weaken the rule it sits beside.
- **Ownership and access cannot both be claimed of one path.** A unit test asserts the resolver
refuses a path one module owns and another accesses, naming both.
- **Absent is refused, not created.** A host unit test applies an access to a path that does not
exist and asserts a clear refusal that names the operator, and that nothing was created.
- **Present is confirmed and nothing moves.** A host unit test applies an access to an existing
directory and asserts the apply reports no change and disturbs nothing.
- **Undeclaring never removes.** A host unit test drops a previously declared access and asserts
the operator's directory and its contents are left exactly as they were — the data-loss failure
[ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md) exists to prevent, on a directory the
mesh never made.
- **Every media manifest is corrected.** No `/services/media/*` path is an owned `directory`
resource in any of the eight; each is an `accesses` entry, and each module's own config and state
directories remain owned. Checked by the control-plane manifest parser, which now understands
`accesses` and refuses a malformed one.
## Consequences
**A shared filesystem between co-located modules now has a vocabulary**, and it is not the media
stack's alone: any pipeline handing files to a neighbour on one machine — an ingest directory, a
spool, a drop folder — says *I access this operator path* rather than *I own this directory*, and
several of them may say it of one path.
**Unassigning a module that reached shared data leaves the data.** Correct, and the same trade
[ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md) made for owned directories: removing
data is a person's act, done knowingly, not a side effect of unassignment.
**The operator must provision the shared paths before the stack is applied**, and a machine that
lacks one is told plainly which. That is a real new obligation, and it is the right one: the mesh
cannot own what predates it, so it cannot create it either, and saying so at apply time beats a
directory conjured as root and a service that half-works.
**The host vocabulary grew by one shape**, which is a cost — every added shape widens what a
compromised control plane can express ([ADR 0005](0005-the-node-host.md)). It is a narrow one: an
`access` is confirmed by a stat and grants the host no new action. It earns its place by letting
the host refuse to create what it must not own, which no existing shape could say.
**A path can be both owned and accessed only by refusal.** If a future manifest declares one path
as an owned directory in one module and an access in another, the resolver refuses it rather than
guessing which is meant — the two assertions about who owns the data cannot both hold.
## References
- [04-ISSUES/036](../04-ISSUES/036-six-modules-own-what-they-must-share/00-report.md) — six (in
fact eight) modules own what they must share; the problem this resolves
- [04-ISSUES/026](../04-ISSUES/026-the-data-directories-are-not-declared/00-report.md) — the two
kinds of mount spelled identically, which deferred this field to a decision
- [ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md) — data outlives the mesh; the host
keeps what it did not make. This record adds the case it had no word for
- [ADR 0005](0005-the-node-host.md) — the host removes what it made and leaves what it merely
configured; the vocabulary is finite and every shape is a security decision
- [ADR 0040](0040-what-a-module-is.md) — what a module is; an access is a new thing a module may
say about the machine it lands on
- mesh-control `feat/shared-data-access`, mesh-catalog `feat/media-access-not-ownership`,
mesh-host `feat/mount-operator-owned` — the mechanism, the corrected manifests, and the host
shape
@@ -0,0 +1,198 @@
---
topic: what runs on it
status: accepted
date: 2026-09-05
deciders: jochen
reconstructed: false
extends: 0005-the-node-host.md
---
# 52. An init step is a container run once to completion, gating what follows
## Context
**A module can declare things that exist; it cannot declare a step that runs.** The host owns a
finite vocabulary of shapes — `file`, `directory`, `service`, `package`, `container`, `action` —
and every one but `action` describes *state*: a thing that should be present, with content or a
mode or an image, which the host reconciles toward ([ADR 0005](0005-the-node-host.md)). That is
right for what it covers. But a real class of modules needs, once, to *run their own code at a
point in their own lifecycle* — and the vocabulary has no word for it
([04-ISSUES/037](../04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md)).
**mosquitto is the sharp case, and it fails silently without this.** Its Dynamic Security plugin
loads at broker start and refuses to come up unless `dynamic-security.json` already holds an admin
client. Seeding that file is a step that must happen *after* the data directory exists and *before*
the broker container starts. The manifest can declare the directory, the config file and the broker
container; it cannot declare "seed this, once, before that container starts." Written as it is
today, the broker starts against an unseeded store and the plugin aborts — and the next reconcile
does not fix it, because nothing in the declaration ever seeds the file.
**It is not one module's defect.** The database providers need the same to run a first-boot
migration, an extension enable, or a health gate before they are announced ready; today that works
only where the *image* happens to seed itself from an environment variable, and anything the mesh
must run once against the server has no home. This is the timing face of the same gap
[04-ISSUES/035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md) records
from the content side: the manifest needed *the file to exist before first start*, and had only
*the file has this content, forever*.
**The obvious answer is the one that already went wrong.** An earlier mesh had exactly this as a
feature — event-driven hooks that ran custom code at phases of build, publish and deploy. It was
powerful and it was *complex to set up and flaky*, and that fragility, not the need, is the content
of the issue. Whatever this becomes must not rebuild that engine.
**The ground has shifted since that engine, in a way that makes a much smaller answer possible.** A
module with tools or events now runs a **process of its own** — a container carrying the module's
compiled code, holding the single broker account the mesh scoped to it, isolated from every other
module ([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). The
code that must run at first boot is *already in that image, under that account*. So the mesh does
not need a way to run a module's code — it has one. It needs a way to say **run this container to
completion, and start the one that depends on it only after it has.**
## Considered Options
1. **A host `run` shape — a command the host executes on the machine.** The direct reading of
"run code at a lifecycle phase." **Rejected.** The host's vocabulary is finite and every added
shape is a security decision, because it widens what a *compromised control plane* can express
([ADR 0005](0005-the-node-host.md)). A general "run this command as the host" is the largest
such widening there is: the blast radius is the whole machine, as root. The mesh already drew
this exact line for `action` — it runs a command, and so it is *permitted from the bundle and
refused from the link*, because the bundle arrives with the binary and the link is a separate
party with an unbounded reach. A new host-command shape usable by an ordinary module would be an
`action` from the link by another name, which is precisely what is refused.
2. **Per-phase lifecycle hooks on a module** — `pre-start`, `post-start`, `pre-remove`, and their
build/publish cousins, each naming code the mesh runs at that phase. The general answer, and the
old feature. **Rejected for now.** It is the flaky engine the issue warns against, and most of
its phases have no present need. Deciding the full set of phases, where each one's code runs, and
how each is made idempotent is a large design taken to buy capability nothing yet asks for. The
three blocked modules all need one phase — *before a container starts* — and a mechanism narrow
enough to be obviously correct beats a general one that is not.
3. **A distinct one-shot resource type** — a new shape, sibling to `container`, that names an image
and runs it once. **Rejected.** It grows the host vocabulary by a whole shape (a `Type`, a
struct, an applier, a place in every host's shape list) to express something a `container` almost
already is. A one-shot *is* a container — a pinned image, an account, volumes, an environment —
that happens to exit. Spending a new shape on the difference is the cost of option 1 in smaller
type, for a capability the existing shape can carry with one modifier.
4. **A modifier on the existing `container` shape: this container runs once, to completion, and the
host gates the apply on it.** **Adopted.** It reuses the shape the host already has, adds no new
host action, and leans on two guarantees the host already gives — *apply in declared order,
never sorted*, and *a failed step fails the apply* — to turn "before that container starts" into
an emergent property of ordering rather than a dependency graph the host must resolve.
## Decision
**A run-once step is an ordinary `container`, marked to run to completion.** The manifest sets
`run-once: true` on a container resource. Everything else about it is a container as before — a
digest-pinned image ([ADR 0006](0006-the-substrate-and-the-control-plane.md)), volumes, an
environment, and for a module's own code the same scoped account its runtime already holds
([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). The mesh adds
no way to run code; it marks a container the host must **run to completion and require to exit 0**,
rather than start and leave running.
**The gate is declaration order, not a named dependency.** The host applies a declaration in the
order it is given, does not sort, and does not resolve dependencies — ordering is a decision, and it
is the control plane's ([ADR 0005](0005-the-node-host.md),
[04-ISSUES/013](../04-ISSUES/013-a-file-arrives-after-the-service-that-needs-it/00-report.md)). A
run-once step is placed *before* the container that depends on it, and **a failed run-once step
halts the apply**, exactly as a failed `action` does — so everything the declaration places after
it, the broker included, is never reached until the step has completed. "Before the broker starts"
is therefore expressed by list position plus completion, and the host cross-references nothing.
**Completion is recorded, and a re-apply does not re-run it.** The host records what it applied only
after the fact, as the digest of the declaration that produced it
([ADR 0018](0018-a-picture-is-read-from-what-runs.md)) — a run-once step no differently. Because the
step leaves nothing running to inspect, that persisted digest, not a live container, is the marker
that it happened. On a later apply the host finds the digest already recorded for this exact
declaration and does nothing; it re-runs only when the declaration's digest has changed, and a step
that exited non-zero recorded nothing and so is retried next apply. This is the reconcilable,
idempotent discipline the state shapes get for free, made explicit for a step.
This is a decision and not a patch because it settles **what a module may say about running its own
code**, which the whole catalogue of providers — a seed before start, a first-boot migration, a
health gate — now and later depends on, and because it draws the line the issue asked for: the
narrowest sound mechanism that unblocks the three modules without rebuilding the hook engine whose
fragility is the warning.
### The security bound, stated plainly
**The host gains no new action and no new shape.** `run-once` is a boolean modifier on the
`container` shape that already exists. A run-once container is strictly *less* powerful than an
`action`: it cannot run an arbitrary host command, only a digest-pinned image under an account the
mesh scoped — which is exactly the capability `container` already grants from the link. A control
plane that is compromised can express nothing through `run-once` it could not already express by
declaring an ordinary `container`. The dangerous expansion of option 1 — a command the host runs on
the machine — is not made.
### How each claim is checked
- **A run-once step runs to completion and its exit 0 is required.** A host unit test applies a
run-once container whose image exits 0, asserts the host ran it to completion (not detached, not
left running) and reported it done; a sibling test applies one that exits non-zero and asserts
the apply fails, naming the step.
- **A failed run-once step gates what follows.** A host unit test places a run-once container that
exits non-zero before another container and asserts the second is never started and the apply is
reported gated — the mirror of the existing test that a failed action stops what follows.
- **It is not re-run once it has completed.** A host unit test applies a run-once step, then applies
the identical declaration again with the first run's record present, and asserts the second apply
runs nothing and reports the step unchanged.
- **A changed declaration re-runs it.** A host unit test applies a run-once step, then applies one
whose image or environment differs, and asserts it runs again — the digest moved, so the marker no
longer matches.
- **The vocabulary carries the field end to end.** A control-plane unit test resolves a module whose
manifest marks a container `run-once` and asserts the rendered host declaration carries the field,
in author order before the container it gates; the manifest parser refuses a `run-once` that is
not boolean.
- **mosquitto seeds before the broker.** mosquitto's manifest declares a run-once init container,
before the `server` (broker) container, that writes the admin client into `dynamic-security.json`
and exits — checked by the control-plane resolver, which now understands the field, and by the
ordering of the rendered declaration. The end-to-end proof that it runs exactly once, at the right
phase, and converges on re-apply is owed to a lab scenario ([04-ISSUES/037] open question), which
this record does not close.
## Consequences
- **The three blocked modules gain a home for their step.** mosquitto seeds its dynsec admin before
the broker; a provider that must migrate or health-gate at first boot declares a run-once step in
its own runtime image, under its own account, before the container that depends on it.
- **The general lifecycle hook is deferred, deliberately.** Only *before a container starts* is
bought here. `post-start`, `pre-remove` and the build/publish phases remain unbuilt, and the day
one is genuinely needed it is decided then, against a need, not speculatively — the same restraint
that kept this from being the old engine.
- **The seed-then-mutate file is safe if the step is written to be.** A run-once seed writes
`dynamic-security.json` only when it is absent and never reconciles it, so what the running plugin
grows in that file afterward is never wiped
([04-ISSUES/035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md)).
The host's marker guarantees the step is not re-run; the step's own code guarantees it does not
clobber on the pass it does run.
- **The host vocabulary did not grow, and that is the point.** The cost of a run-once step is one
boolean and a completion path in the container applier, not a new shape and not a new action. The
mesh expresses ordering and completion; the module runs its own code, where it already runs it.
- **A run-once step that never converges is a stuck apply, loudly.** A step that exits non-zero
every time halts the apply every time, and the container it gates never starts — which is the
correct failure, reported, rather than a broker that half-starts against an unseeded store and a
reconcile that reports success. It is failed forward, not failed silent.
## References
- [04-ISSUES/037](../04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md) — a
module cannot run its own code at a lifecycle phase; the gap this resolves, and the warning about
the old hook engine
- [04-ISSUES/035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md) — the
content face of the same gap: a file needed before first start, that the running program then
mutates
- [04-ISSUES/013](../04-ISSUES/013-a-file-arrives-after-the-service-that-needs-it/00-report.md) — the
order of a declaration is the control plane's, and the host applies it as given; the gate rests on
this
- [ADR 0005](0005-the-node-host.md) — the host's finite vocabulary, the ordered declaration it does
not sort, and `action` as the shape a command already is and why it is refused from the link
- [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) — a module runs
its code as its own process under its own account; a run-once step is that process, run to
completion
- [ADR 0018](0018-a-picture-is-read-from-what-runs.md) — what was applied is recorded after it works;
the completion marker is that record
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — a container is pinned by digest; a
run-once container no differently
- mesh-control `feat/lifecycle-run-once`, mesh-host `feat/apply-run-once`, mesh-catalog
`feat/mosquitto-bootstrap` — the vocabulary, the apply support, and mosquitto's seeded broker
@@ -0,0 +1,181 @@
---
topic: what runs on it
status: accepted
date: 2026-09-06
deciders: jochen
reconstructed: false
extends: 0052-a-step-that-runs-once-before-a-container.md
---
# 53. A scheduled step is a container run on a recurring schedule
## Context
**[ADR 0052](0052-a-step-that-runs-once-before-a-container.md) gave the mesh a step that runs *once*;
a real class of modules needs one that runs *again and again*.** kometa reconciles a media library
against its lists on a timer; a ticketing integration polls its source for new work every few minutes;
a backup, a cache warm, a metrics roll-up all recur. The host's vocabulary describes *state* — a file,
a directory, a container that should be running — and 0052 added *a step that happens once and is
done*. Neither says *this should happen every night at 3, forever*
([04-ISSUES/037](../04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md) named
the lifecycle gap; 0052 closed the once-before-start face of it and left the recurring face open).
**The modules that need it already run their own code.** As with run-once, the ground has shifted
since the old mesh's flaky hooks: a module with tools or events runs a **process of its own** — a
container carrying the module's compiled code under the single scoped account the mesh gave it
([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). The work that
must recur is already in that image, under that account. The mesh does not need a way to run a
module's code on a timer — it has the code and the account. It needs a way to say **run this
container again on this cadence**.
**The shape of the answer is already decided, one modifier over.** 0052 rejected a host `run` command,
per-phase hooks, and a distinct one-shot resource type, and adopted *a modifier on the `container`
shape the host already has*, because a one-shot is a container that happens to exit. A scheduled step
is the same container that happens to exit — run repeatedly. Reusing that shape keeps the host
vocabulary flat and inherits 0052's security bound whole. The only thing 0052's `run-once` does not
carry is *when to run it again*.
**One difference from run-once changes a rule, and it is the reason this is its own record.** A
run-once step **gates the apply**: it is placed before the container that depends on it, and a failure
halts everything after it, because "seed the store before the broker starts" is a correctness
precondition ([ADR 0052](0052-a-step-that-runs-once-before-a-container.md)). A scheduled step is the
opposite: it runs *after* the machine is up and converged, on its own clock, and a single failed run
is an ordinary operational event — the next run comes anyway. A scheduled step that halted the apply,
or that a failed run marked the node not-current over, would make a routine poll into a reason the
whole machine reads as broken. So the gating rule 0052 established is exactly the rule this record
must **not** inherit.
## Considered Options
1. **A host `cron`/`timer` shape — the host installs a system timer that runs a command.** The direct
reading. **Rejected**, for the reason 0052 rejected a host `run` shape: it widens what a
*compromised control plane* can express toward "run this command on the machine, forever," which is
the largest widening there is, and it is `action`-from-the-link by another name
([ADR 0005](0005-the-node-host.md)). A recurring command is worse than a one-off, because it
persists.
2. **Per-phase lifecycle hooks** — `on-schedule` joining `pre-start`/`post-start` as named code the
mesh runs. **Rejected for now**, as in 0052: it is the flaky hook engine the issue warns against,
and the three modules that need this need one thing — *run this container on a cadence* — which a
narrow modifier expresses without deciding a whole hook vocabulary.
3. **A distinct `scheduled` resource type**, sibling to `container`. **Rejected**, as 0052 rejected a
distinct one-shot type: it spends a whole new host shape (a `Type`, a struct, an applier, a place
in every host's shape list) on something a `container` already almost is — a scheduled task *is* a
container (pinned image, account, volumes, environment) that runs on a clock.
4. **A modifier on the existing `container` shape: `schedule`, a cron expression the host runs the
container on.** **Adopted.** It reuses the shape the host has, adds no new host action, and sits
beside `run-once` as its recurring twin — the same container, exited, run again.
## Decision
**A scheduled step is an ordinary `container`, marked with a `schedule`.** The manifest sets
`schedule: "<cron>"` on a container resource — a standard five-field cron expression. Everything else
about it is a container as before: a digest-pinned image
([ADR 0006](0006-the-substrate-and-the-control-plane.md)), volumes, an environment, and for a module's
own code the same scoped account its runtime already holds
([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). The mesh adds no
way to run code; it marks a container the host must **run on that cadence, each time to completion**,
rather than start once and leave running (a service) or run once and gate (a run-once step).
**A run is fired by the clock, not by the apply, and does not gate it.** Applying the declaration
installs the schedule; it does not run the step. The machine converges — reports applied and current —
as soon as the schedule is installed, exactly as it does for a service that is running. Thereafter the
host fires the container when the cron expression is due. This is the deliberate inversion of 0052:
a scheduled step is downstream of convergence, not a precondition of it.
**A failed run is recorded and the next run still comes; it never marks the node not-current.** A run
that exits non-zero is logged against the module — visible, auditable — but it does not fail the apply,
does not halt other resources, and does not flip the node's reported state. A poll that fails at 03:00
and succeeds at 03:05 is the system working, not a machine that needs attention. A step that fails
*every* time is a loud, repeating log entry, which is the correct signal for "this recurring job is
broken" — distinct from "this machine did not converge."
**Runs do not stack.** If a run is still going when the next is due, the host skips the due run rather
than starting a second copy, and logs the skip. A slow nightly job that occasionally overruns must not
spawn a growing pile of concurrent containers competing for the same account and volumes — the failure
mode that made the old timers dangerous.
**Each run is independent and idempotent by the module's own code.** The mesh guarantees only *the
container is run on the cadence*; that a run does the right thing when the previous one half-finished
is the module's contract, the same discipline a run-once seed owes
([ADR 0052](0052-a-step-that-runs-once-before-a-container.md),
[04-ISSUES/035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md)).
This is a decision and not a patch because it settles **what a module may say about running its own
code on a cadence**, which every recurring provider job — a sync, a poll, a roll-up — now and later
depends on, and because it draws the line 0052 left: the recurring twin of run-once, with the gating
rule deliberately reversed so a routine job's failure is never a machine's failure.
### The security bound, stated plainly
**The host gains no new action and no new shape.** `schedule` is a string modifier on the `container`
shape that already exists. A scheduled container is strictly *less* powerful than an `action`: it runs
a digest-pinned image under an account the mesh scoped, which is exactly what `container` already
grants from the link, and it cannot run an arbitrary host command. A compromised control plane can
express nothing through `schedule` it could not already express by declaring a `container` — the cron
string only says *how often*, not *what*. The dangerous expansion of option 1 — a command the host
runs on the machine on a timer — is not made.
### How each claim is checked
- **A scheduled step runs when the schedule is due.** A host unit test installs a container with a
schedule that is due immediately (or advances a injected clock to when it is due) and asserts the
host ran it to completion; a sibling test with a schedule not yet due asserts it has not run.
- **Installing it does not run it, and the node is current without a run.** A host unit test applies a
scheduled container and asserts the apply reports current *before* any run has fired — the schedule
is state that is present, not a step that gated.
- **A failed run does not fail the apply or the node.** A host unit test fires a scheduled container
that exits non-zero and asserts the failure is recorded against the module, the apply is not failed,
and the node stays current — the mirror of the run-once test where a non-zero exit *does* halt.
- **Runs do not stack.** A host unit test fires a scheduled container whose run outlasts its next due
time and asserts the host skipped the due run and logged the skip, rather than starting a second
container.
- **The vocabulary carries the field end to end.** A control-plane unit test resolves a module whose
manifest sets `schedule` on a container and asserts the rendered host declaration carries the field;
the manifest parser refuses a `schedule` that is not a valid cron expression, and refuses a container
that is both `run-once` and `schedule` (a step is one or the other, never both).
- **A real module recurs in the lab.** A converted module declaring a scheduled step (kometa's library
sync, or a poller) is assigned in a lab scenario, and the scenario asserts the scheduled container
fires on its cadence and its account and volumes are the module's — the end-to-end proof this record
owes, as 0052 owed its run-once lab proof.
## Consequences
- **The recurring providers gain a home for their cadence.** kometa reconciles on its schedule; a
poller polls; a roll-up rolls up — each a scheduled step in its own runtime image, under its own
account, on the cron it declares.
- **The general lifecycle hook is still deferred.** Only *run once before* (0052) and *run on a
cadence* (this) are bought. `post-start`, `pre-remove` and the build/publish phases remain unbuilt,
decided when a real need arrives, not speculatively — the restraint that kept both from being the old
engine.
- **A recurring job's failure is loud but not fatal.** The node stays current while a scheduled step
fails and retries; a step that fails forever is a repeating log entry, not a machine marked broken.
This is the correct separation — a machine's convergence and a job's success are different questions —
and it is why this could not simply be `run-once` without the schedule.
- **The host vocabulary did not grow, again, and that is the point.** The cost of a scheduled step is
one string field and a cron loop in the container applier, not a new shape and not a new action. The
mesh expresses cadence; the module runs its own code, where it already runs it.
- **`run-once` and `schedule` are exclusive and complete for now.** A container runs once and gates, or
runs on a cadence and does not, or runs and stays up (a service). A manifest that asks for two of
these at once is refused, because the three are distinct answers to "how does this container run."
## References
- [ADR 0052](0052-a-step-that-runs-once-before-a-container.md) — the run-once step; this is its
recurring twin, reusing the `container`-modifier shape and inheriting its security bound, and
reversing its gating rule
- [04-ISSUES/037](../04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md) — a
module cannot run its own code at a lifecycle phase; 0052 closed the once-before-start face, this
closes the recurring face
- [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) — a module runs
its code as its own process under its own account; a scheduled step is that process, run on a cadence
- [ADR 0005](0005-the-node-host.md) — the host's finite vocabulary, and why a host command (even on a
timer) is refused from the link
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — a container is pinned by digest; a
scheduled container no differently
- [ADR 0018](0018-a-picture-is-read-from-what-runs.md) — what was applied is recorded after it works;
the installed schedule is state, a fired run is an event
- mesh-control `feat/schedule-container`, mesh-host `feat/apply-schedule`, and the converted module
that first declares a scheduled step — the vocabulary, the apply support, and the recurring proof
@@ -0,0 +1,178 @@
---
topic: what runs on it
status: accepted
date: 2026-09-06
deciders: jochen
reconstructed: false
extends: 0050-model-access-is-vendor-agnostic.md
---
# 54. Model usage is a vendor-neutral record, produced by the adapter, at two grains
## Context
**[ADR 0050](0050-model-access-is-vendor-agnostic.md) fixed the *shape* of a usage reading and left
its *home* open.** It decided that an adapter may expose `usage(licence) → normalised rows`, that a
row is `(licence, consumer, period, metric, value)` plus the raw vendor response as jsonb, and that the
metric is vendor-defined (Anthropic's `utilization%` is one metric, a token count another). It did not
say where those rows are stored, how they get produced on a cadence, or whether the only grain is the
whole licence — and the mesh's first real consumer needs all three answered.
**The predecessor mesh recorded usage at two grains, and both are wanted.** It polled the vendor for a
licence-level reading (an account's `utilization%`), and it also attributed **per-session** token and
cost — which model, how many input and output tokens, what it cost — parsed from the agent's own
transcript, including the case where a long session switched the account it billed against mid-way. The
licence-level reading answers "how close is this subscription to its cap"; the session-level reading
answers "what did this piece of work cost, and against which account." A mesh that kept only the first
could not bill a project or notice a runaway session; keeping only the second could not see a cap
approaching. Both are load-bearing and neither subsumes the other.
**A session is already a consumer, so the second grain needs no second vocabulary.**
[ADR 0026](0026-the-mesh-has-a-session-of-its-own.md) and design 15 establish that an agent session is
a thing in its own right, and that **model access is bound to the session, not to the machine** — a
session is a *consumer* of `model-access`, identified by `(node, module)` and its own id. That is
exactly the `consumer` column ADR 0050 already put in the usage row. So the two grains are not two
schemas; they are the same row at two consumer resolutions: the holding module for the licence grain,
the session for the finer one. The design's own still-open worker-naming gap (design 14) is the same
gap here and is left where it is — a session id distinguishes what `(node, module)` cannot.
**The pieces this needs already exist.** A periodic reading is a **scheduled step**
([ADR 0053](0053-a-step-that-runs-on-a-schedule.md)) — the adapter's `usage()` poll is a container the
mesh runs on a cadence, which is precisely what 0053 was built for. A durable audit of what happened is
**an event the audit trail records** ([ADR 0041](0041-events-are-a-relationship.md),
[ADR 0042](0042-the-shape-of-an-event-on-the-wire.md)); the `audit-logger` module already consumes every
event. And a queryable current picture is **a context store**, the same shape the licences themselves
live in ([ADR 0008](0008-a-context-owns-its-store.md)). Nothing new in kind is required; what is missing is
the decision to point them at usage.
## Considered Options
1. **Licence grain only — a single `utilization%` poll, nothing per session.** Rejected: it cannot
attribute cost to a piece of work or catch a session that is burning an account down, which is half
of why usage is recorded at all.
2. **A bespoke `sessions` schema mirroring the old mesh's `*_sessions` / `*_session_account_usage`
tables.** Rejected: it reintroduces a second, vendor-shaped vocabulary for something the mesh
already names — a session is a consumer, and its usage is a usage row. A parallel schema would drift
from the `model-access` vocabulary and force every reader to learn two.
3. **Store usage only as raw vendor blobs, normalise later.** Rejected as the *whole* answer (kept as a
fallback within the chosen one): a reader that must parse Anthropic's response shape to answer "what
did this cost" has the vendor coupling the whole feature exists to remove. The raw blob is kept
beside the normalised row (0050 already requires this), not instead of it.
4. **One vendor-neutral usage record at two consumer grains, produced by the adapter, recorded as
both an event and a queryable row.** Adopted.
## Decision
**Model usage is one vendor-neutral record — `(licence, consumer, period, metric, value)` plus the raw
response — recorded at two grains that differ only in the `consumer`.** At the **licence grain** the
consumer is the holding module and the metric is the vendor's own account reading (Anthropic:
`utilization%`). At the **session grain** the consumer is the agent session — `(node, module)` and its
session id ([ADR 0026](0026-the-mesh-has-a-session-of-its-own.md)) — and the metrics are the ones a
session bills: input tokens, output tokens, model, and cost. The row shape is 0050's, unchanged; the
grain is which consumer the row is *for*.
**The adapter is the only thing that knows the vendor, and it produces both grains.** Reading an
account's cap is the adapter's `usage(licence)` verb ([ADR 0050](0050-model-access-is-vendor-agnostic.md));
attributing a session's cost is the adapter reading that vendor's transcript or usage API and emitting
rows keyed to the session. The mesh defines the row and the plumbing; the adapter fills it from whatever
the vendor exposes, and a static-key vendor that exposes nothing simply produces no rows — usage is an
optional reading, not a requirement of holding a licence.
**A reading is taken on a schedule, not on a request.** The licence-grain poll is a **scheduled
container** ([ADR 0053](0053-a-step-that-runs-on-a-schedule.md)) the adapter runs on a cadence; the
session-grain rows are produced as sessions progress, from the transcript the session already writes.
Neither blocks anything: a poll that fails is a logged, retried scheduled run (0053's rule), and a
session whose cost cannot yet be attributed is a row not yet written, never a session refused.
**Usage is recorded two ways, for two audiences.** Each reading is **emitted as an event**
([ADR 0041](0041-events-are-a-relationship.md)) — an immutable "this was observed at this time" that the
`audit-logger` already records, so the history of what an account did is in the audit trail by default,
under nobody's special arrangement. And the **current** picture — the latest reading per
`(licence, consumer, period, metric)` — is upserted into a **usage context store**, so "how close is
this cap" and "what has this project spent this month" are a query, not a fold over the event log. The
event is the record of what happened; the store is the answer to what is true now.
**Usage is not a credential, and is recorded in the clear.** The one thing the mesh must not read is the
sealed key ([ADR 0050](0050-model-access-is-vendor-agnostic.md), the refresh-token carve-out aside). A
token count and a cost are not secrets; they are the operator's own operational facts, recorded openly
so they can be queried, audited, and charged against. This is the deliberate opposite of the credential
rule, and stating it prevents a later reader assuming usage inherits the key's secrecy and hiding it
from the person who is paying.
This is a decision and not a patch because it settles **where a model's usage lives and at what grain**,
which every reader — a bill, a cap alarm, a per-project report — depends on, and because it closes the
half [ADR 0050](0050-model-access-is-vendor-agnostic.md) explicitly left open, reusing the session
([ADR 0026](0026-the-mesh-has-a-session-of-its-own.md)), the schedule
([ADR 0053](0053-a-step-that-runs-on-a-schedule.md)), the event ([ADR 0041](0041-events-are-a-relationship.md))
and the store ([ADR 0008](0008-a-context-owns-its-store.md)) the mesh already has rather than inventing a
vocabulary beside them.
### How each claim is checked
- **A usage row is vendor-neutral and carries the raw beside it.** A unit test constructs an Anthropic
`utilization%` reading and a token/cost reading and asserts both render to
`(licence, consumer, period, metric, value)` with the vendor response preserved in the raw column;
a reader that answers "what did this cost" touches only the normalised columns.
- **The two grains differ only in the consumer.** A unit test records a licence-grain row (consumer =
the module) and a session-grain row (consumer = a session id) for one licence and asserts both are
the same shape and both are returned when the licence's usage is asked for, distinguishable by
consumer.
- **A poll is a scheduled run and its failure is not fatal.** The adapter's usage container declares a
`schedule` ([ADR 0053](0053-a-step-that-runs-on-a-schedule.md)); a host test (0053's) already proves a
scheduled step runs on cadence, does not gate, and logs rather than fails on a non-zero run — the poll
inherits this and adds nothing to check.
- **Each reading is an event the audit trail records.** An integration check asserts a usage reading
emits an event that the `audit-logger` receives (it consumes `#`), so the history is present without
the usage module and the audit module knowing about each other beyond the event.
- **The current picture is a query.** A store test upserts two readings for one
`(licence, consumer, period, metric)` and asserts the later replaces the earlier, so "what is true
now" is one row, while the event log keeps both.
- **A static-key vendor with no usage reading records nothing, and that is fine.** A test resolves a
`static-key` model-access consumer whose adapter has no `usage` verb and asserts the licence works and
no usage rows or poll are required — usage is optional, holding a licence is not conditioned on it.
- **Usage is readable in the clear; the key is not.** A test asserts a usage row is stored unsealed and
is returned to an ordinary query, while the licence key remains sealed and absent from the same
surfaces — the deliberate inversion of the credential rule.
## Consequences
- **A bill and a cap alarm are both queries.** "What did project X spend this month" reads the
session-grain rows; "how close is account Y to its cap" reads the latest licence-grain metric — both
from the usage store, neither a fold over events or a call to the vendor.
- **The session becomes the unit of cost, which is what it already is.** Because a session is the
consumer, attributing cost needs no new identity — and the worker-granularity gap
([design 14](../03-DESIGN/01-to-be/14-model-access.md)) surfaces here exactly as it does for access,
to be closed once, for both, when a session id is threaded through.
- **The adapter carries the vendor's usage quirks alone.** Anthropic's `utilization%`, its transcript
shape, a mid-session account switch — all live in the Anthropic adapter; the mesh, the store, and
every reader see only rows. A second vendor adds a second adapter and no new table.
- **Usage history is durable and tamper-evident by reuse, not by a new mechanism.** It rides the event
trail the mesh already keeps, so an operator who wants the whole history has it, and one who wants the
current number has the store — without usage owning either mechanism.
- **The refresh-token carve-out is untouched by this.** Usage is read *from* an authenticated adapter;
it neither holds nor exposes the credential, so the one place the mesh reads what it stores
([ADR 0050](0050-model-access-is-vendor-agnostic.md)) is not widened by recording what that credential
was spent on.
## References
- [ADR 0050](0050-model-access-is-vendor-agnostic.md) — model access is vendor-agnostic; fixes the
usage row shape and the `usage(licence)` adapter verb, and leaves its home open — which this closes
- [ADR 0026](0026-the-mesh-has-a-session-of-its-own.md) — the mesh has a session of its own; a session
is the consumer the finer grain attributes to
- [ADR 0053](0053-a-step-that-runs-on-a-schedule.md) — a scheduled step; the licence-grain poll is one
- [ADR 0041](0041-events-are-a-relationship.md) — events are a relationship; a usage reading is one, and
the audit-logger records it
- [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) — the shape of an event on the wire; the form a
usage reading takes to reach the audit trail
- [ADR 0008](0008-a-context-owns-its-store.md) — a context owns its store; the current usage picture
lives in one
- [ADR 0024](0024-model-access-is-a-provision.md) — model access is a provision; usage is a reading of
what that provision was used for
- [03-DESIGN/01-to-be/14-model-access.md](../03-DESIGN/01-to-be/14-model-access.md) — the model-access
design; this fills its usage section and shares its open worker-naming gap
- [03-DESIGN/01-to-be/15-the-agent-session.md](../03-DESIGN/01-to-be/15-the-agent-session.md) — the agent
session; the consumer the session grain is keyed to
@@ -0,0 +1,146 @@
---
topic: what runs on it
status: accepted
date: 2026-09-07
deciders: jochen
reconstructed: false
extends: 0050-model-access-is-vendor-agnostic.md
---
# 55. Model access is answered by a licence, or by a node that hosts the model
## Context
**[ADR 0024](0024-model-access-is-a-provision.md) made model access a provision, and
[ADR 0050](0050-model-access-is-vendor-agnostic.md) fixed what answers it: a record — a licence — that
an adapter turns into a sealed vendor credential.** A consumer requires `model-access`, is put on a
licence, and is delivered a key (a static API key, or an access token a manager refreshes). Every
answer so far has been a credential to reach a vendor's API across the internet.
**But a model need not come from a vendor. A node in the mesh can host one.** An operator with a GPU
runs Ollama or vLLM, which serves an OpenAI-compatible API on that node. A consumer that wants that
model does not need a vendor credential — it needs the model server's **endpoint**: the base URL and
the model name, and a key only if the server is configured to want one. This is model access answered
by a **node**, not by a record.
**The mesh already knows how a node answers a provision — it is the ordinary provider/consumer path.**
A provider `provides` a provision at a scope, `serves` its connection facts, and the mesh fills the
consumer's bound facts with the provider's `at`/`port` and each served fact — exactly how a Postgres
consumer learns where its database is. Nothing reserved `model-access` to records: a node offering
`provides: ["model-access"]` resolves through this path, and the resolver already **prefers a local
answer over a licence** — its own comment names the case, "a model the mesh runs itself." So the
capability exists; what is missing is the decision to use it, and the statement of what a node-answer
delivers and where it stops.
**A node-answer and a record-answer are the same provision with two shapes of answer.** This is the
same move [ADR 0050](0050-model-access-is-vendor-agnostic.md) already made for shapes within the
vendor path (static-key vs refreshable-grant): one provision, more than one way it is answered. A
consumer written against `model-access` should not care whether the model behind it is a vendor's or
the mesh's own — it asks for model access and is given what reaches a model.
## Considered Options
1. **A separate provision for the local case (`local-model`, `model-endpoint`).** A node answers that;
`model-access` stays record-only. Rejected: it splits "where my model comes from" into two
provisions a consumer must choose between in its manifest, when the mesh already models a
record-answer and a node-answer to **one** provision. A consumer would have to know, at authoring
time, whether its model will be a vendor's or the mesh's — the exact coupling the provision was
meant to remove. It is the safer implementation (see the limitation below) but the worse interface.
2. **Model access answered by either a licence or a node, under the one provision.** A consumer
requires `model-access`; the operator answers it with a licence (a vendor) or by assigning a
node that hosts a model. Adopted: one interface, and the answer is an operator's deployment choice,
not a consumer's authoring choice.
## Decision
**Model access is one provision answered two ways: by a licence (a record, turned into a sealed vendor
credential by an adapter — [ADR 0050](0050-model-access-is-vendor-agnostic.md)) or by a node that hosts
the model (a provider that serves an endpoint).** A consumer requires `model-access` and is delivered
whichever the operator assigned; it does not name the kind.
**A node-answer delivers an endpoint, not a credential.** The provider `provides: ["model-access"]`
and `serves` its connection facts — the port it listens on and the model it runs — and the mesh fills
the consumer's bound facts with the provider node's `at`, the served `port`, and the served `model`,
the same way every provider consumer learns where its provider is. The consumer assembles a base URL
(`http://<at>:<port>/v1`) and points an OpenAI-compatible client at it. If the local server wants a
key, the provider mints one the ordinary way (a per-consumer secret, sealed and host-unsealed); if it
does not — the common Ollama case — the consumer lists `model-access` under `binds` and **not** under
`secrets`, and no key is delivered. Secret delivery and fact delivery are already independent, so a
keyless endpoint is expressed by asking for the facts and not a secret.
**A node-answer uses no adapter.** The adapter registry ([ADR 0050](0050-model-access-is-vendor-agnostic.md))
is the vendor-credential machinery — accept-and-seal, refresh, usage. A node-hosted model has no vendor
secret to seal; its endpoint is served, and its key (if any) is minted like any provider's. The adapter
is consulted only for the record/vendor answer. So the vendor-agnostic decision is untouched, and the
node-answer adds no vendor logic anywhere.
**The resolver prefers a local answer.** When a node's own set answers `model-access` — a model the
mesh runs itself — a licence for it is not consulted. This is already the resolver's behaviour and is
made a decision here: a mesh that runs a model uses it, and a licence is the answer for a consumer that
has no local model, not a competitor to one that does.
**One limitation, stated so it is not found as a bug.** The local-preference above is exact for a
**node-scope** provider co-located with its consumer, and for any node-answer in a mesh that holds no
`model-access` licence. It is *not* yet exact for a **mesh-scope** model server — one node serving the
model to others — **while a licence for `model-access` also exists in the same mesh**: the record pass
that turns a licence into an answer keys on same-node satisfaction and would still demand the licence be
used, double-answering. Until that pass is taught to stand down when a brokered node need already
answers, a mesh-scope local model and a vendor licence must not both answer `model-access` in one mesh.
A node-scope local model has no such constraint. This is named because an unstated limitation is
indistinguishable from a bug, and costs more.
This is a decision and not a patch because it settles **what may answer model access** — a question
every model-access consumer's meaning depends on — and because it lets the mesh's own hosted models sit
behind the same provision as the vendors', which is what makes "the mesh can run its own model" a
deployment choice rather than a second interface to build against.
### How each claim is checked
- **A node answers model access without a licence, and is preferred over one.** A resolver test
assigns a consumer and a module that `provides: ["model-access"]` at node scope on the one node, with
a licence also present, and asserts no resolved need is answered by the record — the local model
answers and the licence is ignored. (This test exists; the decision adopts what it proves.)
- **The consumer is delivered an endpoint, not a credential.** A mesh bed assigns a model-server
provider and a consumer that binds `model-access` and does not list it under `secrets`, and asserts
the consumer's config carries `OPENAI_BASE_URL` built from the provider's served `at`/`port`, and
that no key file was delivered to its secret path.
- **The local endpoint is reachable through what the consumer was given.** The bed makes a request to
the base URL the consumer wrote and asserts the model server answers — the wiring, not a model's
output, is what is proven (the server may be a stub; a real model is not needed to prove the mesh
routed the consumer to it).
- **A node-answer consults no adapter.** A node-answered `model-access` need is resolved with the
vendor registry never read — asserted by the absence of any vendor on a node-answered need and the
ordinary served-facts delivery.
- **The scope limitation holds where stated.** The node-scope case is what the bed and the resolver
test exercise; the mesh-scope-plus-licence collision is recorded here and left for the resolver
change that reconciles the two answer passes, not worked around in a module.
## Consequences
- **The mesh can run its own model, and a consumer reaches it through the same `model-access` it uses
for a vendor.** One interface, two answers; a consumer moves between a vendor and a local model by an
operator reassigning its provision, not by a code change.
- **A local model is keyless by default and keyed by the ordinary path when it must be.** Nothing new
is invented for the local server's credential: it either has none, or mints one the way every
provider does.
- **The vendor path is untouched.** Adapters, the refresh carve-out, and usage
([ADR 0054](0054-model-usage-is-recorded-at-two-grains.md)) are the record answer's business; a
node-answer neither uses nor changes them. A local model that exposes usage would serve it as facts,
not as an adapter's usage verb.
- **The two answer passes meet in one place, and must be reconciled there.** The mesh-scope limitation
is the single point where a node-answer and a record-answer to the one provision can collide; it is
named, and its fix is a resolver change, not a per-module workaround.
## References
- [ADR 0024](0024-model-access-is-a-provision.md) — model access is a provision; this decides a node
may answer it, not only a record
- [ADR 0050](0050-model-access-is-vendor-agnostic.md) — model access is vendor-agnostic; the record
answer and its adapters, which the node answer sits beside and does not use
- [ADR 0054](0054-model-usage-is-recorded-at-two-grains.md) — model usage; a vendor's business on the
record path, served as facts (if at all) on the node path
- [ADR 0008](0008-a-context-owns-its-store.md) — a context owns its store; a licence is a record in one,
a node-answer needs none
- [03-DESIGN/01-to-be/14-model-access.md](../03-DESIGN/01-to-be/14-model-access.md) — the model-access
design, which this extends with the node answer
@@ -0,0 +1,115 @@
---
topic: the tiers
status: accepted
date: 2026-09-09
deciders: jochen
reconstructed: false
extends: 0007-connectivity.md
---
# 66. Public routing is name-agnostic, its names are resolved inside the mesh, and an internal authority can certify them
## Context
**[ADR 0007](0007-connectivity.md) and [connectivity §3](../03-DESIGN/01-to-be/08-connectivity.md)
made a public route a grant: a workload that must be reachable requires a route, the proxy provides
it, the consumer contributes the name it wants and the port it listens on.** What was never pinned
is **what that name is** — and building a whole mesh in the lab showed the gap costs more than it
looks.
**The catalogue shipped each route as a full domain.** A module that needed a public name carried
that name, in full, as a literal in its manifest. Running the same catalogue against a different
domain — a lab standing in for production, or a second operator's mesh — meant overriding that
literal on every routed module, per node. The mesh was, in effect, carrying a **map of names to
services**: the one thing it should never hold, because a name is the operator's choice (one runs
the forge at `git`, another at `code`) and the domain is the node's, and neither is the mesh's to
know.
**And a second gap surfaced the moment an internal issuer tried to certify those names.**
[Connectivity §5](../03-DESIGN/01-to-be/08-connectivity.md) already states the issuer must be
configurable and that the lab runs its own ACME authority. With that authority wired to the proxy,
issuance still could not complete: the authority accepted the order and offered a challenge, then
**could not connect to the validation target.** Nothing inside the mesh resolved the public route
name. The mesh publishes each `<node>.internal` name into every container, but not the public names
the proxy serves — so a validator living in the mesh had no address to reach, and a name the mesh
cannot resolve is a name it cannot have certified.
**The two are one problem.** A name the mesh can *compose* from parts it is given, and *propagate*
to whoever needs to resolve it, is exactly a name it can also have *certified* — and the reverse:
without the composition and the propagation, neither the routing nor the certificate is the
operator's to move between meshes.
## Considered Options
**1. Keep the full domain in the manifest, override per node.** The status quo. It works, and it is
wrong in the specific way this repository cares about: the catalogue holds a domain map, lab and
production differ by an override on every routed module rather than one fact, and a module manifest
names something — the public domain — that belongs to the node, not the module. An unowned name in
the wrong place is the shape of a leak.
**2. A module declares a label; the node declares its public domain; the mesh composes.** The route
contribution carries a subdomain the operator chose, the node carries its public domain as
node-level configuration, and the mesh joins `<label>.<public-domain>` and grants exactly that. The
mesh interprets nothing. Lab-versus-production becomes one node setting. Chosen.
**3. For certification, issue only from a publicly reachable node against a public authority.**
This is already true for public meshes and stays true. It is not an option for a lab or an
internal-only mesh: there is no public authority to answer, and no public reachability to validate
against. An internal issuer is required there — and an internal issuer must be able to *validate*,
which it cannot do unless the routed name resolves and is reachable **inside** the mesh. So the
resolution gap is not optional to close; it is what makes an internal authority possible at all.
## Decision
**The mesh core holds no map of hostnames, subdomains or domains.** A route contribution carries a
**label** (the subdomain) chosen by the module's operator. A node contributes its **public domain**
as node-level configuration. The mesh composes `<label>.<public-domain>`, grants exactly that name,
and never interprets what it means. One operator's forge at `git.example.tld` and another's at
`code.other.example` are the same module with two facts supplied around it.
**When the proxy is granted a name, the mesh publishes that name → the node that serves it into
internal resolution, mesh-wide** — the same mechanism, and the same "given by the mesh, not chosen
by a module," that already writes `<node>.internal` into every declared container
([connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md)). The mesh propagates the names it was
told to serve. It still knows nothing about what any of them mean.
**An internal authority certifies those names by the same path a public one would.** The proxy is
pointed at whichever issuer the mesh names — a public ACME authority, or an internal one — and
trusts that issuer's root; nothing else about issuance changes. The internal authority validates by
reaching the routed name, which the clause above has just made resolvable inside the mesh. So the
three are one decision: **compose the name, propagate it, certify it** — each is meaningless without
the one before it.
## Consequences
- **Lab-versus-production is one node-level `public-domain` setting**, not an override on every
routed module. The same catalogue runs against any domain.
- **The manifest layer needs composition it does not yet have.** Today a route name is stored as a
literal, with no interpolation of a node's domain into a module's label. Until that exists, the
composed name is produced by a per-node settings override — a stopgap that reproduces option 1's
per-module cost and is explicitly *not* the design.
- **An internal issuer depends on route-name resolution.** Its challenge validates against the
routed name; without that name in internal resolution, issuance for it cannot complete inside the
mesh. The lab found this as a live failure, not a theory.
- **Nothing about the public path changes.** A publicly reachable node issuing a public name from a
public authority is untouched; this widens the same shape to names and meshes that are not public.
**How each is checked** — an unenforced rule is indistinguishable from a wrong one:
- **Name-agnostic:** the same catalogue resolves against two different public domains by changing
one node setting and nothing else; and no module manifest contains a full public domain. A
manifest that pins an FQDN is the smell the check looks for.
- **Resolution:** a request to a routed name, made from inside the mesh, reaches the workload that
serves it — and, the sharper check, issuance for that name against the internal authority
completes, which it cannot unless the validator resolved and reached the target.
- **Internal authority:** a TLS handshake to a routed name verifies against the internal root and
nothing else, the same shape §5 already uses for internal node-to-node names.
## References
- [ADR 0007 — connectivity](0007-connectivity.md), which made a route a grant and named exposure,
resolution and certificates as one context.
- [ADR 0009 — modules and the graph](0009-modules-and-the-graph.md), the provide/require/contribute
vocabulary a route and an authority both use.
- [Connectivity design §2, §3, §5](../03-DESIGN/01-to-be/08-connectivity.md), amended alongside this
record.
+129
View File
@@ -0,0 +1,129 @@
---
topic: the tiers
status: accepted
date: 2026-09-10
deciders: jochen
reconstructed: false
extends: 0006-the-substrate-and-the-control-plane.md
---
# 67. Genesis is a pivot: a temporary control plane installs the registry that makes it permanent
## Context
**The mesh builds its own modules into its own registry, and there is no public registry for them
— that is the point, not an omission.** A module is cloned from the forge, built, and published
where the mesh can move, replace and back it up. Nothing about that arrangement wants a copy of the
mesh's code hosted by somebody else.
**Every image must be pinned by digest** ([ADR 0006](0006-the-substrate-and-the-control-plane.md)),
and the reasoning is exactly right: a bundle is applied where no mesh exists to check anything
against anything, so what it names must be exact.
Those two sentences are individually correct and together they close a door. The digest a pin means
is a *manifest* digest, and a manifest digest is **assigned by a registry when something is pushed
to it**. The control plane's image is built from source and pushed nowhere, so it has no such
digest, so it cannot be named — and a registry cannot be installed without a control plane to
install it. That is not a pin. It is a dependency the pinning rule created by accident.
**It went unnoticed because the lab hid it.** The lab raised a disposable registry, stocked it from
a workstation, and rewrote every image reference to point at it — so the lab bootstrapped along a
path no real machine has. A first node in the lab always worked, and a first node anywhere else had
no path at all. Every bootstrap fault found this year was found late for the same reason: **the
install procedure existed only as a test fixture**, and a fixture is free to invent what it needs.
## Considered Options
**1. Publish the mesh's own images to a public registry.** Rejected on the premise: there is no
public registry for the mesh's modules and there is not meant to be. It would also make raising a
mesh depend on somebody continuing to host its code, which is the dependency the whole arrangement
exists to remove.
**2. Build the control plane from source on the first machine.** Rejected. A bare machine would
need a toolchain and a working tree — and worse, the source lives in a forge **that runs on the
mesh**. A total rebuild would then need the mesh it is rebuilding. Acceptable for adding a node to
a healthy mesh; useless for the case that matters.
**3. Keep a disposable registry as an install step.** Rejected. It exists in no production, and
concealing this problem is precisely what it has been doing.
**4. Pivot through a temporary control plane.** Chosen.
## Decision
**An image may be named by the digest of its own configuration.** A bare `sha256:…` names an image
the machine already holds — content-addressed, immutable, unforgeable, and requiring nothing to
have served it. It satisfies what the pinning rule asks for; the rule simply never contemplated an
image that no registry had ever seen. It is legal exactly where nothing could have served one.
**Genesis is a pivot**, in this order:
1. the installer **carries the control-plane image** and loads it onto the machine
2. a **temporary** control plane is raised from it, named by that image's own digest
3. the **registry module is installed** — its image is upstream and it is never built, which
[`04-ISSUES/029`](../04-ISSUES/029-the-artifact-store-cannot-be-delivered-by-the-artifact-store/00-report.md)
already settled: a module that provides the artifact store cannot be delivered through it
4. the control-plane image is **pushed into the mesh's own registry**, which assigns it a manifest
digest — the first one it has ever had
5. the control plane is **reinstalled as an ordinary module** pinned to that digest
**The host performs the replacement, not the control plane.** Tier 0 outlives tier 2: the control
plane composes a declaration naming the registry-pinned image, and the host applies it and recreates
the container. Nothing is asked to replace itself while running, and the control plane is stateless
— what it knows is in the store.
**The installer is tier 0, and a separate program from the host.** Bootstrapping is by hand and
changes the machine, which is tier 0's definition. But the host states that it *connects to nothing
and listens on nothing*, and that claim is what makes the one thing running forever on every
machine auditable. An installer connects to plenty. Same tier, same delivery, different program.
## Consequences
- **The control plane stops being a special case.** It becomes an ordinary module with an ordinary
image in the mesh's own registry — so the mesh can build and roll out **its own upgrades**, which
is what a mesh that runs itself was always reaching for.
- **The bundle's job shrinks** to raising a temporary control plane exactly once.
- **The source builds the installer; it does not run it.** Cloning moves to a release machine, where
a forge being available is an ordinary working assumption, and leaves the disaster-recovery path
where it very much is not.
- **The lab's disposable registry is deleted.** The lab bootstraps by running the same program a
bare machine runs — the only arrangement in which the installer cannot quietly drift out of truth
again.
- **Two public images remain at genesis** — the store and the broker. An air-gapped install would
embed those too, at a much larger artifact; that is a build variant, not a different design.
- **A control-plane module manifest must exist**, and did not.
- **The registry may require nothing.**
[`04-ISSUES/029`](../04-ISSUES/029-the-artifact-store-cannot-be-delivered-by-the-artifact-store/00-report.md)
states that a module providing the artifact store may not *build* artifacts, because there is
nowhere to put them until it runs. The pivot shows that is the narrow case of a wider rule: **it
may not require anything the store is needed to deliver.** Found the hard way — an unrelated
change gave the registry a public name and, with it, a route requirement. At genesis nothing
provides a route, and nothing can, because the routing stack needs images and images need the
store. The same cycle, re-entered through a door the existing wording did not cover.
- **The handover is the sharp edge.** For one moment the bundle and the module both describe the
same container, and the host tracks what it owns. If a safe handover is not expressible with what
exists, the install **stops before it** and says what is missing. A machine left without a control
plane cannot be fixed remotely, so a partial install that halts cleanly is the better outcome.
**How each is checked** — an unenforced rule is indistinguishable from a wrong one:
- **Naming by its own digest:** a machine that can reach no registry at all raises a control plane.
- **The pivot completed:** after installing, the running control plane's image is pinned by a digest
**the mesh's own registry assigned** — not by an image id. If it is still the image id, the pivot
did not happen and the mesh cannot upgrade itself.
- **No fiction left in the lab:** the scenario declares no registry machine, and the bed bootstraps
through the installer rather than around it.
- **The handover:** a machine whose control plane has been replaced still has one, and it answers.
- **The registry requires nothing:** its manifest is resolvable on a mesh that has no other module
in it. A requirement added to it later is caught where it is written, rather than by a genesis
that cannot complete — which is how this one was found.
## References
- [ADR 0006 — the substrate and the control plane](0006-the-substrate-and-the-control-plane.md),
which pinned images by digest and named what a first node fetches.
- [ADR 0005 — the node host](0005-the-node-host.md), which makes tier 0 the one thing installed by
hand and the only thing that changes a machine — the property this keeps true by shipping the
installer beside the host rather than inside it.
- [`04-ISSUES/029`](../04-ISSUES/029-the-artifact-store-cannot-be-delivered-by-the-artifact-store/00-report.md),
the same cycle one layer down, and the rule that a registry module is named and never built.
+112
View File
@@ -0,0 +1,112 @@
---
topic: building it
status: proposed
date: 2026-09-12
deciders: jochen
reconstructed: false
extends: 0016-the-lab.md
---
# 68. The lab takes requests, one at a time, and runs each from its own copy
## Context
**The lab is exclusive hardware, and today a person holds it.** Raising a scenario takes over
addresses and names on the workstation for as long as it stands, and only one scenario can stand
at a time. So a run is not merely slow — it occupies the machine and the person who started it,
who then waits rather than works.
**Running it in the background against the working copy is worse than waiting.** The obvious fix
is to start a run and carry on editing. But a run reads the working copy as it goes: binaries are
rebuilt from it, manifests are read out of it, and the bed's own code is loaded from it. Edit
while it runs and the result describes a state that never existed — a mixture of what was there
when each file happened to be read. A green result obtained that way is not evidence, and a red
one costs a day to disbelieve.
**Nothing today records what was asked for.** A run is a command line in somebody's terminal. What
commit it exercised, what it was trying to find out, and what it answered all live in scrollback,
which is why the same question gets re-run rather than looked up.
**Most of the parts already exist.** The lab writes a receipt of its last run. The mesh already
carries messages between nodes and can notify a person. The machine already runs work on a
schedule. What is missing is the thing in the middle.
## Considered Options
**1. Leave it as it is — a person drives the lab and waits.** Rejected. It is the loop
[ADR 0010](0010-delivery.md) removed everywhere else, kept here by habit rather than by argument,
and the cost compounds: because a run is expensive to start and blocks the person, fewer are run,
so faults are found later and in larger batches.
**2. Run in the background against the working copy.** Rejected on the reasoning above. The
failure is silent, which is the kind this repository exists to refuse.
**3. Put the lab behind the ordinary build pipeline.** Rejected for now. The pipeline builds
artifacts and does not own a machine that can raise virtual machines; giving it one makes the
pipeline's slowest job the lab's, and couples every push to hardware only one machine has. This
may become right later; it is not the smallest thing that works.
**4. A queue in front of the lab, and an isolated copy behind it.** Chosen.
## Decision
**The lab accepts requests rather than commands.** A request is recorded, queued, and answered.
The person who made it is told when it is answered and does not wait.
**A request names a bed and a commit, and nothing else.** This is the load-bearing restriction. A
request may say *run this bed, at this version of these repositories*. It may not say what to
install, on which machine, or with which settings — because a request that could say those things
would be a second way of installing a mesh, and the whole reason the installer exists is that the
lab already was one ([ADR 0067](0067-genesis-is-a-pivot.md)). The bed decides what is installed;
the request only decides which bed and which version.
**Requests are released one at a time.** The hardware admits one standing scenario, so the queue
enforces what the hardware already requires, rather than leaving it to whoever remembers.
**Every run happens in a copy the lab owns.** The lab checks the requested commit out into its own
path and builds and runs from there. A working copy is never read by a run. This is what makes the
queue safe to use while work continues, and without it the rest of this record is not worth
having.
**The lab is reached through tools, not only a command line.** A command line is available only
to whoever is sitting at the machine, which is the constraint this record exists to remove. The
lab answers three questions to anything that can reach the mesh — *what is standing now*, *what is
queued or running*, and *what did this request answer* — and accepts a request and a cancellation.
An agent can therefore start a run, stop attending to it, and come back; and somebody who did not
start a run can still see it, which is the difference between a shared lab and a private one.
**The restriction holds at every door.** A tool submits a bed and a commit, exactly as a command
line does. A tool that could name a module, a node or a setting would reintroduce the second
installer through a different entrance, and the entrance is not what made it dangerous.
**Every run leaves a record that outlives the terminal**: what was asked, which commit, when it
ran, what it answered, and where its output went. A question already answered is looked up rather
than re-run.
## Consequences
Work continues while the lab runs, which is the point. A second session may edit freely, because
nothing it edits is what the lab is reading.
A request is reproducible by construction: it names a commit, so the same request can be asked
again and compared. Today two runs of "the same thing" are only as alike as the tree happened to be.
The lab gains a second copy of every repository it exercises, costing disk and needing to be kept
from drifting into a place people edit by hand.
Anything that can reach the mesh can now see what the lab is doing, including an agent working on
something else. That is the intended gain and also the obvious hazard: a thing that is easy to ask
is easy to ask too often, and the hardware still admits one scenario at a time.
The queue becomes a thing that can fail — stuck, backed up, or lost — and a queue nobody watches
is worse than no queue, because it absorbs requests silently.
## How this is checked
| Rule | Checked by |
|---|---|
| A run never reads a working copy | The runner is given a path it owns and no other; a run started while a working copy is deliberately dirtied produces a result matching the commit, not the edits. |
| One scenario stands at a time | A second request submitted while one runs is observed to wait, not to raise. |
| A request cannot say what to install | The request format admits a bed and a commit only. A request naming a module, a node or a setting is refused, and the refusal is exercised. |
| A request is answered | Every queued request reaches a terminal state with a record. A request that vanishes is a failure of the queue, not a quiet nothing. |
| The lab can be asked from elsewhere | What is standing is asked from a session that did not raise it, and the answer matches the machine. A lab that only answers its own caller has not left the terminal. |
@@ -0,0 +1,89 @@
---
topic: building it
status: accepted
date: 2026-09-12
deciders: jochen
reconstructed: false
extends: 0009-modules-and-the-graph.md
---
# 69. A module is a repository and a path within it
## Context
**The builder clones one repository and reads `module.json` at its root.** The to-be design says
so in as many words — *"one file at the root"* — and the code implements it: clone, read the root
manifest, build what it declares.
**Nothing that exists is shaped that way.** The catalogue holds sixty-seven modules, each in its
own directory, and has no manifest at its root. None of the five code repositories has one either.
So today the builder cannot be asked to build any module that exists: pointed at the catalogue it
finds no manifest, and pointed at a module's source it finds no manifest.
**The system being replaced already works the other way**, and has for years: a monorepo with one
directory per piece of software, and the coordinator builds a module from a repository and a path
inside it. The root-only assumption is not a simplification of that — it is a different model that
was never reconciled with it.
**And it splits what a build needs into two places.** The control plane's manifest sits in the
catalogue; the source it describes sits in the control plane's own repository. A build must read
one tree, so under the root-only model neither location can be built from.
## Considered Options
**1. One repository per module.** Rejected. Sixty-seven repositories for sixty-seven modules, most
of which are a single manifest naming a public image, and every one needing its own creation,
permissions and lifecycle. It also contradicts [ADR 0015](0015-applications-live-in-their-own-repository.md),
which put *applications* in their own repositories precisely because modules do not need one.
**2. Keep manifests in the catalogue and source elsewhere, and have a build fetch both.** Rejected.
A build would clone two trees whose versions can disagree, so "what commit is this module?" stops
having one answer — and that question is the whole basis of knowing when to rebuild.
**3. A module is a repository and a path within it.** Chosen. It is what the current system does,
what the catalogue already looks like, and it keeps a module's description beside the thing it
describes.
## Decision
**A module is named by a repository and a path within it.** The path holds `module.json`, and
everything that manifest declares is produced from that path. A module whose path is the root is
the ordinary case of this, not a separate one.
**A module's manifest lives beside its source.** Where a module has code, its directory holds both,
so one commit answers "what is this module, and what is it made of". Where a module has no source —
a manifest naming a public image — the directory holds only the manifest, and there is nothing to
build.
**This moves the core modules.** The control plane and the builder are built from the control
plane's repository, so their manifests belong in that repository at their own paths, not in the
catalogue. The catalogue keeps the modules whose source it holds, and the modules that are only a
manifest.
**One commit, one module version.** Because a module is one path in one repository, the commit that
built it identifies it exactly, and "the source has moved ahead of what the mesh holds" stays a
question with a yes or no answer.
## Consequences
The builder gains a path alongside the repository and the ref. A build is `repository, path, ref`,
and the manifest it returns is the module the mesh records.
The catalogue stops being the place every manifest lives, and becomes the place manifests live
*when their module has no other home*. That is a smaller claim than it sounds: most of the
sixty-seven stay exactly where they are.
Two repositories change shape — the control plane's gains manifests for the modules built from it.
Nothing else moves.
A repository can hold modules that are built and modules that are not, and no rule distinguishes
them beyond whether their manifest declares anything to build.
## How this is checked
| Rule | Checked by |
|---|---|
| A module is buildable from its repository and path | The builder is asked for a module by repository and path, and returns a manifest whose artifacts are pinned to digests the mesh's registry assigned. |
| A manifest sits beside what it describes | A module declaring something to build, whose path holds no source to build it from, is refused at build time rather than producing an empty result. |
| One commit identifies one module | Two builds of the same repository, path and commit produce the same digests. |
| The core modules are built like any other | The control plane is rebuilt from its own repository and path, and the running mesh is upgraded to it — the same path an ordinary module takes. |
@@ -0,0 +1,107 @@
---
topic: the tiers
status: accepted
date: 2026-09-12
deciders: jochen
reconstructed: false
extends: 0067-genesis-is-a-pivot.md
---
# 70. The catalogue owns the module graph, and genesis builds rather than carries
## Context
**The module graph has no owner.** What modules exist, what each requires and provides, what each
claims, what each is made of — all of it lives inside the control plane because that is where it
was first written, not because anything decided it belonged there.
**The control plane's own test says it does not belong there.**
[`06-the-control-plane`](../03-DESIGN/01-to-be/06-the-controller.md) defines the tier as
*everything that needs to know about more than one node*, and states the corollary plainly:
anything a single machine could answer alone is not the control plane's. What a module is, and
what it needs, requires no knowledge of any node whatsoever.
**And nothing can query it.** The graph is the thing that answers *what must be rebuilt when this
changes*, *what would break if this were removed*, and *what can be installed here* — and today it
is reachable only as control-plane internals. [ADR 0009](0009-modules-and-the-graph.md) already
decided that a build edge is derived, that an artifact is stale when anything it was built against
moved, and that the rebuild set is therefore computable. None of that has anywhere to live.
**Genesis currently carries an image, and cannot produce a builder at all.**
[ADR 0067](0067-genesis-is-a-pivot.md) has the installer carry the control plane's image. That
works, but it leaves the builder with no route onto a fresh mesh — it cannot be fetched from the
public internet, because the mesh builds it, and the installer carries one image only. So a raised
mesh cannot build anything, including the modules it is made of.
## Decision
**The catalogue is a core module, beside the control plane and the builder, and it owns the module
graph.** Modules, what they require and provide, what they claim, their dependencies and their
build edges, assignments and configuration — the catalogue holds them and serves tools over them:
install a module, query what is available and what it needs, query the graph, query and update
settings.
**It is one per mesh**, expressed the way the control plane already expresses it — a claim scoped
to the mesh, not a new mechanism.
**The control plane consumes it.** Resolving what a module requires into an actual binding, and
composing what a machine should be, both need to know what modules are. So the dependency runs
from the control plane to the catalogue, which is the opposite of what the tiers suggest and is
therefore written down here rather than left to be inferred.
**Genesis builds the core modules rather than carrying them.** The installer ships an *init
builder* — the one thing carried — which is started, clones the source, and builds the control
plane, the catalogue and the builder. The installer then raises a temporary control plane, which
installs the catalogue, registers the permanent control plane and the builder, and assigns all
three to the first machine. The temporary control plane stops; the installer verifies that the
permanent one answers and can query the graph. The builder then sees a catalogue in its initial
state and builds the core modules into it.
**So exactly one thing is carried, and it is a builder rather than a result.** That is the
difference from [ADR 0067](0067-genesis-is-a-pivot.md), which carried the control plane's image:
carrying a builder produces every core module on the machine, including the builder itself, so
there is no component left without a route.
## Consequences
The catalogue joins the small set of things that cannot arrive through the ordinary path, because
it cannot be installed by something that needs it in order to install anything. It arrives the same
way everything else does under this record — built by the init builder before the mesh can install
anything — so the set is answered by one mechanism rather than three special cases.
The rebuild fan-out gains a home. *What was this built against* and *what must rebuild now* are
questions about the graph, and the graph now has an owner to hold the edges and answer them.
The catalogue holds state, so it owns a store in the substrate's database, the same way the control
plane's contexts do. That is the mesh's own store and not the `postgres` module, which is a
provider other modules consume.
A mesh without a catalogue cannot resolve anything, where previously it merely lacked an interface.
That is the cost of ownership over surfacing, and it is deliberate: one owner beats two copies.
## Open, and to be settled before this is built
**Where the init builder clones from.** [ADR 0067](0067-genesis-is-a-pivot.md) rejected building
from source at genesis partly because the source lives in a forge that runs on the mesh, so a
total rebuild would need the mesh it is rebuilding. Carrying a builder answers the toolchain half
of that objection and not this half. Genesis must therefore name a source that exists before the
mesh does.
**What the init builder publishes into.** Building produces artifacts that must be pinned by a
digest a registry assigned, and today the registry is installed after the machine has joined. If
the core modules are built first, the registry has to exist first, so the substrate's order needs
restating rather than assumed.
**Where the line falls between the catalogue and the control plane.** Claims and assignments need
to know about every node, which by the control plane's own test is its work. Whether the catalogue
holds them and asks, or the control plane holds them and the catalogue surfaces them, is not
settled here.
## How this is checked
| Rule | Checked by |
|---|---|
| The catalogue owns the graph | The control plane answers *what does this module require* by asking the catalogue, and a mesh whose catalogue is stopped cannot resolve — observed, not assumed. |
| One per mesh | A second catalogue assigned anywhere in the mesh is refused by the claim, and the refusal is exercised. |
| Genesis builds rather than carries | The installer carries exactly one artifact, and after installing, every core module is pinned to a digest the mesh's own registry assigned. |
| The builder has a route | A mesh raised by the installer, with no hand-placed image, can build a module. |
@@ -0,0 +1,80 @@
---
topic: the tiers
status: accepted
date: 2026-09-12
deciders: jochen
reconstructed: false
extends: 0070-the-catalogue-owns-the-module-graph.md
---
# 71. Genesis clones from a mesh, and checks what it got
## Context
**[ADR 0070](0070-the-catalogue-owns-the-module-graph.md) has the init builder clone the source,
and does not say from where.** [ADR 0067](0067-genesis-is-a-pivot.md) had already rejected building
at genesis partly for that reason: the forge holding the source runs *on* the mesh, so a total
rebuild would need the mesh it is rebuilding.
That objection is real but narrower than it reads. It only binds when the mesh being raised and the
mesh holding the source are the same one, which is true exactly once.
## Decision
**Genesis clones from a mesh's forge, reached by name.** Any mesh that holds the source can serve
it. The first mesh is not structurally special — it is simply the only one that existed when there
was nothing else to clone from.
**If the mesh serving the source is lost, the name moves to another mesh that holds a copy.**
Recovery is a name pointing somewhere else, not a backup being restored. This is what makes the
source's survival a property of there being more than one mesh, rather than a property of somebody
having remembered to take a copy. A mesh that has installed from that name holds the source
afterwards, so every installation adds a place the name could point.
**Genesis names a commit and checks what it got.** It does not clone whatever a branch happens to
point at. The forge a mesh installs from is the trust anchor for everything that mesh will ever
run, and a branch is a moving target that somebody else controls.
*This is not hypothetical.* On 2026-09-11 the forge that would serve this role was running a
cryptominer, and its git operations were being tampered with in flight — output injected into the
protocol stream by a hook that fired on every fetch. Nothing was altered: the repositories were
verified against local copies and found byte-identical. But a mesh installing from that name during
those hours had no way to establish that for itself, and would have had none.
## Consequences
The init builder needs a name it can resolve and a commit it can verify, and nothing else. It does
not need to know which mesh answers.
Whoever operates the mesh that name points at carries a responsibility to everyone installing from
it, and should know that. It is not merely a convenience host.
A mesh that cannot reach any forge cannot be raised. That is a real limit and it is accepted: the
alternative is carrying the whole source in the installer, which makes the installer a release
artifact that goes stale rather than a program that fetches what it was told to.
## Open — what relationship a mesh keeps afterwards
**Not decided, and named here so it is not decided by accident** by whoever writes the init
builder. Two shapes, and they are meaningfully different:
**A snapshot, and then independence.** A mesh installs once, mirrors the source into its own forge,
and has no upstream afterwards. It is fully self-hosted, in the sense that nothing it needs lives
anywhere else. Updates are then something an operator does deliberately, by pulling changes in —
tooling for which is possible and is not a priority.
**A continuing upstream for core modules**, the way a distribution serves packages and a separate
collection serves everything else. A mesh keeps looking at the origin for the modules that make a
mesh a mesh, and holds its own for the rest.
The first is more obviously aligned with the rest of this design, which is arranged so nothing a
mesh needs depends on somebody else continuing to host it. The second is more convenient and makes
a security problem in one forge everybody's problem. Neither is chosen here.
## How this is checked
| Rule | Checked by |
|---|---|
| Genesis needs only a name and a commit | A mesh is raised with the name pointed at a different mesh than the last time, and the result is identical. |
| What was cloned is what was asked for | Genesis is pointed at a commit and refuses a forge serving different content under it, rather than building what it received. |
| Losing the serving mesh is survivable | The name is repointed at a mesh that installed from it earlier, and a raise succeeds. |
@@ -0,0 +1,86 @@
---
topic: the tiers
status: accepted
date: 2026-09-12
deciders: jochen
reconstructed: false
extends: 0070-the-catalogue-owns-the-module-graph.md
---
# 72. Two graphs, and a build chain that orders itself
## Context
[ADR 0070](0070-the-catalogue-owns-the-module-graph.md) gave the catalogue the module graph and
said the control plane consumes it — that resolving a requirement and composing what a machine
should be *"both need to know what modules are, so the dependency runs from the control plane to
the catalogue."*
**That paragraph is wrong, and this record corrects it.** It was written before the two graphs had
been told apart, and it creates a dependency that does not need to exist: a catalogue that is down
would leave the control plane unable to compose the declaration that would repair it.
## Decision
**There are two graphs, with different owners, and they meet only when something is installed.**
| Graph | Owner | What it links | Answers |
|---|---|---|---|
| the module graph | the **catalogue** | module-versions to each other | what was this built against · what must rebuild now · what does this need |
| the runtime graph | the **control plane** | module-versions to nodes | what runs where · who consumes this provision · what breaks if this machine goes |
The catalogue does not know nodes exist. The control plane holds module-versions and nodes, along
with capabilities and claims, because deciding whether a machine qualifies needs every node.
**So the control plane never asks the catalogue anything.** It holds what it needs to compose a
declaration. A catalogue that is down stops new installs and stops the rebuild fan-out, and does not
touch anything already running or the ability to repair it.
**Everything between them travels as events over the broker**, like all module-to-module
communication. The builder finishes and announces that a module was built. The catalogue registers
it, places it among what it depends on, and announces that a module was upgraded. The control plane
reacts to *that*, not to build output — a semantic fact rather than an artifact.
**Nothing is lost if a receiver is down.** A consuming module's queue is durable with a dead-letter
exchange, declared by the mesh rather than by the module, so an event waits for a consumer that is
not there.
**Build order is not computed. It emerges from the chain.** The builder never consults the graph: it
builds what it is asked for, one at a time. The catalogue asks for the next build after the previous
registration, so *"do not start this until that is registered"* holds by construction rather than by
a schedule somebody maintains.
**The catalogue's rule is a condition, not an order:** ask for a module to be rebuilt once everything
it was built against is current. That handles a chain and a diamond with one rule, where an
order-based approach needs to know the shape in advance.
**Two refusals belong to the catalogue.** A cycle, because the chain would never settle. And a
rebuild whose artifacts are identical to what it replaced, which is not an upgrade and must not be
announced as one — or a single change ripples outward forever through modules that did not change.
**Genesis does none of this.** The init builder has a fixed, short list — control plane, catalogue,
builder — in a written order, because there is no catalogue yet to ask.
## Consequences
The builder stays simple, and independent of the catalogue. That is what makes genesis possible at
all: the thing that builds the catalogue cannot require the catalogue.
The two sides can be briefly out of step — the catalogue may hold a module-version a moment before
the control plane knows of it. Assigning in that instant fails, and should say why rather than
report that no such module exists.
A module's declared events stop being documentation and become its permissions: an account is
scoped from what a module emits and consumes, so a builder that announces what it built is granted
what it needs by the ordinary mechanism rather than by a special case.
## Open
**Whether an upgrade is applied or merely noticed.** Today the system this replaces deploys
automatically, and that is a defensible default for core modules on a mesh its operator runs. But
the design as it stands does the opposite: it records that the source moved ahead, makes it visible,
and waits to be told. This must become a setting with a chosen default rather than inherited
behaviour — and the choice matters most on the day a bad commit reaches something that carries mail.
**Whether a module assigned to several machines upgrades on all of them at once.** Doing so makes
one bad commit simultaneous everywhere. Doing one machine and pausing turns it into one casualty.
@@ -0,0 +1,94 @@
---
topic: the tiers
status: accepted
date: 2026-09-13
deciders: jochen
reconstructed: false
extends: 0070-the-catalogue-owns-the-module-graph.md
---
# 73. The installer carries a builder, and the registry stays where it is
## Context
[ADR 0070](0070-the-catalogue-owns-the-module-graph.md) decided that genesis builds rather than
carries, and [ADR 0071](0071-where-genesis-gets-its-source.md) settled where it clones from. Two
questions were left open, and the design record names them as the one gap that stops a fresh mesh
from being able to produce anything at all: **how the builder arrives**, and **what it publishes
into**.
Today the installer carries the control plane's image inside itself. That works, and it is why
genesis needs no registry: nothing is ever fetched, because the one image that matters is already
present. The cost is that the mesh which results holds an artifact it did not make, cannot rebuild,
and knows nothing about — no version, no source, no edges. That is the same shape as the fault
[issue 044](../04-ISSUES/044-the-runtime-every-module-builds-on-cannot-be-built-by-the-mesh/00-report.md)
recorded for the shared runtime, and fixing it there while shipping it here on every new mesh would
be a strange place to stop.
## Decision
**The installer carries a builder, and nothing else.** One artifact, not a growing set. It clones
the source at a named commit, checks what it got ([ADR 0071](0071-where-genesis-gets-its-source.md)),
and produces the control plane from the same repository and path that any later rebuild of it would
use. What raises the mesh is therefore the same thing that will maintain it, and there is no second
mechanism kept in step with the first.
**The registry does not move, and the argument for moving it does not survive being made.**
It was put this way: a produced image has to be put somewhere before anything can fetch it, so the
registry must now precede the control plane, and
[ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md)'s answer to that question has to flip.
It does not, because the premise is false. **The thing that builds the image and the machine that
runs it are the same machine.** A built image is already in that machine's container runtime, and
the temporary control plane names it exactly as it names a carried one — by the digest of its own
configuration, a local identity that requires nothing to have served it. Building changes where the
bytes came from. It does not change where they are.
| | is it substrate? | must it precede the control plane? |
|---|---|---|
| the store | yes | yes — there is nowhere else to put the control plane's state |
| the broker | yes | yes — the control plane reaches a machine only over it |
| the image registry | yes — it cannot grant itself a repository | **still no** — the first machine neither fetches the control plane nor needs to, whether the image was carried in or made here |
So the registry stays where [ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md) put it:
substrate by role, ordinary by delivery, installed by the temporary control plane as its first act.
The bundle carries two services and a control plane, as it did. **What publishes into the registry
is unchanged too** — the existing step that pushes the control plane's image into it, which is the
moment that image first receives a digest assigned by something other than itself. It now pushes
something this mesh built rather than something it was handed.
## Consequences
**Genesis gains one step and changes no others.** A build happens before the image is loaded. The
pivot described in [ADR 0067](0067-genesis-is-a-pivot.md) survives exactly as written, because the
step it pivots on never cared where the image came from.
**A fresh mesh can produce from the moment it exists.** The builder is present before the control
plane is, so the core modules, the catalogue and the builder's own module can be built in the
ordinary way rather than waiting for somebody to carry them in. The paragraphs in
[`17-raising-a-mesh`](../03-DESIGN/01-to-be/17-raising-a-mesh.md) that describe this were describing
something that could not start; they can start now.
**Genesis needs more of the outside world.** Carrying an image needed nothing but the installer.
Building one needs the source, and whatever the build itself reaches for. This is a real cost and
is not waved away: it makes genesis fail in more ways, all of them at a step that says what it was
doing. It is accepted because the alternative is a mesh that cannot rebuild its own control plane,
which fails in exactly one way, silently, later, and for ever.
**A pre-built bundle remains possible and is not this.** Nothing here forbids delivering artifacts
rather than building them; it fixes where they may come from. A bundle of pre-built core modules is
an **export of a mesh that built them**, carrying what the catalogue knows about each alongside the
artifact itself — so that loading one leaves the graph in the state building would have left it. A
bundle that carries images without that is the thing this decision rejects, whoever ships it.
## What this does not decide
**Whether the builder's own module is carried or built.** It builds everything else; what installs
*it* as an ordinary module afterwards, so that it too can be upgraded, is the same closed-list
question [`12-a-module-repository`](../03-DESIGN/01-to-be/12-a-module-repository.md) already holds,
and is unchanged by this.
**How a machine authenticates to a registry that asks it to.** Genesis raises its own and reaches it
over the loopback, so this remains a joining problem
([issue 042](../04-ISSUES/042-nothing-gives-a-node-an-account-for-a-registry/00-report.md)).
@@ -0,0 +1,172 @@
---
topic: the tiers
status: accepted
date: 2026-09-15
deciders: jochen
reconstructed: false
extends: 0039-what-the-sdk-holds-and-refuses.md
---
# 74. The mesh defines a module protocol; an SDK is an implementation of it
## Context
[ADR 0039](0039-what-the-sdk-holds-and-refuses.md) says what the SDK holds: the tool-serving
harness, the messaging and event framework, the contracts, and core primitives. It settles what
belongs in *an* SDK. It does not say what happens when there is more than one.
There is already more than one. **The contracts are expressed twice** — as Go types in the control
plane and the host, and as TypeScript types in the SDK — and nobody has felt it because both live
in one repository and one head.
**A correction, made after inspecting the wire rather than the types** (2026-09-16). This record
first claimed the two implementations already disagreed — `resource` vs `Provision`, `consumer`
meaning the module in one and the node in the other, headers declared on one side and emitted by
neither. **On inspection the live wire agrees**, and the claim was wrong:
- The grant types that disagreed (`Grant`, `Interface`, `Credential` in the SDK's `contracts`)
were **dead** — exported and imported by nothing. The live provisioning wire is the contributions
file, whose shape (`as`, `secret`, `node`, `at`, `values`) is the same on both sides. Those dead
types have been removed.
- The envelope agrees too: Go emits all five required headers, and `x-causation-id`/`x-schema` are
**optional** — the SDK sets them when a handler has a causation or a schema, and a bare event
carrying neither is correct, not a drift.
So the danger was never live disagreement. It was **dead types that contradicted the live wire**,
which read as the contract and were not — and are exactly what led this record to assert a drift
that inspection did not find. That is a sharper reason for the decision below, not a weaker one: a
type is only as good as its being the wire, and the way to guarantee that is to specify the wire and
check implementations against it, rather than to trust a hand-kept type to still describe it.
A failure of this kind does not announce itself. Two implementations that disagree about an
envelope do not fail to compile — they ignore each other's messages, and a mesh where a module
stops reacting looks exactly like a mesh where nothing happened.
## The question this settles
A module may be written in any language the mesh can build
([`18-building-a-module`](../03-DESIGN/01-to-be/18-building-a-module.md)). Every language needs an
SDK. What is an SDK *of*?
Two answers were available, and the obvious one is wrong.
**Shared types, generated.** Write the shapes once — a schema, an IDL — and generate Go, TypeScript,
Rust. It is the familiar answer and it solves the smaller half of the problem. The shapes are not
where the difficulty is.
**A specified wire, with a conformance suite.** The shapes are a consequence; what an SDK must get
right is *behaviour*.
## Decision
**The mesh defines a module protocol. An SDK is an implementation of that protocol in one
language, and nothing more.**
That is the whole of what an SDK is. Not a library a language happens to have, not a convenience
layer, not a place for helpers to accumulate — an implementation of a specified protocol, finished
when it implements it and correct when it agrees with every other implementation.
### The protocol is split per capability
**A module does not use all of it, so an SDK need not implement all of it.** A module that only
consumes events uses the event capability. One that serves tools uses the tool capability. A
provider uses provisioning. Nothing about consuming an event requires knowing how a grant is
answered.
So the protocol is a floor plus capabilities:
| part | what it covers | who needs it |
|---|---|---|
| **connection** — the floor | reading the sealed credential, pinning the certificate fingerprint, taking identity from the credential rather than the environment | everything |
| **events** | the envelope and its headers, the durable per-consumer queue, binding, at-least-once with dedup on `x-event-id` | a module that emits or consumes |
| **tools** | registration, the shared durable `serve.<key>` queue, request and reply | a module with a surface |
| **provisioning** | a grant in, a credential out, and what each carries | a module that provides something |
**This is the same shape the host already has.** A host declares which resource kinds it can apply,
and a partial host — one that can write files and run things but not manage users or containers —
is a real thing rather than a broken one ([ADR 0005](0005-the-node-host.md)). An SDK that implements
the floor and events is exactly as legitimate, and a module written against it is a module that
does events.
**So a language arrives in pieces rather than all at once.** A Rust SDK implementing connection and
events is useful the day it exists; tools and provisioning follow when something needs them. The
alternative — a language is unsupported until it is entirely supported — is what makes adding one a
project rather than a contribution.
**And what a language can be used for is then a fact the mesh can state**, rather than something an
author discovers by writing a module that cannot be built: the toolchain list says which languages
exist, and the conformance results say what each can do.
### What the specification covers
Per capability, what two implementations can disagree about:
- **the exchanges and queues** — which exchanges exist, that a consumer's queue is durable and
named `<node>.<module>.events`, that a tool is served from a shared durable `serve.<key>`
- **the envelope** — every header, which are required, what an unknown `x-` header means, and that
ignoring one is correct rather than lax
- **identity** — that a module's node and module name come from its sealed credential and not from
its environment, so what it emits matches what the mesh authorised
- **delivery** — at-least-once, and that dedup is on `x-event-id`, which only the emitter can make
- **the credential** — the sealed document's fields, and that a connection pins a certificate
fingerprint rather than trusting an authority
- **provisioning** — a grant in, a credential out, and what each carries
- **the vocabulary** — that `consumer` is one thing, named once
### Conformance is per capability
**A suite per part, and an SDK claims the parts it passes.** A monolithic pass/fail would make a
partial implementation indistinguishable from a broken one, which is the distinction this is built
on.
**And the suite is executable, not prose.** A specification nobody can run is a document two
implementations drift from while both believe they conform. Conformance is a set of fixtures — an
emitted event, a served tool call, a grant and its answer — that every SDK must produce and consume
byte-for-byte.
**The existing two implementations are the first two to be made to pass it.** Not a future language:
the drift above is present, and a suite that only new SDKs must satisfy would leave the disagreement
that already exists in place while certifying everything added afterwards against it.
## Why not generated types
Generation makes the shapes agree and leaves everything that matters unspecified. Two SDKs
generated from one schema can still name their queues differently, take identity from the
environment, dedup on the wrong field, or omit a header the other requires — and every one of those
is a mesh that runs and quietly does not work.
It also makes the contract into whatever the generator supports, which is a decision nobody made
about a boundary everything else depends on.
**The shapes are worth generating once the wire is specified.** That is a convenience, and it comes
second.
## Consequences
**A language is a commitment, and now a divisible one.** Adding one means implementing the protocol
and passing the suites for the parts it claims. That is more work than transliterating types, and
it is the work that was always there — the difference is that it can be finished, and finished in
pieces, rather than believed.
**Versioning becomes possible.** `x-schema` exists for it and is never written. A specified envelope
with a version on the body is what lets a mesh hold a module built against an older SDK, which is
the ordinary state of any mesh that has been running for a while.
**The two current implementations agree on the live wire** — inspection showed it. What was wrong
was a set of dead types beside the wire, now removed. The suite's job here is therefore prevention:
to keep that agreement true as the wire changes, and to hold a new language's SDK to it, rather than
to repair a break that exists today.
**This does not make the mesh polyglot by itself**, and should not be reported as though it does. It
makes polyglot possible to do correctly. A Rust SDK is still a Rust SDK.
## How this is checked
| Rule | Checked by |
|---|---|
| One vocabulary | A word means one thing across implementations, checked by the fixtures using it. |
| The wire is what is specified | Both existing SDKs run the conformance suite in their own test suites, and a change to one that breaks a fixture fails there rather than in a mesh. |
| A new SDK is a passing SDK | A language is not listed as buildable for a capability until its SDK passes that capability's suite; the toolchain list and the conformance results name the same set. |
| A partial SDK is a real thing | An SDK implementing the floor and one capability passes, is listed for that capability, and a module using another is refused with the reason — rather than failing at runtime in a language nobody said was finished. |
| An unknown header is ignored | A fixture carries one, and every implementation accepts it. |
| Identity comes from the credential | A fixture sets an environment that disagrees with the credential, and the emitted event carries the credential's. |
@@ -0,0 +1,129 @@
---
topic: the tiers
status: accepted
date: 2026-09-15
deciders: jochen
reconstructed: false
extends: 0014-no-npm-workspace.md
---
# 75. An artifact store is a provision; a package registry is a different one
## Context
Two questions have been circling, and they turn out to be one question asked twice.
**"Should gitea be the mesh's registry?"** It serves OCI images and a dozen package ecosystems, it
is already needed — genesis clones from one — and the mesh's own registry has neither
authentication ([issue 042](../04-ISSUES/042-nothing-gives-a-node-an-account-for-a-registry/00-report.md))
nor a transport a runtime will accept over a network
([issue 048](../04-ISSUES/048-nothing-makes-a-machine-trust-the-mesh-registry/00-report.md)). Gitea
has both.
**"Where does the SDK come from?"** [ADR 0014](0014-no-npm-workspace.md) already answers it — each
module consumes its dependencies, the mesh's own shared library included, *from the private
registry* — and nothing installs one, so today it comes from a git URL, which is
[issue 053](../04-ISSUES/053-the-sdk-is-pinned-twice-and-the-two-disagree/00-report.md).
**The framing that dissolves both:** `artifact-store` is already a provision, and `registry`
already provides it. So "should gitea be the registry" is not a question about replacing a
component. It is a question about **a second provider of an existing provision** — which this mesh
has a mechanism for, and uses for certificate authorities and VPNs already.
## Decision
**Two provisions, because they are two jobs.**
| provision | is | for |
|---|---|---|
| `artifact-store` | content-addressed blobs, pinned by digest, no versions, no ranges | what the **mesh** delivers to **machines** |
| `package-registry` | an ecosystem's own registry — npm, cargo, PyPI, Go | what **code** resolves when it is compiled |
They are not the same store with different clients. One is addressed by digest and immutable by
construction; the other is addressed by name and version, and resolves ranges. Conflating them is
how a mesh that pins everything ends up rebuilding one commit into two different things.
**`registry` remains the provider genesis installs.** Not because it is better, but because of what
it is: a directory and one container, no database, no control plane, installable at step 8 of an
install where neither exists yet. Gitea needs a store and provisioning, which means a control plane,
which means the pivot has already happened — and the pivot needs somewhere to publish to.
**Gitea also provides `artifact-store`, and a mesh may choose it.** Two providers of one provision
is a thing the mesh understands: it refuses, names both, and choosing is assigning the one you want.
**And it does not claim `the-artifact-store`.** That claim is node-scoped, so a module holding it
cannot share a machine with another that does — and a machine running gitea for git and packages
*alongside* a registry serving artifacts is an ordinary arrangement, not a conflict. They are
different ports doing different jobs.
The exclusivity that matters is mesh-wide and is already expressed: `provides` at mesh scope means
two providers are two answers, and the resolver refuses until one is assigned. Forbidding
co-residence adds nothing to that and forbids something reasonable. **Whether `registry` should
still hold that claim is left open here** — it may be protecting something about the port or the
data directory that is not written down, and removing a claim is not a thing to do from the outside
of a manifest.
A mesh that assigns gitea gets authentication and TLS for its artifacts — which is to say, **issues
042 and 048 are answered by choosing a provider that already solved them**, rather than by
reimplementing accounts and certificates in a registry that has none.
**Gitea provides `package-registry`.** That is ADR 0014's private registry, and it is one service
rather than one per ecosystem. `verdaccio` may provide it too, for npm alone, and is then a choice
somebody makes rather than the answer.
**The registry is not removed at the end of installing.** A mesh that never runs gitea still has an
artifact store. Retiring it is a migration a mesh performs, not a step an installation ends with.
## Why not simply gitea, from the start
Because genesis would need a control plane before the thing that stores the control plane's image,
and that is circular rather than merely awkward. It would also make one of the three things the
build loop cannot produce for itself into a stateful application with a database — the pivot is the
hardest part of this design already.
And it puts every artifact in the service that is also the trust anchor for everything the mesh will
ever run ([ADR 0071](0071-where-genesis-gets-its-source.md)), which records that forge serving a
cryptominer with tampered git operations. Two blast radii are better than one.
## Moving from one provider to the other is a designed act
**Not a removal.** Every image a machine runs is pinned to a digest at a named store, the control
plane's own included. Changing the provider means:
1. gitea installed, reachable, and holding an account the builder may publish with
2. every artifact mirrored
3. every declaration re-pinned, the control plane's **last**, because it is what performs the others
4. **every machine verified to have converged and to be able to pull from the new store**
5. only then the old provider unassigned, and its volume kept ([ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md))
Step 4 is the one that is easy to skip and the only thing between this and a mesh that cannot
restart its own control plane. A machine that reboots mid-migration pulls from a store that no
longer exists, and a local image cache hides that until exactly the moment it matters.
## Consequences
**The bootstrap is unchanged**, which is the point of keeping the small provider.
**042 and 048 gain a second answer.** They can be fixed in the registry, or dissolved by choosing a
provider that already has accounts and TLS. The second is less work and more service.
**ADR 0014 becomes satisfiable.** There is a provision for the private registry, something that
provides it, and a module may depend on it — so the SDK can be published and consumed rather than
cloned, and issue 053 has somewhere to go.
**A mesh can be minimal or complete, and both are legitimate.** One with the small registry and no
gitea builds and runs modules and cannot serve packages. That is a real configuration, not a broken
one — the same way a partial host is real.
**And the bootstrap still has no package registry.** The first build of the shared base happens
before anything has installed one. That is the same pivot as everything else and it is **not solved
here**: it is named, so the next person does not discover it.
## How this is checked
| Rule | Checked by |
|---|---|
| Genesis needs no database | The installer raises a mesh of one on a machine with nothing, and the artifact store it installs has no store of its own. |
| Two providers are a choice, not a conflict | A mesh holding both is asked to resolve `artifact-store` and refuses, naming both, until one is assigned. |
| Providers may share a machine | A node is assigned both gitea and a registry, and both run — only one of them answers `artifact-store`. |
| The two stores are not interchangeable | A module depending on `package-registry` is not satisfied by `artifact-store`, and the refusal says why. |
| A migration is verified before it is finished | The old provider cannot be unassigned while any machine's declaration still names it. |
@@ -0,0 +1,87 @@
---
topic: building it
status: accepted
date: 2026-09-16
deciders: jochen
reconstructed: false
extends: 0075-two-stores-and-which-provides-what.md
---
# 76. The SDK is a published package, and the toolchain resolves it by version
## Context
[ADR 0014](0014-no-npm-workspace.md) decided a module consumes its dependencies — the mesh's own
shared library included — from the private registry. [ADR 0075](0075-two-stores-and-which-provides-what.md)
decided the private registry is a `package-registry` provision, and that gitea provides it. What
neither settled, and what [issue 053](../04-ISSUES/053-the-sdk-is-pinned-twice-and-the-two-disagree/00-report.md)
left open, is the one build where the rule cannot simply be obeyed: **the first one.**
The TypeScript toolchain image is built *from* the SDK — it carries the SDK so that every module
compiled inside it resolves the shared library without each build fetching it. So the thing that
compiles TypeScript and the thing that contains the SDK were the same object, and that object
cannot be what builds the SDK. Stated as a question — "how does the SDK reach the registry before
the toolchain exists, when the toolchain is what builds it?" — it reads as a paradox.
It is not one. The paradox exists only because the toolchain *bakes a git-cloned copy* of the SDK.
The SDK itself is plain TypeScript: it needs `node` and `tsc` and nothing the mesh makes. A public
base image can build it. The circularity is a property of the workaround, not of the SDK.
## Decision
**The SDK is an ordinary published package in the mesh's `package-registry`, consumed by version.**
The git URL in the toolchain's manifest and the sibling-path lock beside it — the two halves of
issue 053 — are both removed. A build resolves the SDK the way it resolves any dependency, with a
lock that agrees with its manifest, so `npm ci` is the command and reproducibility is by
construction rather than by the machine the build ran on.
**The SDK is built with a public base image, not with the mesh's toolchain.** It is *not* one of
the components the loop cannot build — the control plane, the registry, the builder, the catalogue
([`12-a-module-repository`](../03-DESIGN/01-to-be/12-a-module-repository.md)), which arrive by
carrying an init builder because they are the loop's own machinery. The SDK is machinery for
nothing; it is an ordinary dependency the loop builds and publishes like any other. The only
constraint is narrow: it cannot be compiled *in the mesh toolchain*, because that toolchain is built
from it. So it is compiled on a public base image instead — which needs nothing the mesh makes — and
published before the toolchain that consumes it. It is not carried, because building it does not
wait on a mesh existing first.
**The toolchain base stays, thinned.** mesh-tools remains the image bundles are compiled in and the
one place the SDK is resolved — but it `npm ci`s the SDK by version from the registry instead of
baking a copy cloned from a git URL. Bundles keep borrowing its resolved dependencies; what changes
is that the version they borrow is named and honest. This was the shape chosen over dropping the
shared base entirely and having every bundle resolve the SDK itself: one resolution point, one
place to be right about the version.
**Genesis orders the publish before the first compile.** The package-registry provider is a public
image (gitea), so it comes up needing no toolchain; the SDK is published into it; only then is the
toolchain built, so the first `npm ci` has a registry to read from. Nothing in that chain is
circular, because the only thing that needed the toolchain — baking the SDK — is gone.
## Consequences
Each language's toolchain repeats the shape: its own SDK, built from that language's public base
image, published to the same registry, resolved by version with that ecosystem's lockfile-honest
install (`npm ci`, `cargo` against a vendored or registry source, `pip` against a pinned set). The
warning in issue 053 — that whatever the TypeScript repository does the others will copy — is
answered by making the copied thing the correct one.
A change to the SDK is publish-then-consume, exactly as [ADR 0014](0014-no-npm-workspace.md) already
priced it: publish the new SDK version, then bump the toolchain (and any module pinning it directly)
to consume it. There is no shortcut that resolves an unpublished SDK, which is the property that was
missing.
mesh-tools is no longer an SDK carrier in the sense that mattered — it does not contain a copy
whose provenance is a branch head somebody force-pushes. It contains a version.
A mesh with no package-registry cannot build TypeScript. This is accepted and is not new: it is the
same shape as a mesh that cannot reach a forge being unable to be raised
([ADR 0071](0071-where-genesis-gets-its-source.md)). Installing brings the registry up first.
## How this is checked
| Rule | Checked by |
|---|---|
| The SDK a build compiles against is named, not cloned from a branch | The toolchain manifest pins `@novox/mesh-sdk` to a version, and the build runs `npm ci`, which refuses a lock that disagrees with the manifest. Issue 053's two checks become this one. |
| The SDK builds without the mesh's own toolchain | The SDK's build recipe names a public base image. A recipe that named the mesh toolchain would reintroduce the cycle and is refused in review. |
| The registry is up before the first compile | The genesis bed asserts the package-registry answers, and the SDK is published, before the base build runs. A base build that ran first would fail its `npm ci` with no registry, which is the positive control. |
| A second language repeats the shape, not a new one | When a second SDK is added, its recipe is compared to this one: public base, publish by version, lockfile-honest install. |
@@ -0,0 +1,61 @@
---
topic: the mesh
status: accepted
date: 2026-09-16
deciders: jochen
reconstructed: false
extends: 0006-the-substrate-and-the-control-plane.md
---
# 77. The parts are named controller, foundation, node — not control plane, substrate, master
## Context
The words drifted. In conversation and in code the same thing was called *control plane*,
*controller*, *master*, and *hub*; the store-and-broker pair was called *substrate* and
*foundation*; a machine was a *node*, a *worker-node*, a *peer*, a *slave*. A mesh named
differently by two people is a mesh they describe differently, and the drift was worst on the
parts talked about most.
Three of the terms carried wrong ideas. *Control plane* is borrowed from networking's
control-plane/data-plane split and means nothing here. *Master/slave* and *hub/peer* imply a
subordinate — but no node is: a node applies its own declaration and keeps running when the
control-node dies, so it is as much its own machine as any other. *Substrate* is a biology
metaphor that landed for no one.
## Considered Options
1. **Keep the inherited words.** Rejected: they are the source of the drift, and two of them
(control plane, substrate) are metaphors that teach the wrong shape to anyone reading them cold.
2. **master / slave, or hub / peer, for the nodes.** Rejected: both name a hierarchy the mesh does
not have. The control-node owns no other node; lose it and the rest keep running what they were
last told.
3. **controller / foundation / node + control-node.** Chosen.
## Decision
The component that decides what each node should be, holds the mesh's records, and tells nodes is
the **controller** — the module `mesh-controller`, which claims the mesh-scoped `the-controller`
seat. The store and broker raised at genesis are the **foundation**. Machines are **nodes**;
there are 0..n of them, and exactly one — the one running the controller — is the **control-node**.
Retired: *control plane*, *substrate*, *master/slave*, *hub/peer*, *worker-node*.
[`00-META/glossary.md`](../00-META/glossary.md) is the authority, and a new name for an existing
thing lands there in the change that introduces it in code.
## Consequences
`mesh-control` became `mesh-controller` across the module, container, image, binary, `cmd/` dir and
the git repository; `substrate` became `foundation` in the embedded base bundles, the default
template and the example lock; the seat `the-control-plane` became `the-controller`. The 03-DESIGN
prose and 00-META follow the new words.
What got harder: the records under `02-DECISIONS/` are immutable, so they keep the words they were
written with — this record included, whose own title names what it retires. A term retired here
still appears there, and the glossary is how to read it. The git repository on the forge was
renamed `mesh-control` → `mesh-controller`.
## References
- [`00-META/glossary.md`](../00-META/glossary.md) — one name per thing, and the words retired.
- The rename shipped across all six code repositories and hq (main).
@@ -0,0 +1,63 @@
---
topic: the tiers
status: accepted
date: 2026-09-16
deciders: jochen
reconstructed: false
extends: 0033-the-substrate-is-a-store-and-a-broker.md
---
# 78. The store and the broker are ordinary modules
## Context
[ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md) settled that the foundation is a store
and a broker, raised at genesis; [ADR 0006](0006-the-substrate-and-the-control-plane.md) settled
that the controller cannot grant itself either, because it consumes them and is not running yet to
ask. Both were raised as bundle resources — plumbing, with no record in the mesh's module graph.
That left two costs, named in [issue 051](../04-ISSUES/051-the-mesh-cannot-update-what-it-depends-on/00-report.md).
The foundation's own store and broker could not be upgraded — nothing owned them as modules. And a
mesh that wanted a database or a queue for its modules installed the `postgres`/`lavinmq` modules,
each of which raised a **second** server: a mesh ran two postgres and two brokers.
## Considered Options
1. **Leave them as bundle-only plumbing.** Rejected: they cannot be upgraded, and the second
server stays. The floor keeps a permanent specialty in it.
2. **The control-plane pivot verbatim — raise a temporary one, install the module, retire the
temporary** ([ADR 0067](0067-genesis-is-a-pivot.md)). Rejected for a *stateful* server: it means
a handover with real downtime, tearing down the store the controller is mid-read of.
3. **Adopt in place.** Chosen.
## Decision
The foundation's store and broker are **adopted in place** as the ordinary `postgres` and `lavinmq`
modules. Genesis still raises them first (nothing else can — ADR 0006), then each module declares a
container with the **same name, image and spec** the foundation raised, so the applier — which keys
on the container name and compares a spec digest — reconciles it rather than raising a second. The
credentials are the foundation's, made at genesis and carried in through `secret accept`, because
the mesh cannot invent a credential that already made the databases. The servers bind mesh-wide so
a consumer on any node can reach the one shared server.
A mesh runs **one postgres and one lavinmq**, and each is upgradeable through a stated window: the
store's is a connection-pool reconnect; the broker's is the harder case of recreating the bus the
push travels over, so the mesh reconnects to the one that returns.
## Consequences
The twelve-module floor has no specialty left in it — the store and broker are moments in a
module's life, not a separate kind of thing. Two follow-ups are tracked:
[issue 054](../04-ISSUES/054-the-adopted-store-and-broker-are-open-before-the-filter/00-report.md)
(the adopted servers bind `0.0.0.0` before the packet filter is installed) and
[issue 055](../04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md)
(whether a consumer on another node reaches them over the overlay).
What got harder: a foundation upgrade recreates the very server the controller reads from, or the
bus the instruction to upgrade travels over — a window that a stateless module upgrade does not have.
## References
- [issue 051](../04-ISSUES/051-the-mesh-cannot-update-what-it-depends-on/00-report.md) — the gap this closes.
- [`03-DESIGN/01-to-be/07-the-foundation.md`](../03-DESIGN/01-to-be/07-the-foundation.md) — the amended design.
- Shipped across mesh-host, mesh-catalog and mesh-lab (main); proven 22/22 in the one-node lab.
@@ -0,0 +1,63 @@
---
topic: the tiers
status: accepted
date: 2026-09-17
deciders: jochen
reconstructed: false
extends: 0078-the-store-and-broker-are-modules.md
---
# 79. The foundation seats are named after their servers
## Context
[ADR 0078](0078-the-store-and-broker-are-modules.md) settled that the store and broker are the
ordinary `postgres` and `lavinmq` modules, adopted in place on the control-node, and that a mesh
runs **one postgres and one lavinmq**. But that singularity held only by convention: genesis
assigns them to the control-node alone. [Issue 056](../04-ISSUES/056-an-adopted-module-assigned-to-a-second-node-raises-a-second-server/00-report.md)
recorded the gap — a second `assign` to another node finds no container of that name there and
raises a *second* server, holding none of the first's data, and nothing refuses it.
The mesh already has the mechanism for "there is one of me": a mesh-scoped exclusive **seat**.
[ADR 0077](0077-the-controller-and-the-foundation.md) gave the controller one, which it named
`the-controller`. Nothing gave the store and broker theirs.
## Decision
Each foundation module claims a mesh-scoped seat named **after the server it guards**:
`mesh-controller`, `mesh-store`, `mesh-broker`. So the `postgres` module claims `mesh-store` and
the `lavinmq` module claims `mesh-broker`; the resolver refuses a second holder mesh-wide, the same
way it keeps one hub and one controller. A second `assign` is now a refusal at resolution — *one
per mesh* — not a silent second server.
The controller's seat, which [ADR 0077](0077-the-controller-and-the-foundation.md) named
`the-controller`, is renamed `mesh-controller` under this same rule, so all three foundation seats
follow one convention: the seat is the server. (The module `mesh-controller` and its seat now share
a name, which is the point — there is one of that server, and the seat says so.)
This does not tie a foundation module to the control-node — a mesh-scoped seat forbids a *second*
holder, not a wrong single one. Adoption still requires the container to already be running where
the module lands, which genesis arranges; the seat closes the "two servers" gap, and the
control-node convention remains what puts the one holder in the right place.
## How this is checked
- The resolver's `checkClaims` refuses two holders of a mesh-scoped seat
(`mesh-controller/internal/catalogue/resolve.go`).
- `TestAFoundationModuleCannotBeRaisedOnASecondNode` asserts each foundation module's second
assignment is refused with *one per mesh*.
- The `postgres`, `lavinmq` and `mesh-controller` manifests declare the seat.
## Consequences
"One store, one broker, one controller" is now a property the mesh enforces rather than a
convention it hopes for. The silent operational edge that remains — a cross-node consumer
provisioned only when the provider is pushed again — is separate, and tracked as
[issue 057](../04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md).
## References
- [issue 056](../04-ISSUES/056-an-adopted-module-assigned-to-a-second-node-raises-a-second-server/00-report.md) — the gap this closes.
- [ADR 0077](0077-the-controller-and-the-foundation.md) — named the controller's seat, here renamed.
- [ADR 0078](0078-the-store-and-broker-are-modules.md) — asserted one postgres and one lavinmq, here enforced.
- [`00-META/glossary.md`](../00-META/glossary.md) — seat and claim.
@@ -0,0 +1,55 @@
---
topic: how we work
status: accepted
date: 2026-09-17
deciders: jochen
reconstructed: false
extends: 0019-how-this-repository-works.md
---
# 80. The development cycle is checked, not trusted
## Context
[ADR 0019](0019-how-this-repository-works.md) made this repository the source of truth, and the
process overview drew the flow work must follow: an idea or a symptom, a decision, a to-be design,
a build in a code repository, an as-is update on shipping. The playbooks describe every step, and
frontmatter carries every status.
But the flow itself was enforced by nothing. A design could appear citing no decision; a design
could sit `in-progress` naming no code; an issue could be `fixed` by nobody knows what. Each is
indistinguishable from correct work until somebody reads carefully — and the whole point of the
playbooks is that nobody should have to hold this repository in their head. A session that starts
cold (or an agent after a context clear) must be able to *find* the chain by following frontmatter
pointers, which only works if the pointers are reliably there.
## Decision
The development cycle is enforced mechanically, to the extent frontmatter can carry it:
- **No design without a decision** — every to-be design names at least one record in `decisions:`.
- **No development without a design that says where** — an `in-progress` or `implemented` design
names its owning code in `code:`.
- **No owner-less diagnosis, no fix-less fix** — an issue marked `located` or `fixed` names
`located-in:`; one marked `fixed` or `resolved` says `fixed-by:` (prose counts — "nothing, the
capability existed" is an answer).
- **No silent graduation** — a `graduated` research overview says what it `became:`, and the
targets exist.
[`00-META/checks/cycle.py`](../00-META/checks/cycle.py) refuses violations, beside `records.py`
and `index.py`; all three run before any merge here. What frontmatter cannot see — that code work
actually started from a handoff — remains held by playbooks 04 and 07: a feature branch exists
because a design or an issue sent it, and a merge is a human checkpoint.
## Consequences
A `/clear` costs little: [`AGENTS.md`](../AGENTS.md) now carries the cycle and a where-to-look
table, and the chain a fresh session needs is guaranteed present in frontmatter rather than
reconstructed from memory. The checks are the floor, not the ceiling — they verify pointers exist,
not that their content is true; reading remains the job.
## References
- [`00-META/process/00-overview.md`](../00-META/process/00-overview.md) — the flow, and its new
"The cycle is checked" section.
- [ADR 0019](0019-how-this-repository-works.md) — the repository this disciplines.
@@ -0,0 +1,51 @@
---
topic: how we work
status: accepted
date: 2026-09-17
deciders: jochen
reconstructed: false
extends: 0080-the-development-cycle-is-checked.md
---
# 81. A decision nothing cites is not yet in the chain
## Context
[ADR 0080](0080-the-development-cycle-is-checked.md) made the development cycle checked — but its
checks covered designs, issues and research, not the decisions themselves. Measuring showed why
that matters: 19 of 70 records were cited by nothing — no design doc's `decisions:`, no research
`became:`, no issue, no other record. Among them sat load-bearing decisions (the credential flow,
the module-runtime cluster), and the cost had already been paid once in practice: a stale premise
about an orphaned decision survived in working memory precisely because no pointer led to the
record that had settled it.
## Decision
Every **accepted** decision must be reachable from the cycle: cited by a design doc's frontmatter
(`decisions:` — a governing citation, not a prose mention), a research overview, an issue report,
a `00-META` document, or another record's `extends`/`supersedes` chain.
[`cycle.py`](../00-META/checks/cycle.py) refuses orphans. Proposed records are exempt — a record
under consideration has no home yet — and superseded records are reachable through their
supersession chain by construction.
The 19 orphans were given true homes in the same change: the module-runtime cluster
(0044–0049, 0053–0055) into the connectivity, controller, protocol, writing and model-access
designs; the build decisions (0072, 0076) into the building design; 0021 into the playbook that
implements it; the process records were already reachable once `00-META` counted as a source.
Taken together with 0080, the practice has a name the industry will recognise:
**spec-driven development, with provenance** — a decision is the *why*, the design doc is the
spec, `code:` names the implementation, and the lab beds are the conformance tests. What the
common form leaves implicit, the cycle makes checked: the spec itself must trace to a decision,
and the decision must be findable from the work it governs.
## Consequences
Following pointers now reaches every accepted decision, so a cleared session (or a person) can
trust the frontmatter graph as the whole map. The check is reachability, not truth: a citation
placed wrongly still lies, and reading remains the job.
## References
- [ADR 0080](0080-the-development-cycle-is-checked.md) — the cycle this completes.
- [`00-META/process/00-overview.md`](../00-META/process/00-overview.md) — the flow.
@@ -0,0 +1,84 @@
---
topic: building it
status: accepted
date: 2026-09-17
deciders: jochen
reconstructed: false
extends: 0072-two-graphs-and-the-build-chain.md
---
# 82. The registry is reached by name, and the overlay is its security
## Context
Delivery ends at a node pulling an image, and for every node but the one that built it, that step
has never worked. Three facts conspired, each recorded separately:
- Artifact references are written under `127.0.0.1:5000` — a deliberate parking (the catalogue
says so in the commit that reverted the mesh-reachable name: *"the runtime refused it: 'http:
server gave HTTP response to HTTPS client' … the binding expression returns then"*), so a joined
node is told to pull from its own loopback
([issue 048](../04-ISSUES/048-nothing-makes-a-machine-trust-the-mesh-registry/00-report.md)).
- A container runtime treats any non-loopback registry as HTTPS, and the mesh's registry serves
plain HTTP; nothing the mesh writes tells any runtime otherwise — the one place that file
existed was the lab's, which is why this never failed in a bed (issue 048).
- Nothing gives a node an account for the registry
([issue 042](../04-ISSUES/042-nothing-gives-a-node-an-account-for-a-registry/00-report.md)) —
and nothing has ever said whether one is required.
The tempting fix is TLS from the mesh's own authority. The machinery even exists — the mesh CA
issues leaves for `<node>.internal` names, and a `certificate:` manifest field delivers them. But
the registry is reached **over the overlay**, and the overlay is WireGuard: every byte is already
encrypted and already authenticated to a peer the mesh admitted. The mesh's own code has carried
this position for weeks: *"the registry is reached over the mesh's own private network … a second
layer inside it would be certificates to issue and rotate for no property the first does not
have."* TLS inside the tunnel would also re-order genesis (a certificate needs the network, the
registry precedes it) and put CA handling into three clients (the runtime, the archive fetcher,
the builder) — cost with no new property.
## Decision
**The overlay is the registry's transport security, and the mesh writes the trust it means.**
1. **References name the registry by its mesh name.** The builder publishes to
`<provider>.internal:<port>` (the binding's `at` and `port`) — the parked one-line change
lands. Genesis still publishes to loopback on the first machine, before any overlay exists;
references minted at genesis stay loopback and are valid where they matter — on that machine.
A rebuild re-pins to the mesh name.
2. **Every node on the private network is told the mesh's registry speaks plain HTTP.** The
controller injects, into every such node's declaration, a merged `/etc/docker/daemon.json`
naming `<provider>.internal:<port>` under `insecure-registries`, and a service resource that
restarts the runtime when that file changes — the `/etc/hosts` pattern for the content, the
nftables pattern for the reload. No module author is involved; being on the network is what
grants the trust, because being on the network is what the trust *is*.
3. **No accounts (issue 042), recorded as the position it always was.** Reading and pushing
require presence on the overlay and nothing else. The boundary is enforced, not assumed: the
registry's `listens` is `from: mesh`, the firewall derives from it, and the overlay admits only
peers holding mesh-issued keys. An operator's *external* registry credential is an ordinary
operator-supplied secret (`secret accept`), owned by whichever module names that registry.
Per-node accounts return as a decision, not a patch, if the boundary assumption ever changes —
the shape would be the existing provision flow.
## How this is checked
- The no-fake lab bed: a joined node pulls a mesh-built image by the registry's `.internal` name
over the overlay — an address outside every range the lab's own runtime configuration trusts,
so the mesh-written trust is what makes it work or nothing does.
- The firewall half is generated from `listens` and visible in the node's ruleset (`from: mesh`).
- The plaintext-inside-tunnel position holds exactly as long as the registry is unreachable off
the overlay; the `listens` stanza and the derived ruleset are the check.
## Consequences
A runtime restart when the trust first lands on a node — at joining, before workloads, where it
is free. The registry's HTTP is exposed to whatever stands on the overlay, which is the stated
boundary; a mesh that wants defence in depth inside its own tunnel reopens this record rather
than bolting certificates on quietly. Issues 042 and 048 close on this record.
## References
- [issue 042](../04-ISSUES/042-nothing-gives-a-node-an-account-for-a-registry/00-report.md),
[issue 048](../04-ISSUES/048-nothing-makes-a-machine-trust-the-mesh-registry/00-report.md).
- [ADR 0072](0072-two-graphs-and-the-build-chain.md) — the build chain this completes.
- [ADR 0029](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md) — the overlay as a
boundary.
@@ -0,0 +1,59 @@
---
topic: the mesh
status: accepted
date: 2026-09-18
deciders: jochen
reconstructed: false
extends: 0010-delivery.md
---
# 83. One push leaves the mesh consistent
## Context
A provision is minted while composing the *consumer's* node; the provider's grant list is a pure
read of secrets already issued from it. So assigning a cross-node consumer and pushing its node
produced a consumer that retried forever against a provider that had never heard of it, until the
provider's node was pushed a second time — an action with no signal to take, documented nowhere,
and invisible whenever consumer and provider share a machine
([issue 057](../04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md)).
Two remedies were on the table: **cascade** — a push also delivers to the machines its compose
changed — or **report** — a push says "now push the provider" and leaves the act to the operator.
## Decision
A push finishes what it starts: after composing and sending the named node, the controller flushes
every *other* machine that is now behind — whose declaration differs from what it was last sent —
by name, in the push's own output, converging over a bounded number of rounds (a flushed send may
itself mint).
Behind is measured against what a machine was last *sent*, not against a before/after snapshot of
this push. The mint that makes a provider behind happens when the consumer is assigned or its
account issued — before `push` runs at all — so by push time the provider already differs from
what it holds, with no in-command delta to detect. The only durable signal is "what it should be"
versus "what it last received", which is the same comparison `push --behind` already makes.
A machine behind for an unrelated reason is flushed by this too, and that is correct rather than a
cost: a named push that knew a machine was behind and left it so would be the very silence this
decision removes. The narrower reading — flush only what this push provably changed — was
rejected because it cannot see a mint that a prior command performed, which is precisely the 057
case.
Reporting alone was rejected because it converts a derived fact the controller already holds into
an operator obligation, and an obligation enforced by nothing is issue 057 restated. The
declaration is computed from the whole mesh; delivering a mesh that is knowingly inconsistent and
merely saying so would make "push succeeded" mean less than it says.
## Consequences
- One push is sufficient for a cross-node consumer: the provider's grants arrive from the same
act that minted the provision. The undocumented rule "push the provider node too" ceases to
exist rather than becoming documentation.
- A named push delivers to every machine that is behind, not only the one named — each named in
the output, never silent. `push --behind` remains the way to reconcile the mesh without naming
a node; a named push now carries the same guarantee for the machines its work touched and any
others already waiting.
- How this is checked: the built-store-cross-node bed registers a cross-node consumer, pushes
only the consumer's node, and asserts the provider minted its vhost — the workaround push is
removed, so a regression fails the bed.
@@ -0,0 +1,108 @@
---
topic: what runs on it
status: accepted
date: 2026-09-20
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md
---
# 84. Which provider serves a consumer, when the mesh runs more than one
## Context
[ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) settled what a provision
is *named* for — the thing the consumer's code is coupled to, so `postgres-database` and
`mssql-database` are different provisions and a wrong match is refused at resolution. It said one
thing more, in passing, and left it: *"Two providers of `postgres-database` — a container on this
node and a managed instance elsewhere — are interchangeable and should both match."* Naming was
decided; **which of several providers serves a given consumer was not.**
The mesh assumes there is only one to choose. Provisions are mesh-scoped: the adopted store is a
single `mesh-store` a consumer on any node reaches, and `scope: "mesh"` is written into the
provision definitions. **That assumption is false on day one, and was always meant to be.**
Node-specific services delivered to the mesh is the plan, not an edge case:
- Both control-capable nodes already run their **own** general-purpose relational store (the same
engine, one instance each), serving that node's own applications.
- Each runs its **own** SQL server, its **own** cache, its **own** object store. One node alone
runs six separate relational-store instances, each raised by the module that needed it.
- The one provision that is currently single — the identity provider, one instance on one node —
already authenticates applications whose home is a **different** node.
So several providers of one provision name genuinely coexist, and they are **not** interchangeable
the way 0027's aside supposed. They differ by node, by the data they hold, and by locality. A
consumer bound to the wrong one reads the wrong database, or takes a cross-node hop it did not
need, or cannot be moved without silently rebinding. The model has no field in which to say which
one. This is 0027's own fault — *a match that resolves and is wrong* — one level up: 0027 refused
the wrong **dialect**; nothing refuses, or even asks about, the wrong **instance**.
## Considered Options
1. **Keep `scope: "mesh"` — one provider per provision, mesh-wide.** Rejected: it is false on day
one, and making it true would force every node's applications onto one node's server — the
exact opposite of node-specific services delivered to the mesh, and a single point of failure
the topology was built to avoid.
2. **Resolve to any provider of the name (0027's "both match").** Rejected: when providers hold
different data and live on different nodes they are not interchangeable, and picking one
arbitrarily is a wrong-instance match — the confidently-wrong answer 0027 exists to prevent,
restated at the level of the instance rather than the dialect.
3. **Always require the consumer to name the provider explicitly.** Rejected: needless ceremony in
the common case, where the consumer wants the provider on its own node; and a field every
manifest must carry is a field an author forgets, which then matches everything again — the
failure 0027 warned about for qualifiers.
4. **A provision is node-scoped; the consumer selects the provider, defaulting to co-location.**
Adopted.
## Decision
**A provision is served by a provider identified by its node, and the consumer selects which one.**
A provider is a (node, module) pair, not a mesh-wide singleton. A consumer's binding resolves to a
specific provider, and the selection is part of the assignment
([ADR 0046](0046-a-module-configuration-is-its-assignments-not-its-manifest.md)), not the
manifest.
**The default is co-location.** A consumer that names no provider is served by the provider of
that provision on **its own node**. This is the common case and needs nothing said. A mesh with
one provider of a kind is just the case where co-location and "the only one" coincide — expressed
as *there happens to be one*, not as a scope.
**A consumer coupled to a provider's data names it.** Where two consumers must share one database,
or a consumer must reach a provider on another node, the assignment names that provider — because
that coupling is exactly what may not be guessed, and naming it is what makes a later move safe.
**A module need not consume a provider at all.** It may carry its **own** instance inside its own
composition — on its own module network, publishing no host port, **not** declared as a provision —
when a genuine engine fork or a pinned server version makes the shared provider unusable. Such an
instance is invisible to resolution and can be bound by nothing else. The rule is *share by
default; embed only when a fork or a version forces it* — most of the per-module stores that exist
today are vanilla engines on stale pins that a consolidation onto the node's provider would absorb.
## Consequences
The mesh can carry its real topology **deliberately** rather than by the accident of which
provider happened to be the single one. A consumer's data-coupling becomes a stated fact, which is
what lets a provider be moved without a consumer silently following the wrong one — provided a
provider keeps its identity across a relocation, which is a follow-up this record opens rather than
closes. Rotation ([13](../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md)) addresses a
specific provider's holders, not "the provision's".
What got harder: an assignment now may carry a provider selection, and a wrong one is a new way to
misconfigure. It is mitigated the way 0027 mitigated its own: the co-location default removes the
choice in the common case, and genuine ambiguity — several providers, none named, none co-located —
is refused with the candidates named, never resolved by picking.
## References
- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — the naming this extends;
its "both match" aside is the gap closed here.
- [ADR 0031](0031-the-control-plane-authenticates-nobody.md) — identity is exactly such a
provider, and already serves consumers on another node.
- [ADR 0046](0046-a-module-configuration-is-its-assignments-not-its-manifest.md) — where the
selection lives.
- [ADR 0078](0078-the-store-and-broker-are-modules.md) — the adopted store, whose mesh-wide
binding is the single-provider assumption this record replaces.
- [issue 067](../04-ISSUES/067-a-provision-cannot-name-which-provider-serves-it/00-report.md) — the
gap, and the day-one evidence.
- [`03-DESIGN/01-to-be/23-choosing-a-provider.md`](../03-DESIGN/01-to-be/23-choosing-a-provider.md)
— the design.
@@ -0,0 +1,153 @@
---
topic: what runs on it
status: accepted
date: 2026-09-20
amended: 2026-09-20
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0031-the-control-plane-authenticates-nobody.md
---
# 85. A secret is a provision, and the vault is the module that provides it
## Context
[ADR 0031](0031-the-control-plane-authenticates-nobody.md) decided that identity *runs on the
mesh, not of it* — a module other modules require, rather than a privileged part of the
controller. [ADR 0078](0078-the-store-and-broker-are-modules.md) did the same for the store and
the broker: the twelve-module floor has no specialty left in it. **Secrets are the exception that
survived.** No module owns a secret.
Secret handling is smeared across three built-in parts of the runtime:
- the **controller mints** one credential per consumer↔provider pair and seals it to both node
keys ([ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md));
- the **mesh generates local secrets** — *"generated secrets are the mesh's, never authored"*
(as-is, `06-configuration-and-secrets.md`) — for a value a single module needs for its own use;
- the **synchroniser injects** both into a node's generated files.
Nothing is the owner of "a secret" the way the store module is the owner of "a database", and the
cost is recorded rather than hypothetical. The as-is design names the weakness in its own words:
*"Rotation is not a mesh operation… there is no mechanism that rotates one and informs everything
holding it."* A `rotate` command has since been built and proven, but it reaches **only** the
provisioned pairs; a secret a module generates for its own fully-local use — the password of a
version-pinned embedded store ([ADR 0084](0084-which-provider-serves-a-consumer.md)), an internal
token — is minted by the mesh and then has no operation that can remake it. Three species of
secret, and only the first has an owner:
| species | minted by | rotates? |
|---|---|---|
| a provisioned credential (a database login) | controller, sealed to nodes (0048) | yes — `rotate`, per pair |
| a module's own local secret | the mesh, as a generated value | **no owner, no rotation** |
| an operator-delivered secret (an external key) | a person, sealed in (`secret accept`) | no rotation, no audit |
## Considered Options
1. **Leave it a property of the controller.** Rejected: it is the smear above — no owner, local
secrets that cannot be rotated, operator secrets that cannot be audited — and it is exactly the
specialty 0078 removed for the store and broker, kept here for no reason anyone recorded.
2. **A dedicated vault built into the foundation, not a module.** Rejected: it reintroduces a
privileged built-in, the thing 0031 and 0078 went out of their way to remove, and a mesh that
wants none would still carry it.
3. **Fold all minting, the controller's provisioning credentials included, into the vault.**
Rejected: the controller must mint in order to **deliver** any provision — the vault's own
credential among them — so making the vault mint the credential of its own delivery is the
store/broker chicken-and-egg for no gain. The provisioned-pair credential already has an owner
and a rotation ([13](../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md)); this record
does not disturb it.
4. **A secret is a provision; the vault is an ordinary module that provides it.** Adopted.
## Decision
**Secret-holding is a module, parallel to identity.** A module that needs a secret **for its own
use** — a local service's password, an internal token, an external key it was handed — requires a
`secret` provision from a vault provider, exactly as it requires a database from the store. The
vault generates the value (or holds one it was given), and because the credential belongs to the
consumer↔vault pair it **rotates, backs up and is audited through the same per-pair machinery**
[13](../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md) already defines. That is what
gives the second and third species the rotation and audit they lack. A module's own secret stops
being a generated value that nothing owns and becomes an ordinary provision with a provider.
**The controller's minting of provisioning-pair credentials is unchanged** ([ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)):
it is how every provision, the vault's own included, is delivered. The vault does not mint the
mesh's delivery credentials; it provides secrets to modules, and is itself provisioned the ordinary
way.
**The vault is node-scoped like every provider** ([ADR 0084](0084-which-provider-serves-a-consumer.md)):
each node its own, selected the same way, so a module's own secret is held by the vault on the
module's node. **A mesh that wants no vault runs none** — a module requiring no secret needs
nothing, which is the same test 0031 applied to identity.
## Consequences
The gap that opened this — a generated local secret with no rotation — **closes without new
machinery**: rotating such a secret is the vault's provisioner remaking a pair credential, the
operation 13 already specifies. Backup, audit and break-glass gain an owner — the vault module —
and become things a design specifies rather than absences. The as-is sentence *"rotation is not a
mesh operation"* is already false for provisioned pairs and, once the vault ships, for local
secrets too; the as-is document is updated when it does, not before.
What got harder: a break-glass path — recovering a secret when the sealed delivery path is
unavailable — must not reintroduce a key that one place holds, which is the property
[ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md) was built to preserve; how
the vault offers recovery without it is left to the design as an open question. And a module that
today bakes a password into its own composition must instead require it from the vault — a
migration taken module by module, not a flag day.
## Amendment — 2026-09-20, before anything shipped
Recorded on the record itself rather than as a supersession, by its decider, on the day it was
accepted and before any code was merged against the sentences that change. The original text above
is left as written; this section says what it got wrong and what stands instead.
**What it got wrong.** The decision treated the vault as a provider like the store — optional, and
one per node — and left the mesh's own root secrets outside it: the store's superuser, the broker's
administrator, the controller's contexts, sealed to a node key and nothing else. Those are the
secrets with no rotation and no recovery, and they are the ones a vault exists for. At genesis they
are not even secret: the foundation raises its store and broker with fixed, well-known credentials
and carries those into the mesh. Leaving that floor in place gave module secrets an owner and the
root secrets none.
**What stands instead.**
- **The vault is a foundation module.** It is installed at genesis as part of the foundation
([ADR 0078](0078-the-store-and-broker-are-modules.md) is the precedent: a foundation piece is
still an ordinary module), not assigned later by a mesh that happens to want one. *"A mesh that
wants no vault runs none"* is withdrawn. A mesh has root secrets, so a mesh has a vault.
- **One per mesh, on the control-node.** *"The vault is node-scoped like every provider"* is
withdrawn. Node scoping ([ADR 0084](0084-which-provider-serves-a-consumer.md)) exists because a
store holds data a consumer is coupled to; the vault holds nothing a consumer is coupled to, and a
second one would be a second place to lose. Consumers on other nodes reach it as they reach
identity.
- **The vault holds the mesh's root secrets under an operator-held key.** The controller mints and
delivers exactly as before; in addition, every secret a module holds for itself is sealed a second
time, to an **operator sealing key** whose private half never enters the mesh. The vault keeps
those operator-sealed copies on its own disk, outside the store, and can hand them out — they are
ciphertext to everything but the operator. This is the break-glass path the original text left
open, and it does **not** reintroduce a key one place holds: the mesh holds blobs it cannot open,
and the operator holds a key with nothing to open until given a blob. Recovery needs both.
- **Genesis mints real root secrets and seals them to the operator key first**, so the fixed
credentials the foundation is raised with are replaced before the mesh is handed over.
**Unchanged.** A module's own secret is a `secret` provision the controller mints and the vault
records; the provisioned-pair path of [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)
is untouched; the vault stores no plaintext, ever.
**What it costs.** An operator key is a thing a person must keep, and a mesh whose operator key is
lost has root secrets that can be rotated but not recovered — the same standing as today, stated.
Sealing every own secret twice is a column and a call. Genesis grows a step. A module's
vault-provided secret (a pair credential) is not yet sealed to the operator key; that is the next
increment, not this one.
## References
- [ADR 0031](0031-the-control-plane-authenticates-nobody.md) — identity is a module; this is the
same move for secrets.
- [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md) — the provisioning-credential
path, left unchanged.
- [ADR 0078](0078-the-store-and-broker-are-modules.md) — the de-specialisation this completes.
- [ADR 0084](0084-which-provider-serves-a-consumer.md) — the node-scoping the vault obeys.
- [issue 068](../04-ISSUES/068-secrets-have-no-owning-module/00-report.md) — the gap.
- [`03-DESIGN/01-to-be/24-the-secrets-vault.md`](../03-DESIGN/01-to-be/24-the-secrets-vault.md) —
the design. The source mesh's `secret_locate` / `secret_backup` / `secret_verify` /
`secret_breakglass` subsystem is the prior art it draws on.
@@ -0,0 +1,75 @@
---
topic: building it
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0085-a-secret-is-a-provision.md
---
# 86. A secret reaches a process as a file, and an exception is declared
## Context
The mesh seals a secret to the machine that uses it and discards the plaintext; the host unseals it
into a file at 0600 ([ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)). Then
the module hands it to its container through an env-file, and the runtime puts it where anything
on the machine that can talk to the runtime can read it: `docker inspect` prints it, and
`/proc/<pid>/environ` holds it for the life of the process
([issue 041](../04-ISSUES/041-a-sealed-credential-ends-up-in-the-process-environment/00-report.md)).
Two facts made this a decision rather than a fix. **It is a supported shape**: placing
`${secret:…}` in a file a container reads as its environment is what playbook 06 shows, and 36
containers in 25 catalogue modules do it, the controller's own among them. And **nothing says which
secrets are exposed**: a reader cannot tell from a manifest whether a credential is protected from
`inspect` or not, because the manifest looks the same either way. The vault
([ADR 0085](0085-a-secret-is-a-provision.md)) rests on the seal this undoes.
## Considered Options
1. **Leave it: the environment is where configuration goes.** Rejected — it gives the sealed value
back to a routine operation, and the care taken to seal it is then theatre.
2. **Refuse every secret in an environment.** Rejected — some software reads its configuration
from the environment and nothing else, and a rule the catalogue cannot obey is a rule that gets
switched off.
3. **A secret reaches a process as a file; an environment exception is declared, with a reason,
and refused otherwise.** Adopted.
## Decision
**A secret reaches a process as a file.** A module mounts the file the host wrote and points the
program at it; the mesh's own programs accept a `_FILE` twin for every variable that carries a
credential, the way the controller's store connections already did.
**A container that reads a secret from its environment says so.** The manifest key
`secrets-in-environment` on the container carries the reason. It is catalogue-level — the host
never sees it — and it exists so a reader can tell from the manifest which secrets are exposed
that way and why.
**Everything else is refused.** A `${secret:…}` placeholder inside a container's `env` is never
filled and is refused outright. A file carrying a secret that a container names in `env-file` is
refused unless the container declares the exception.
## Consequences
The property *a sealed credential is readable only where it is used* becomes a manifest-level
fact: true where no exception is declared, stated where one is. The controller reads all six of
its credentials from files; the 35 other containers are marked with their reason and convert one
by one where their software accepts a path, which is per-module work under
[issue 041](../04-ISSUES/041-a-sealed-credential-ends-up-in-the-process-environment/00-report.md).
What got harder: a module author meets one more refusal, and the reason they write is only as
honest as they are. What is not decided: the reach of the exposure on a node — which identities can
talk to the runtime — which decides whether a declared exception is a hardening item or something
sharper.
## How it is checked
The catalogue engine refuses at composition, before anything reaches a machine, and its tests
refuse both shapes and prove the reason is stripped from what is sent.
## References
- [issue 041](../04-ISSUES/041-a-sealed-credential-ends-up-in-the-process-environment/00-report.md)
- [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md), [ADR 0085](0085-a-secret-is-a-provision.md)
- [`03-DESIGN/01-to-be/13-credentials-and-their-rotation.md`](../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md)
@@ -0,0 +1,63 @@
---
topic: what runs on it
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0010-delivery.md
---
# 87. A seeded file is created once, and what grows in it is not the mesh's
## Context
A declaration is complete for what the host owns, and the host reconciles what is declared
([ADR 0010](0010-delivery.md)): a file with this content, held to it. That is the only thing a
manifest could say about a file, and it is the wrong thing for a file a module needs to **exist
before first start** and something else then legitimately writes into — an access list a
provisioner appends consumers to and the program persists back, a bootstrap configuration a
program rewrites. Every reconcile restored the seed behind the running program, erased what had
grown in it, and reported success
([issue 035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md)).
A run-once step ([ADR 0052](0052-a-step-that-runs-once-before-a-container.md)) can write a seed
only if absent, and that closed the instance for a broker whose seed is a program's job. It left
the general case: a plain file the mesh writes and never overwrites.
## Considered Options
1. **Two owners never share a file: the provisioner owns it, and first-start ordering is solved
another way.** Rejected as the only answer — some software refuses to start without the file,
and a module that must ship a program merely to write an empty file has been made to write a
program to say one word.
2. **A create-once semantic on a file.** Adopted.
## Decision
A file resource may say `create-once`. The host writes it when it is absent and, when it is
present, leaves it entirely alone — content, mode and owner — and reports it as **kept**, not
corrected. What is in the file then is somebody else's work the mesh asked for. The mesh removes
nothing it did not create ([ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md)); it now
also does not overwrite what it created once and handed over.
On the security question ADR 0010 asks of every new resource behaviour: this **narrows** what a
declaration can do to a machine. A create-once file gives a compromised control plane one fewer
way to change a machine repeatedly — it can seed, once, and never again.
## Consequences
A module says which of its files are seeds, and the difference is visible in the manifest rather
than in whether the file happened to be revisited. A later change to a seed's declared content
does not reach a machine that already has the file; that is the meaning of a seed, and a module
that needs the new content ships it as a run-once step that migrates the existing file.
## How it is checked
The host's apply tests: a seed is created, grown into by hand, reconciled, and the growth survives
with the outcome `kept`. The vault bed declares one on a real node, grows it, pushes again, and
reads it back.
## References
- [issue 035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md)
- [ADR 0010](0010-delivery.md), [ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md), [ADR 0052](0052-a-step-that-runs-once-before-a-container.md)
@@ -0,0 +1,60 @@
---
topic: the mesh
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0078-the-store-and-broker-are-modules.md
---
# 88. The foundation filters before anything listens
## Context
Adopting the store and broker as ordinary modules
([ADR 0078](0078-the-store-and-broker-are-modules.md)) needs them reachable by consumers across
the mesh, so genesis raises them bound to every interface. The packet filter that decides who may
reach them is a module too, installed a dozen steps later. Between the two, a control-node facing
the network has its store and its bus open to anyone who can reach the machine, for the length
of the install ([issue 054](../04-ISSUES/054-the-adopted-store-and-broker-are-open-before-the-filter/00-report.md)).
The design's rule — what a port is reachable from is decided by the filter — is enforced by
nothing for that window.
## Considered Options
1. **Accept the window**: the machine is mid-bootstrap and the exposure matches what the
pre-adoption modules had in steady state. Rejected — the whole point of deriving the filter
was to stop accepting that.
2. **Bind narrowly at genesis and widen once the filter exists.** Rejected — a bind change is a
recreate of the mesh's store during install, and adoption in place needs the same spec.
3. **The foundation carries a filter of its own, applied before the store.** Adopted.
## Decision
The foundation bundle installs the packet filter and loads a base ruleset **before the store and
broker are raised**: drop by default; keep loopback, replies, ping and ssh; keep the mesh's own
ports a node must reach before it is on the private network — the bus it enrols over and the
registry it pulls from; and let the container runtime's own networks through the forward chain so
containers keep working. It is written into the **same table** the filter module later derives, so
that module replaces it wholesale the moment it can compute one from what the mesh knows, and
nothing of the base survives to contradict it.
## Consequences
From its first resource a machine being made into a mesh refuses what it will refuse when
finished; the window closes. What got harder: the base ruleset is static and names two ports the
mesh's derived one also names — a change to which ports the foundation needs is now made in two
places, and the bundle's own test says which.
## How it is checked
The installer's bundle test asserts the filter and its load precede the store and broker and that
the rules name ssh, the bus and the registry and not the store's or broker's client ports. The
genesis bed probes the machine from outside throughout the install: the store's port is never
reachable, while the bus becomes reachable.
## References
- [issue 054](../04-ISSUES/054-the-adopted-store-and-broker-are-open-before-the-filter/00-report.md), issue 047
- [ADR 0078](0078-the-store-and-broker-are-modules.md)
- [`03-DESIGN/01-to-be/07-the-foundation.md`](../03-DESIGN/01-to-be/07-the-foundation.md)
@@ -0,0 +1,69 @@
---
topic: checking it
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0016-the-lab.md
---
# 89. A bed reads the catalogue it proves
## Context
A lab bed installs a catalogue module and asserts what the mesh does with it; "proven in the
lab" is the standard a module must meet before it ships. The beds built the manifests they
install inline — a literal copied from the catalogue when each bed was written, and never since.
Six modules were converted to file-delivered secrets ([ADR 0086](0086-a-secret-reaches-a-process-as-a-file.md))
and not one bed ran the converted shape; each ran its copy, and the copies still delivered
secrets the way the catalogue engine now refuses
([issue 073](../04-ISSUES/073-beds-carry-copies-of-catalogue-manifests/00-report.md)). A run's
receipt named the commits of the host, the controller and the lab, and said nothing about the
catalogue, so a catalogue change and a proven catalogue change were indistinguishable.
The end-to-end design already has the rule this breaks — *the run rebuilds what it tests; an
artifact rebuilt from memory is one rebuilt sometimes* — for binaries and images. A manifest is
an artifact too.
## Considered Options
1. **Keep the copies and check them against the catalogue** — a test that diffs each literal
against the module's manifest, ignoring what the lab must rewrite. Rejected: it keeps two
sources of truth and adds a third thing that can drift, the list of what to ignore.
2. **Read the catalogue, rewriting only what the lab must.** Adopted.
## Decision
A bed that installs a catalogue module reads that module's manifest from the catalogue checkout
the run was pointed at. It may rewrite what the lab must and nothing else: a build artifact
becomes the image the machine holds, an image is pinned to what the machine holds, a host port
is remapped where one machine carries colliding modules, and an address may point at a stand-in
the bed raises in place of an upstream. Everything else is the catalogue's, verbatim.
A bed that needs less than the catalogue declares — no upstream server, a secret in the
environment, a requirement edge removed — is not testing that module. It is a mesh test, and it
carries a name of its own (see [issue 074](../04-ISSUES/074-a-mesh-test-wears-a-catalogue-modules-name/00-report.md)).
The run's receipt names the catalogue's commit alongside the other repositories', so a receipt
taken before a manifest changed says so.
## Consequences
A catalogue change is proven by the beds that install the module, or it is not proven, and the
receipt says which. What got harder: a bed can no longer trim a module to the shape it finds
convenient; it meets the module's declared requirements or gives its fixture another name.
The beds that still carry a copy are declared, each with its reason, and the declared list
only shrinks.
## How it is checked
A unit test in the lab refuses an inline manifest literal that names a catalogue module unless
the bed is declared, with its reason, in the test's own list; a declaration for a bed that no
longer carries the copy is refused too, so the list cannot outlive the debt. The receipt test
asserts the catalogue is claimed whenever the run is pointed at one.
## References
- [issue 073](../04-ISSUES/073-beds-carry-copies-of-catalogue-manifests/00-report.md), [issue 074](../04-ISSUES/074-a-mesh-test-wears-a-catalogue-modules-name/00-report.md)
- [ADR 0016](0016-the-lab.md), [ADR 0086](0086-a-secret-reaches-a-process-as-a-file.md)
- [`03-DESIGN/01-to-be/01-end-to-end-testing.md`](../03-DESIGN/01-to-be/01-end-to-end-testing.md)
@@ -0,0 +1,65 @@
---
topic: the mesh
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0010-delivery.md
---
# 90. A failure that repeats is said to be stuck
## Context
Delivery is a comparison, not a one-shot ([ADR 0010](0010-delivery.md)): a node re-applies the
declaration it holds on a steady interval and reports each time. That is right for a failure that
goes away by itself — the overlay not up yet, a registry briefly unreachable — and it makes a
failure that will never go away look exactly the same. A resource nothing can ever apply is
attempted, fails, is reported, and is attempted again every few minutes, indefinitely; the mesh
keeps one report per machine, replaced, so each attempt arrives as "failed" at a fresh time.
Nothing distinguished "failed once, will succeed when its dependency arrives" from "failed
identically for ever", and nothing escalated the second
([issue 065](../04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md)).
## Considered Options
1. **The host gives up** after some number of attempts. Rejected: the host does not know whether
a failure is permanent — that a registry has not answered three times is not evidence it never
will — and a host that stops trying is a node that must be pushed to again by hand.
2. **A duration** since the failure was first seen. Rejected as the signal: a laptop shut for a
week has had one attempt, and a week is not evidence of anything.
3. **The controller counts identical reports**, and says when there are enough of them. Adopted.
## Decision
The mesh keeps, beside each machine's last report, when the current failure was first reported
and how many reports in a row have said it — the same outcome, the same refusal, the same failed
resources by id. Not by the host's words: an error carrying a duration or a counter would read as
new on every report, and the resource looping on it is exactly what this is for. A report that
says something different starts the count again; a clean apply clears it. **Three identical reports in a row make a machine stuck**: `status` says
so beside the failure, with the count and the time it began, and the machine-readable status
carries the same three facts. The host keeps retrying; being stuck is a statement about the
mesh's knowledge, not an instruction to the machine.
The controller counts rather than the host, because only it sees every node: one stuck machine
and a mesh-wide fault are different situations, and the host cannot tell them apart.
## Consequences
A resource that will never apply is visible from `status` after three reconcile intervals, to
anyone who looks, without being asked for. What got harder: nothing on the machine changes — a
gating failure that stops what follows still stops it, and this only makes the wait visible.
Whether a stuck machine should also be raised as an event, and what a gating failure should do,
stay open in the issue's own questions.
## How it is checked
An inventory test records the same failure three times and asserts the count and the unchanged
start; the same resource failing in other words, and asserts the count went on; a different
failure, and asserts it restarted; a clean apply, and asserts both cleared. The status command's own test asserts a stuck machine is said to be one.
## References
- [issue 065](../04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md)
- [ADR 0010](0010-delivery.md)
- [`03-DESIGN/01-to-be/10-delivery.md`](../03-DESIGN/01-to-be/10-delivery.md)
@@ -0,0 +1,68 @@
---
topic: what runs on it
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0051-shared-data-is-the-operators.md
---
# 91. A mount is declared, and there are three things it can be
## Context
A container's bind mount whose source does not exist is created by the container runtime, as
root, with whatever mode it picks. So `owner` and `mode` — which exist so a module can say who
its data belongs to — never reach the directories that hold data, and the rule that keeps a
directory when a module goes away ([ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md))
does not cover them, because the mesh has never heard of them
([issue 026](../04-ISSUES/026-the-data-directories-are-not-declared/00-report.md)). Fourteen
such mounts were declared by hand; a check that every mount is declared was then written and
withdrawn, because it refused the builder: the builder mounts the container runtime's socket,
which is not its data, already exists, and belongs to the machine. Declaring it as the module's
own directory would be a lie the host would act on.
## Considered Options
1. **Declare the socket as a directory anyway.** Rejected: the host would create, own and
protect a path that is the machine's.
2. **A new manifest field** naming machine paths a module may mount. Rejected: the manifest
already says the module needs the container runtime, and a second field would say the same
thing in paths.
3. **Three declarations, one for each kind of path a mount can be.** Adopted.
## Decision
A container may not mount a path the module never declared, and a path is declared in one of
three ways, which are the three things a path can be:
- **the module's own** — a directory or file resource, or where a secret, a grant or a
contribution lands. Created and owned by the mesh for this module, kept when the module goes;
- **the operator's** — an `accesses` entry ([ADR 0051](0051-shared-data-is-the-operators.md)):
pre-existing, shared, granted for use, never owned;
- **the machine's** — a facility a declared capability grants. `container-runtime` grants its
socket. The path exists, the machine owns it, and the capability is the declaration.
A mount under a declared directory is declared. The check runs where the manifest is parsed,
and names the path and the three remedies.
## Consequences
Every directory that holds a module's data is one the mesh created with the module's owner and
mode, and one ADR 0030 protects. What got harder: a manifest borrowed from a compose file no
longer passes on the strength of its volume lines; each must say what kind of path it mounts.
The table of what a capability grants is small and in the catalogue's parser; a new capability
that grants a path adds a row.
## How it is checked
Manifest tests refuse an undeclared mount, accept one under a declared directory, accept one
the module accesses, and accept the runtime's socket with the capability and refuse it without.
A test parses every manifest in the catalogue beside the checkout and fails on any that breaks
the rule, so the catalogue cannot drift back.
## References
- [issue 026](../04-ISSUES/026-the-data-directories-are-not-declared/00-report.md)
- [ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md), [ADR 0051](0051-shared-data-is-the-operators.md)
- [`03-DESIGN/01-to-be/18-building-a-module.md`](../03-DESIGN/01-to-be/18-building-a-module.md)
@@ -0,0 +1,61 @@
---
topic: the tiers
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0085-a-secret-is-a-provision.md
---
# 92. An operator delivers a pair credential, and the mesh never replaces it
## Context
[ADR 0085](0085-a-secret-is-a-provision.md) names three species of secret and gives the vault
two of them: a module's own secret, which the mesh mints, and an operator-delivered secret — a
credential for something outside the mesh, which only a person can supply. Under 0085 a secret
from the vault is a pair credential between the consumer and the vault. The controller's one
command that takes a value from a person wrote only a module's own secret, sealed to one node.
Nothing could put a value into a pair, so the third species had no entry
([issue 070](../04-ISSUES/070-an-operator-cannot-deliver-a-pair-credential/00-report.md)), and
an operator's credential could be held only as an own secret — un-audited, un-rotatable, the
gap 0085 opened to close.
## Considered Options
1. **A new verb on the pair.** Rejected: `secret accept` already means "a value a person
supplied, sealed on the way in, plaintext discarded"; a second verb would mean the same.
2. **`secret accept` grows a provider end.** Adopted.
## Decision
`secret accept <node> <module> <name> --provider <node>` seals the supplied value to the
consumer's node, to the provider's node, and to the operator's key when the mesh has one, and
records the pair as `accepted`. Every pair credential now says where it came from: `made` or
`accepted`.
An accepted pair is never replaced by a made one. When a sealing key at either end changes, the
mesh cannot re-seal a value it does not hold, so the read is refused and names the remedy —
accept it again. `rotate` refuses an accepted pair for the same reason: the mesh cannot make its
replacement, and deleting it would have the next read mint one, delivered and reported as
applied while failing to authenticate somewhere else entirely. Rotating an accepted credential
is accepting a new value.
## Consequences
A credential for something outside the mesh lives in the vault's ledger with the others, sealed
to both ends and recoverable by the operator. What got harder: a mesh whose node keys change
cannot heal an accepted pair by itself; a person is asked. That is the honest shape — the value
was never the mesh's to make.
## How it is checked
An inventory test accepts a value into a pair, reads it back twice unchanged with origin
`accepted`, asserts `rotate` refuses it naming the remedy while a made pair still rotates;
another changes a node's key and asserts the read is refused, then accepts again and reads.
## References
- [issue 070](../04-ISSUES/070-an-operator-cannot-deliver-a-pair-credential/00-report.md), [issue 069](../04-ISSUES/069-one-secret-provision-yields-one-value/00-report.md)
- [ADR 0085](0085-a-secret-is-a-provision.md)
- [`03-DESIGN/01-to-be/24-the-secrets-vault.md`](../03-DESIGN/01-to-be/24-the-secrets-vault.md)
@@ -0,0 +1,53 @@
---
topic: checking it
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0089-a-bed-reads-the-catalogue-it-proves.md
---
# 93. A fixture that runs a module's runtime carries the module's name
## Context
[ADR 0089](0089-a-bed-reads-the-catalogue-it-proves.md) said a bed that needs less than a module
declares is a mesh test, and carries a name of its own. Twelve beds were to be renamed on that
basis ([issue 074](../04-ISSUES/074-a-mesh-test-wears-a-catalogue-modules-name/00-report.md)).
Reading how a module's tools are reached showed why they cannot be: the account the mesh issues
a module is scoped to `serve.<module>.*` from the manifest's name, and the runtime binds its tool
queues under the name baked into its image. A fixture named otherwise but running the real
runtime would be refused its own queues. The name is not a label; it is the tool namespace and
the broker scope.
## Considered Options
1. **Rename anyway, and rebuild each runtime under the fixture's name.** Rejected: a runtime built
under a false name proves nothing about the module and costs a build per bed.
2. **A fixture that runs a module's runtime carries the module's name, and therefore reads the
catalogue.** Adopted.
## Decision
A bed that runs a module's runtime installs the catalogue's manifest for that module and what it
requires — the vault for a `secret`, the route module for a route — and proves the mechanism
against the real module. A name of its own is for a fixture that runs no real runtime: a
declaration-only stub, a bare upstream image.
## Consequences
The mechanism beds become module beds with a mechanism inside them, which is more than they
were. What got harder: a bed that wanted a cut-down redis now raises the vault beside it; one that
wanted a sidecar without its server raises the server. Each conversion is a lab run, and the beds
still carrying a copy are declared with this reason until converted.
## How it is checked
The lab's inline-copy check refuses undeclared copies as before; the declared list's reason for
these beds names this record. Two beds converted with the vault beside them ran green.
## References
- [issue 074](../04-ISSUES/074-a-mesh-test-wears-a-catalogue-modules-name/00-report.md)
- [ADR 0089](0089-a-bed-reads-the-catalogue-it-proves.md), [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)
- [`03-DESIGN/01-to-be/01-end-to-end-testing.md`](../03-DESIGN/01-to-be/01-end-to-end-testing.md)
@@ -0,0 +1,63 @@
---
topic: the tiers
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0085-a-secret-is-a-provision.md
---
# 94. A module may hold several secrets from one provider, each a pair of its own
## Context
[ADR 0085](0085-a-secret-is-a-provision.md) makes a module's own secret a provision: the module
requires `secret` from the vault and reads the pair credential minted for that consumer↔vault
pair. A pair has one credential, a module requires a provision once, and so a module received
one value. Read against the catalogue, nine modules hold two or more secrets besides their
broker account; seven of them hold genuinely independent values with independent lifetimes — a
root certificate, its key and that key's password; an admin password beside an API token. None
derives from another, so "one value, derivation the module's business" answers nothing
([issue 069](../04-ISSUES/069-one-secret-provision-yields-one-value/00-report.md)).
## Considered Options
1. **One value per module; the module derives the rest.** Rejected: the values are independent.
2. **Require the provision several times.** Rejected: `requires` is a list of names, and a
requirement is matched by name everywhere.
3. **The `secrets` map names several files under local names, and each local name is a pair
credential of its own.** Adopted.
## Decision
A module's `secrets` entry for a requirement may be a path, as before, or an object of local
names to paths. Each local name is its own pair credential, keyed on it beside the provision,
the consumer node, the consumer module and the provider; its own file on the consumer, referred
to as `${secret:<local name>}`; its own holder at the provider, named the consumer's identity
with the local name after it; and rotated apart from the others. A local name may not be one of
the module's own secrets or something it requires, so what a placeholder means is never
ambiguous. The plain shape is unchanged, and every credential that exists is the one it was.
The holder's suffix is not a login any backend checks — a secret is not a login — so the
identity limit that binds a database role or an access key does not apply to it.
## Consequences
The ten modules that could not move onto the vault can. What got harder: `rotate secret` for a
consumer rotates every local name it holds from that provider; rotating one of several is a
finer command than the mesh has, and waits for a case that needs it.
## How it is checked
Manifest tests read both shapes, write them back, and refuse a colliding or unusable local
name. A resolver test asserts two local names are two needs, two files with two credentials,
and two holders at the provider. An inventory test asserts two local names are two rows, that
rotating one leaves the other, and that the provider is told both. The vault bed installs a
consumer that keeps two secrets and asserts two values delivered, two holders in the vault's
ledger, and both rotated by one command.
## References
- [issue 069](../04-ISSUES/069-one-secret-provision-yields-one-value/00-report.md)
- [ADR 0085](0085-a-secret-is-a-provision.md), [ADR 0049](0049-a-consumers-identity-fits-the-tightest-backend.md)
- [`03-DESIGN/01-to-be/24-the-secrets-vault.md`](../03-DESIGN/01-to-be/24-the-secrets-vault.md)
@@ -0,0 +1,60 @@
---
topic: the tiers
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md
---
# 95. The control plane is the way to ask a module
## Context
A module serves tools over the broker under an account scoped to what it emits, consumes and
serves ([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). A
tool call is a request and a reply: the caller creates a reply queue and publishes to the serving
module's request key, and no module's scope grants either — nor should it, since a module that
only publishes events has no business declaring queues. So a module could serve tools and nothing
in the mesh could call them
([issue 049](../04-ISSUES/049-a-module-can-serve-tools-and-nothing-can-call-them/00-report.md)):
not an operator at a terminal, not an agent acting for one.
## Considered Options
1. **Calling is a grant**: a module declares it may be asked, and a consumer is issued an
account that may create a reply queue and publish to that module's request key. Deferred: a
module-to-module call is the only caller that resembles what the mesh mints today, and none
asks for one yet.
2. **The control plane is the way in.** Adopted. It holds a connection that may already, so a
person or an agent asks through it, and every question passes one process — which is where
an audit of who asked what belongs.
## Decision
`mesh-controller ask <module> <tool> [json]` publishes the request on the tool exchange under
`<module>.<tool>`, with a private reply queue bound under its own name, waits for the answer
whose correlation matches, and prints it as the module gave it. A tool that answered with an
error has answered: the answer is printed and the exit status says so. A module that never
answers is said to have not answered, with where to look.
A module declares nothing about being asked: serving a tool is being askable through the control
plane. A module-to-module call, if one is wanted, is a grant like any other and a later decision.
## Consequences
Anything with the control plane in reach can ask any module anything it serves. What got
harder: nothing outside the control plane can, and the control plane's connection is one more
thing on the path of every question — a cost accepted for the audit it buys.
## How it is checked
A tools-only bed asks a served tool through the control plane and asserts an answer arrived —
an error, since the lab has no upstream and no token, which is an answer where a timeout would
not be.
## References
- [issue 049](../04-ISSUES/049-a-module-can-serve-tools-and-nothing-can-call-them/00-report.md)
- [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)
- [`03-DESIGN/01-to-be/19-the-module-protocol.md`](../03-DESIGN/01-to-be/19-the-module-protocol.md)
@@ -0,0 +1,66 @@
---
topic: building it
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0006-the-substrate-and-the-control-plane.md
---
# 96. An upstream image is copied between registries, never through a machine's image store
## Context
A module may declare that an artifact is an image published elsewhere, to be copied into the
mesh's own registry so machines fetch it by a digest this mesh assigned rather than by a name
somebody else controls. The builder pulled it into the build machine's image store and pushed it
under the mesh's name, and the push was refused: a published image is an index over several
architectures, the runtime's store keeps the index, and pushing one platform out of it fails
however the platform is asked for
([issue 046](../04-ISSUES/046-an-upstream-image-cannot-be-mirrored-into-the-mesh/00-report.md)).
Every variant of pull-then-push was tried and failed the same way.
## Considered Options
1. **Resolve the index to one platform and push that.** Tried, reverted: it did not make the
push work, and a workaround for a store's behaviour is a thing nobody removes later.
2. **Tooling that copies between registries**, installed on the build machine. Rejected: one
more thing the builder's image carries, for a protocol the builder already speaks for blobs.
3. **Copy over the registry API**, in the builder. Adopted.
## Decision
The builder copies an upstream image between registries and never through a machine's image
store: it reads the index and every manifest it names, moves each blob by digest into the mesh's
registry — skipping what is already there, since blobs are content-named — puts the manifests
and then the index under the module's repository, and pins the index's digest. Public images are
read with the anonymous bearer token the registry hands out on challenge, which is how the
public hub and the others the catalogue names serve them. The mesh mirrors the whole index, so
what a machine fetches is the image for its own architecture; that every machine on one mesh is
the same architecture is an assumption this mesh makes and had not written down until now.
Genesis has no registry to copy into and keeps the pull: the image stays in the first machine's
store, named by its own id, as every artifact does before there is anywhere to publish.
## Consequences
An upstream artifact builds. What got harder: the builder now holds a registry client of its
own, some two hundred lines, where a runtime command used to do; and a private upstream that
demands a credential is refused, since the copy is anonymous by design.
## How it is checked
A test raises a fake upstream registry serving an index over two platforms behind a bearer
challenge, and a fake mesh registry that records what arrives: every blob of both platforms
arrives once, two manifests and the index are put under their digests, the reference returned
pins the index under the module's repository, and a second copy uploads nothing. A reference
test reads names the way a runtime does. *Proven against the real thing the same day:* the genesis
bed built the tool runtime through the mesh's builder with its node base copied out of the public
hub into the mesh's registry by this code — after one finding the fake could not give: the builder
ran on the default bridge, where loopback is not the machine, and now runs on the host network.
## References
- [issue 046](../04-ISSUES/046-an-upstream-image-cannot-be-mirrored-into-the-mesh/00-report.md)
- [ADR 0006](0006-the-substrate-and-the-control-plane.md)
- [`03-DESIGN/01-to-be/18-building-a-module.md`](../03-DESIGN/01-to-be/18-building-a-module.md)
@@ -0,0 +1,65 @@
---
topic: building it
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0096-an-upstream-image-is-copied-between-registries.md
---
# 97. A vendor image is a declared build input, and a recipe fetches nothing undeclared
## Context
Three modules could not be built by the mesh's builder because their recipes reached for what
no manifest named: a public package, or a binary copied out of a public image
([issue 064](../04-ISSUES/064-a-mesh-build-cannot-fetch-a-modules-external-dependencies/00-report.md)).
A module already names the bases it stands on — another module's artifact, by name — and the
builder answers with what the mesh holds; a vendor's image had no such declaration, so a recipe
named it directly and the build worked when the public registry answered, which is sometimes.
## Considered Options
1. **Let the build environment reach public registries.** Rejected: a build that fetches from
somebody else's registry on its own is one the mesh cannot rebuild the same way twice.
2. **A vendor image is a base like any other**, declared under `build.on` with the argument the
recipe reads it from, pinned by digest, copied into the mesh's registry before the build
([ADR 0096](0096-an-upstream-image-is-copied-between-registries.md)). Adopted.
## Decision
A build's `on` entry is either a module's artifact or an image published elsewhere, pinned by
digest, read from one build argument. Before the build the image is copied into the mesh's
registry under the module's repository and the recipe is handed the copy; genesis, with no
registry, pulls it into the first machine's store. A recipe whose `COPY --from` names a registry
image the manifest did not declare is refused before the build, naming the image and the remedy;
its own stages, declared arguments and `scratch` are not fetches. An unpinned vendor image is
refused: a tag is what somebody else can move.
A recipe whose `FROM` names an undeclared base was at first said, not refused: the mesh's own
images — the control plane's, the tool runtime's, the route proxy's — started from a public base
and declared none, and refusing those refuses genesis. *Amended the same day:* those three declare
their bases now, and an undeclared `FROM` is refused like an undeclared copy. The builder's own
image and the examples are built by `make`, not by the mesh, and take arguments with defaults.
The package half of the issue is not decided here: the mesh's package registry already proxies
the public one, and the failure the report saw has to be run again to be placed.
## Consequences
A module's build inputs are all in its manifest, and every one of them is something the mesh
holds a copy of. What got harder: a recipe that used to name a base image on its first line now
names an argument, and the manifest names the image.
## How it is checked
A builder test declares a pinned vendor image, asserts it is copied under the module's
repository and handed to the recipe as the argument, and asserts an unpinned one is refused. A
recipe test asserts an undeclared `FROM`, an undeclared `COPY --from` and an undeclared argument
are named, and that stages, declared arguments and `scratch` are not.
## References
- [issue 064](../04-ISSUES/064-a-mesh-build-cannot-fetch-a-modules-external-dependencies/00-report.md)
- [ADR 0096](0096-an-upstream-image-is-copied-between-registries.md)
- [`03-DESIGN/01-to-be/18-building-a-module.md`](../03-DESIGN/01-to-be/18-building-a-module.md)
@@ -0,0 +1,61 @@
---
topic: the tiers
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0085-a-secret-is-a-provision.md
---
# 98. A fact a provider makes at first start is fetched from it, not carried in its manifest
## Context
The catalogue's certificate authority declared its root certificate, its root key and that key's
password as its own secrets, and told the container to initialise from them. The mesh mints an
own secret nobody delivers as random bytes, and random bytes are not a certificate: issued, the
authority could not start; only an operator hand-making its root could raise it
([issue 076](../04-ISSUES/076-a-served-fact-made-at-first-start-cannot-be-served/00-report.md)).
The authority can make its own root at first start. What it could not do then was tell the mesh
what that root is: a consumer was given `${bound:acme-ca:root}` from the provider's `serves`,
which is written in the manifest before anything runs.
## Considered Options
1. **A secret the module makes**, with the mesh taking custody once the file exists. Rejected
for now: a node would have to send a value up to the mesh, which no channel does today, and
a root key is the one thing the mesh has no reason to hold.
2. **A served fact the provider contributes at run time.** Rejected for now: the same new
channel, for a fact that is not secret at all.
3. **The consumer fetches it from the provider**, over the mesh network, through a gate before
the thing that needs it starts. Adopted.
## Decision
A provider's `serves` names where a fact made at first start can be fetched — the authority
serves its root at a path beside its ACME directory — and a consumer fetches it in a `run-once`
step declared before the resource that needs it, from the provider's bound address. The mesh
network is where the fetch happens, which is what makes fetching without a prior trust
acceptable: it is the network the mesh itself authenticates. The mesh mints only what it can
make: the authority's password. The root key stays where it was made.
## Consequences
The catalogue's authority starts, and the proxy that requires it trusts what it fetched. What
got harder: a consumer of such a fact carries one more resource, the gate that fetches it, and
a fact that changes after first start is refetched only when the declaration changes.
## How it is checked
The route-forwarding bed installs the authority, the proxy and a consumer from the catalogue and
asserts a routed name is served through the proxy. The proxy refuses to start on a bundle that is
not a certificate, so the name being served proves the gate fetched one; the gate itself refuses
a body that is not a certificate. That the proxy obtains a certificate from this authority through
that root is the certificate bed's proof, against the same authority with the same proxy. The
catalogue-wide manifest test parses both manifests.
## References
- [issue 076](../04-ISSUES/076-a-served-fact-made-at-first-start-cannot-be-served/00-report.md)
- [ADR 0053](0053-a-step-that-runs-on-a-schedule.md), [ADR 0085](0085-a-secret-is-a-provision.md)
- [`03-DESIGN/01-to-be/08-connectivity.md`](../03-DESIGN/01-to-be/08-connectivity.md)
@@ -0,0 +1,94 @@
---
topic: what runs on it
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 0052-a-step-that-runs-once-before-a-container.md
---
# 99. A step that runs once names what it reads, and runs again when it changed
## Context
[ADR 0098](0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md) has a consumer
fetch a fact its provider made at first start through a run-once step: the route proxy fetches
the certificate authority's root before it starts. A run-once step runs once per declaration
([ADR 0052](0052-a-step-that-runs-once-before-a-container.md)): its marker is the digest of its
own declaration, and a re-apply that finds the marker does nothing.
The provider can move. When the authority is assigned to another node it makes a new root there,
and the mesh rewrites the consumer's binding file with the new address — but the step's own
declaration has not changed, so the step does not run again, the proxy keeps the old root, and it
refuses every certificate the new authority issues
([issue 077](../04-ISSUES/077-a-fact-fetched-at-first-start-is-fetched-once/00-report.md)). A
restart trigger was the natural remedy and was refused on a run-once step, on the ground that a
step does not stay running to be restarted.
Two things the host already does point at the answer. What a container reads is part of what it
is: a container's digest includes the digest of every resource it names under `restart-on`, so a
rewritten file it reads is a changed container
([issue 045](../04-ISSUES/045-a-container-keeps-the-values-it-started-with/00-report.md)).
And a run-once step's marker *is* its digest. Nothing new is needed for the step to run again when
what it reads changed; only the refusal stands in the way.
## Decision
**A run-once step may name what it reads under `restart-on`. For a step the word means *run
again*: when a named resource changed in this apply, the step's digest has moved, its marker no
longer matches, and it runs again — gating what follows, as it did the first time.** Nothing
about the marker changes; the refusal of the pair is lifted, in the control plane and on the host.
**The container that consumes what a step made names the step.** A step that ran counts as a
change, so a service that names it under `restart-on` is recreated after it, holding what the step
fetched. Without this the step fetches a new root and the service keeps serving with the old one.
The route proxy's gate names the binding file it reads; the proxy's server names the gate and the
binding. When the authority moves, the binding is rewritten, the gate fetches the new root, and the
server is recreated with it — in one apply.
## Considered Options
1. **A provider epoch in the binding — the mesh raises a number when a provider is re-issued or
moved, and the consumer's file carries it.** Rejected: the binding already changes when the
provider moves (its address does), and a re-issue does not change what the authority serves —
its state persists. An epoch would be a second signal for a change the file already shows.
2. **The step runs before every start of the service, with no marker.** Rejected: every reconcile
would run it, and a step that runs on every apply reads as a change on every apply, so the
service naming it would be recreated every few minutes.
3. **The proxy fetches the root itself, at start.** Rejected as the general answer: it fixes the
proxy and leaves the next consumer of a fact made at first start to fix itself. The step is the
general shape ([ADR 0098](0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md)).
4. **Lift the refusal and read `restart-on` as *again* on a step.** Adopted: it is what the
digest already does, and it needs no new word.
## Consequences
A fact fetched at first start follows its provider when the provider moves. What is not covered:
a provider whose state is wiped behind the mesh's back, on the same node, makes a new fact that
nothing the mesh knows reflects. That is not a change the mesh can see, and it is not claimed.
The one contradiction the refusal named is real and is now a documented reading: on a running
container `restart-on` means recreate, on a step it means run again. Both are "this must reflect
what it reads". This record is about a run-once *container*, the step ADR 0052 defined. A
run-once process is a different shape whose marker does not carry what it reads; it is not
covered here.
## How it is checked
- mesh-host: a unit test declares a run-once step naming a file, records its marker against the
file's old content, applies with the new content and asserts the step ran; applies again with
nothing changed and asserts it did not. A second test declares a container naming a run-once
step and asserts the container is recreated after the step ran, with the step as the stated
reason.
- mesh-controller: the manifest parser accepts a run-once step with `restart-on`; the
catalogue-wide manifest test parses the route proxy's manifest, whose gate and server name what
they read.
- The route-forwarding bed still passes with the host that accepts the pair. No bed moves the
authority: the mechanism is proven by the unit tests, the declaration by the manifest test.
## References
- [issue 077](../04-ISSUES/077-a-fact-fetched-at-first-start-is-fetched-once/00-report.md)
- [ADR 0052](0052-a-step-that-runs-once-before-a-container.md), [ADR 0053](0053-a-step-that-runs-on-a-schedule.md), [ADR 0098](0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md)
- [`03-DESIGN/01-to-be/08-connectivity.md`](../03-DESIGN/01-to-be/08-connectivity.md), [`03-DESIGN/01-to-be/20-writing-a-module.md`](../03-DESIGN/01-to-be/20-writing-a-module.md)
@@ -0,0 +1,221 @@
---
topic: the mesh
status: accepted
date: 2026-09-22
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0078-the-store-and-broker-are-modules.md
---
# 100. A node in use is adopted before it is converged
## Context
The mesh replaces a predecessor mesh that is running, on the same machines, with the services
people use. The control-node is the machine that already carries the predecessor's broker and
build pipeline. The operator proposed the migration's shape: stop the predecessor's control on a
machine, bring the mesh up there in an adoption mode that keeps the machine's configuration in
force, migrate its modules one at a time, move to the next machine, and flip adoption off when every
machine is done.
Measured on the control-node ([research 012,
*migrating a node that is in use*](../01-RESEARCH/012-the-minimum-viable-node/migrating-a-node-in-use.md)):
60 predecessor containers; the predecessor's control in user-level daemons separate from its
services; its firewall active, with 54 incoming and 52 forwarding rules allowing each served port
explicitly; 12 files its configuration sync writes, 3 of them system files.
Four things in the mesh as it stands break that shape:
1. **The base ruleset closes the machine.** Genesis loads a drop-by-default table
([ADR 0088](0088-the-foundation-filters-before-anything-listens.md)). Every base chain at a hook
runs in priority order; an accept ends only its own chain and a drop in any is final — whether
the other firewall's chains are nftables or legacy iptables. So the predecessor's allowed ports
would close at genesis.
2. **The host replaces what it finds.** A declared file is written whatever is at its path, except
a file declared create-once. A declared container replaces a running one of the same name. So
assigning a module the predecessor also runs replaces the predecessor's service and its files at
once.
3. **Some foundation ports are held.** The foundation's store and bus ports are free — the
predecessor publishes its own elsewhere — but the registry's port and the broker's management
port are held, and on a control-node that is the private network's hub, so is the private
network's port. The foundation's ports are fixed in the installer's bundle and the catalogue's
manifests, so a collision surfaces as a container that fails to bind, and a port changed at
genesis would be changed back when the foundation is adopted as modules.
4. **Published container ports are forwarded, not received.** The foundation publishes its ports
on every interface; a predecessor firewall that filters only incoming traffic never sees them.
The base ruleset is what keeps the store unreachable from outside, and it is the thing that
cannot be loaded.
The mesh already adopts in one place: the foundation's store and broker are taken over in place as
modules, keyed on the container that is already running
([ADR 0078](0078-the-store-and-broker-are-modules.md)). The research this record draws on already
settled the conflict rule for adoption: *on conflict, what is on the machine stays*.
## Considered Options
1. **Cut each machine over in one go** — stop the predecessor's store, broker, registry and proxy,
raise the foundation in their place. Rejected: every predecessor service on the machine is down
until it has migrated, the predecessor's other machines lose their broker, and the rollback is
restarting the predecessor — a recovery, not a step.
2. **Put the control-node on a separate machine.** Rejected: the control-node is decided.
3. **Converge on joining, as today.** Rejected: the base ruleset closes the machine at genesis, and
the predecessor's services and files are replaced as soon as any module naming them is assigned.
4. **Only make the foundation's ports configurable.** Rejected as insufficient: it answers the bind
collisions and neither the firewall nor the files.
5. **Treat assigning a module as migrating it.** Considered and rejected on review: it makes the
rule that keeps found files never fire — the host only ever sees files of assigned modules — and
it makes an assignment on an adopted node an outage rather than a preparation.
6. **Adoption as a mode per node; each module taken explicitly; the node converged by an explicit
flip.** Adopted.
## Decision
**A node is adopted or converged, and the controller records which.** The operator says so: at
genesis for the control-node, and in the enrolment token for the others. The controller is
authoritative, and every declaration it sends says whether the node is adopted and which modules
have been taken on it. A node stays adopted until the operator converges it. An adopted node is
said to be adopted wherever the mesh reports a node's state.
**A converged genesis refuses a machine in use.** A machine is in use when a container is running
on it, or a port is listening on an address other than loopback that is not ssh's. Raised without
saying adopted on such a machine, genesis refuses and names every container and listener it
counted — a forgotten flag must not close a working machine.
**Before a node is adopted, its predecessor's control is stopped by the operator** — the daemons
that write its configuration. Its services keep running on what they have.
**Found means present with no record.** A file at a declared path, or a container at a declared
name, that the host's store has no record of writing is *found*. A file the host wrote in an
earlier life of the node is not found; its record says so.
**On an adopted node, what is found is kept until its module is taken.** The host keeps a found
file and a found container as they are, records the file's original content before anything else
happens to it, and reports each as held. Assigning a module on an adopted node prepares it: what
the module declares that is not found is created; what is found is held. **Taking a module** on a
node is its cutover — the operator's act, done when that module's data has moved — and from then
on the module's resources converge on that node like any other. What is held is never removed,
even when its module is unassigned, and a held file or container that changes while held — a file
rewritten, a container stopped or replaced — is reported as changed by something else, not reverted
or restarted: that is how a predecessor still writing is caught.
**The firewall found on the machine stays in force.** The mesh loads no table on an adopted node
that drops by default or holds an accept — neither genesis's base ruleset nor the filter module's
derived one. What the mesh needs reachable is declared as **openings**: a resource that says a port
is reachable, from where, on the incoming path or the forwarded path — a published container port is
forwarded. The controller derives them from the same inputs as the filter, each from where the
filter would admit it: the `listens` of the modules assigned there, the private network's hub port
and the bus and the registry from anywhere, the store's port and the broker's management port from
the private network. The host
converges an opening through the found firewall in that firewall's own terms, marks it as the
mesh's, and removes only what it marked; it re-checks each opening on every reconcile, so a reload
or a reboot of the found firewall does not lose it for longer than one reconcile. An opening is a
state, not a command, which is what lets it travel over the link. The host reports which firewall
it found. A machine with no firewall needs no openings; a machine with a kind no host speaks is
refused adoption.
**The mesh guards its own ports itself, in a table of its own that only refuses.** It passes
everything by default and holds nothing but refusals, and the two ports it refuses are the
foundation's own, checked free at genesis, so it cannot close anything the machine serves; it is
the mesh's, so the found firewall reloading does not touch it. It refuses the store's port and the
broker's management port except from the private network and from the machine itself — its
loopback and the container runtime's own networks, known by the interface a packet arrives on and
never by its source address alone — at the prerouting hook, ahead of the runtime's
destination translation, so it matches the port the packet was sent to, for both address
families. The bus and the registry stay reachable from anywhere, as a node
enrols over the bus and pulls from the registry before it has a private-network address
([ADR 0088](0088-the-foundation-filters-before-anything-listens.md)); so does the private network's
hub port. The store is unreachable from outside whatever the found firewall does, and on a machine
with none.
**The foundation's ports are the node's.** Every port the foundation binds is an input to genesis,
checked free before anything is raised, refused with the name of what holds it. The ports given
become that node's settings for the foundation's modules — the catalogue's numbers are only their
defaults — and every place that uses them reads them from there: the modules' containers, the
filter, the base ruleset, the private network's endpoint, and the addresses consumers are given.
The private network's address range must not overlap a tunnel the predecessor still runs; genesis
checks that too.
**Converging a node is one act, previewed.** It refuses while an assigned module still holds a
found container: each service is taken on its own, when its data has moved, never by the flip. The
preview lists what is reachable on the machine now — every listening socket and every published
container port — and for each whether an assigned module declares it or it will close, and every
module the flip will take, with the held files each will replace. The flip then takes those modules, loads the
mesh's derived filter in place of its refusal-only table, and retires the found firewall by
disabling it, never by flushing: the container runtime's rules and the found firewall's own
configuration stay on disk. Returning a converged node to adopted unloads the derived filter,
restores the refusal-only table, enables the found firewall again and converges the openings
through it once more; what was taken stays taken. A
node converges when its migration is done; the mesh is migrated when every node has converged.
**The order is the operator's:** the control-node first, adopted, its modules assigned and taken
one at a time; then each other machine, adopted, migrated, converged in turn. The predecessor's
pipeline runs on the control-node, so its updates stop for every machine while the migration runs;
that is accepted.
## Consequences
Each step says what it changes before it changes it. Adopting a node changes nothing that serves;
assigning a module adds what is not there; taking a module replaces one service; the flip replaces
the firewall, after naming every port it will close. Two steps are not undone by the mesh: taking a
module replaces the predecessor's container, and the kept original of a file is recorded but not
yet restored by any act of the mesh
([research 012](../01-RESEARCH/012-the-minimum-viable-node/00-overview.md) leaves where it lives
open).
What got harder:
- **The mesh must speak a firewall it did not install**, on both the incoming and the forwarded
path. One kind is found on the machines measured; another is refused until a host speaks it. The
mesh's own guard does not depend on it: that table is the mesh's.
- **The host gains a guard it did not have** — keep what you found — and its report must say
which files and containers it holds, or an adopted node reads as converged.
- **The declaration gains a node's mode and its taken modules**, and a resource, the opening.
- **Genesis grows inputs, and they outlive genesis.** The foundation's ports stop being constants;
every reader of them reads the node's settings.
- **A node can sit adopted indefinitely.** Nothing forces the flip; the mesh's status says which
nodes are adopted, so one left behind is visible.
- **This narrows [ADR 0088](0088-the-foundation-filters-before-anything-listens.md) for adopted
nodes**: the base ruleset is not loaded on a node raised adopted, and its duty — the store never
reachable from outside — passes to a table of the mesh's that only refuses.
## How it is checked
A lab bed prepares a machine the way the predecessor leaves one: its firewall allowing a served
port and denying the rest, a service container listening on that port under a name a catalogue
module also uses, a file at a path that module declares, a stand-in for the predecessor's control
that would rewrite that file, and a container holding the registry's port. Then:
- **Genesis converged** on it refuses and names every container and listener it counted.
- **Genesis adopted, with the registry's port held**, refuses and names the holder; with another
port given, the foundation comes up — and adopting the foundation as modules leaves it on that
port.
- **Nothing that serves changed**: the service is reachable from a second machine, the file is byte
for byte what it was, and the found firewall's rules differ only by rules marked as the mesh's.
- **The store is unreachable from outside** — probed from a machine off the private network, and
again after the found firewall is reloaded — and reachable over it and from a container on the
node itself; the bus is reachable from a machine that has not yet enrolled.
- **The mesh works through the found firewall, and keeps working after it is reloaded and after the
machine reboots**: the second machine enrols, and the openings are there again.
- **A predecessor still writing is caught**: with the stand-in left running, the held file's change
is reported and not reverted.
- **Assigning prepares, taking cuts over**: the module assigned holds the found container and file;
taken, it replaces them and its port is opened.
- **Converging previews, then changes**: it refuses while the service's module holds its found
container; once that module is taken, the preview names the service's port and a published port no
firewall rule mentions, and the modules it will take; after the flip the mesh's derived filter is
loaded, the found firewall is disabled with its configuration still on disk, the declared port is
open and the undeclared one closed. Returned to adopted, the found firewall is enabled again and
the derived filter is gone.
Unit tests hold the host to keeping a found file and container on an adopted node, converging them
once taken, never removing what it holds, and reporting a held file or container that changed;
genesis to refusing a held port and a converged raise on a machine in use; the controller to
carrying the mode and the taken modules in every declaration, deriving the openings, and refusing
a flip while a found container is held.
## References
- [research 012 — the minimum viable node, and adopting what is already there](../01-RESEARCH/012-the-minimum-viable-node/00-overview.md),
and its document [*migrating a node that is in use*](../01-RESEARCH/012-the-minimum-viable-node/migrating-a-node-in-use.md)
- [ADR 0005](0005-the-node-host.md), [ADR 0011](0011-managed-files-are-generated-never-edited.md),
[ADR 0078](0078-the-store-and-broker-are-modules.md), [ADR 0088](0088-the-foundation-filters-before-anything-listens.md)
@@ -0,0 +1,70 @@
---
topic: the mesh
status: accepted
date: 2026-09-22
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
---
# 101. A machine's own resolver does not make it in use
## Context
[ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md) has a converged genesis
refuse a machine in use. It defines *in use* as a container running, or a port listening on an
address other than loopback that is not ssh's. That definition was written before anything was
measured.
Measured on a freshly installed lab machine, running nothing but its operating system:
| Listening | Held by |
|---|---|
| TCP and UDP on every address, the link-local name resolution port | the system's name resolver |
| UDP on every address, the multicast name resolution port | the system's name resolver |
| UDP on a link-local address, the address-configuration client port | the system's network manager |
| TCP and UDP on loopback | the resolver's stub and the container runtime |
By ADR 0100's words, the resolver's TCP listener on every address makes **every** freshly
installed machine a machine in use. A converged genesis would refuse them all, and `--adopted`
would become the only way to raise anything. The refusal exists to catch a forgotten flag on a
working machine. Refusing an empty one defeats it, and teaches operators to pass the flag by
habit.
## Considered Options
1. **Keep the words.** Rejected: every fresh machine is refused.
2. **Count TCP only, ignore UDP.** Rejected: the resolver listens on TCP too, and a machine that
serves over UDP alone, a resolver or a tunnel, is in use.
3. **Ignore listeners held by the operating system's own network daemons**, a short named list,
on both protocols. Adopted.
## Decision
**A listener held by one of the operating system's own network daemons does not make a machine
in use.** The daemons are the ones the measurement found: the name resolver and the network
manager, named in the installer's code beside that measurement. Everything else in ADR 0100's definition stands: a running container, or any other
listener on an address other than loopback that is not ssh's, makes the machine in use, and
genesis still names every one it counted.
A daemon is added to the list only with a measurement of a fresh machine that holds it.
## Consequences
- A converged genesis on a fresh machine goes ahead, as it did before ADR 0100.
- A machine whose resolver is also serving other machines is not counted as in use by its
resolver alone. It is one of the daemons that serves nobody on a fresh machine, and the one
kind of service this lets through.
- The list is code, not configuration, so it changes by review.
## How it is checked
A unit test in the installer feeds the listeners captured from the fresh machine, as the
listening-socket tool printed them, and asserts the machine is not in use. The existing tests
still assert that a serving machine is in use and that every container and listener is named.
The adoption lab bed raises a converged genesis on a fresh machine and asserts it is not refused.
## References
- [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md)
- [research 012, *migrating a node that is in use*](../01-RESEARCH/012-the-minimum-viable-node/migrating-a-node-in-use.md)
@@ -0,0 +1,93 @@
---
topic: the mesh
status: accepted
date: 2026-09-22
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
---
# 102. The mesh writes into a shared file, never over it
## Context
[Issue 084](../04-ISSUES/084-taking-networking-on-an-adopted-node-restarts-every-container/00-report.md)
found that the networking module declares the container runtime's configuration file to state
one fact in it: the mesh's registry is trusted over the private network. Reading the code
closer showed it worse than reported. The controller merges an operator's settings into the
module's own content, but the host writes the result **whole**. Whatever the machine had in that
file is replaced, including where the runtime keeps its data. On a machine in use, that is every
image and container gone from the runtime's view at its next start. The runtime's service is then
restarted, which stops every container on the machine.
Measured on a lab machine: the runtime takes a new list of trusted registries on a reload, with
no restart, and a running container with no restart policy keeps running through it. The
runtime's log says it reloaded its configuration, and the registry reads as trusted afterwards.
Under [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md), the file is found
on an adopted node and held until networking is taken. Holding it is safe, but it means an
adopted machine cannot pull the mesh's images until then, and taking networking would restart
the runtime.
## Considered Options
1. **Keep writing the whole file; take networking as a cutover of its own.** Rejected: it
still replaces the machine's settings, and the cutover stops everything on the machine.
2. **A drop-in the runtime reads beside its main file.** Rejected: the runtime has no such
directory for its daemon settings.
3. **Write into the file: set the mesh's keys, keep the rest; reload, don't restart.** Adopted.
## Decision
**A file the mesh shares with software it did not install is written into, never over.** A
file resource may say it is written *into* a structured file. The host then reads what is there,
sets only the keys the mesh declares, keeps every other key as it found it, and records what each
of its keys held before. **A list is added to, never replaced**: where the machine already has a
list under a key the mesh declares, the mesh's members are added to it and the host records
exactly which members it added — setting the key would replace the operator's own list, the harm
this record exists to prevent. Undeclared later, each key goes back to what it held and each
added member is taken out again, and a file the mesh created is removed only if nothing but its
own keys is left. A file written into replaces nothing, so on an adopted node it is never held:
it is written whether or not its module has been taken.
**Whatever the host writes over without a record of it, it keeps first.** On any node, adopted
or converged, before the host writes a file over one it has no record of making, it keeps the
original once and says where; if it cannot keep it, it does not write. A file the mesh takes over
is then never lost, whatever put it there.
**A service that re-reads its configuration on a reload is reloaded, not restarted.** A service
resource may name what it must be *reloaded* on, beside what it must be restarted on. The
container runtime is reloaded for the registry's trust.
**The networking module writes the runtime's trust into its file and reloads it.** An adopted
node therefore trusts the mesh's registry as soon as it is on the private network, and taking
networking no longer touches the runtime. The hosts file networking writes is still written whole
and stays held until networking is taken; a converge preview names it among the files it replaces.
## Consequences
- The runtime's file on a machine in use keeps its data directory, its logging settings and
everything else the predecessor set.
- A host that does not know *into* or *reload-on* refuses a declaration carrying them, so hosts
are upgraded before the controller that emits them — the same order ADR 0100 needs.
- One shape more for every host: a file written into a structured document. Only JSON is spoken;
another format is refused until written.
- The hosts file remains a whole file. Writing a marked block into it is the same idea for a text
file and is not decided here.
## How it is checked
Unit tests hold the host to setting only the declared keys and keeping the rest, restoring each
key and removing only a file it created when the resource is undeclared, adding to a list and
removing only the members it added, keeping the original of a file it writes over without a
record, refusing a file that is not a JSON object rather than overwriting it, never holding a file written into on an adopted
node, and reloading rather than restarting a service whose reload-on resource changed. A test in
the controller holds the networking module to declaring the runtime's file written into and the
runtime reloaded. The adoption lab bed gives the machine a runtime file with a setting of its own
and asserts it survives adoption with the registry trusted added, and that a container without a
restart policy is still running afterwards.
## References
- [Issue 084](../04-ISSUES/084-taking-networking-on-an-adopted-node-restarts-every-container/00-report.md)
- [ADR 0082](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md), [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md)

Some files were not shown because too many files have changed in this diff Show More