Author SHA1 Message Date
jschoubben a97feeefe5 Renumber to 114: 113 collided with an issue merged independently to main
Both used the next number available when opened, and the object-store
withdrawal report merged first. No content change beyond the number.
2026-09-24 15:48:56 +02:00
jschoubben a18550e460 Issue 113: should the controller be a container or a process the host supervises
Filed after a session where every mesh-controller interaction went through
docker exec — its manifest runs it as a container with network: host, using
none of the isolation that resource type usually buys, while ADR 0006 makes
it the mesh's single point of coordination. Open question, not a claimed
defect: does type: container get the controller anything type: process
(supervised the way the host supervises its own unit, per ADR 0005) would not.
2026-09-24 15:48:22 +02:00
jschoubben f10dce4f9e Merge pull request 'Issue 113 and research 015: the object store's images are gone upstream, not access-restricted' (#103) from storage/113-the-object-store-lost-its-upstream into main 2026-09-24 13:45:34 +00:00
jochen 4beb6629db Issue 113: ground the rebuildability point in what the design actually says
Review of my own text found an unattributed claim — "the mesh's claim that a node
can be rebuilt from its declarations" — which is not a stated principle anywhere.
Replaced with the design position that genuinely covers it: to-be 07 chooses
references over payload because "reproducibility comes from pinning the identity of
a thing rather than carrying its bytes". This incident is that choice's failure mode
when the identity stops resolving, which is a sharper point than the one I made.

Scope stated honestly: the passage is about the foundation bundle and this module is
not in it, but pin-identity-fetch-bytes is how every module gets third-party images.

Also names the tension the mirroring question actually carries — mirroring is a move
away from references-over-payload, so it is a decision, not a fix.
2026-09-24 15:44:44 +02:00
jochen d497b37e43 Issue 113 and research 015: the object store's images are gone upstream, not access-restricted
The symptom arrived diagnosed as "the registry disabled anonymous pulls for the
whole vendor namespace". It did not hold: sibling repositories in that namespace
pull normally, the "$disabled" token field appears on every repository including
working ones and describes signing rather than access, and "actions": [] with a
401 is byte-identical to what an invented repository name returns. Both registries'
own APIs establish deletion instead.

Recorded because the correction is the expensive part to rediscover, and because
the instance was harmless while the standing condition is not: no node that does
not already hold the images can ever provision the module again, and nothing
detects that until one tries.

Research 015 scopes the replacement. It is not a redesign — the foundation design
already commits to S3 the protocol rather than the product, and the object store
is an ordinary module, so this instantiates an existing principle. The live OIDC
wiring is the requirement that gates the choice, and it is checked first.
2026-09-24 15:27:03 +02:00
jschoubben 183b22997c Merge pull request 'Issue 112: diagnose — the predecessor's own DNS config already names the carried peers' (#101) from issue/112-diagnosis into main 2026-09-24 12:49:50 +00:00
jschoubben 8d67cf63c5 Issue 112: status located, not diagnosing
Playbook 03 step 2: move status to diagnosing, then located once the
owner is known. located-in is filled with four confirmed packages —
the owner is known.
2026-09-24 14:27:25 +02:00
jschoubben 75c104c355 Issue 112 diagnosis: correct located-in attribution
The carried-peer record (CarriedPeer/TunnelPeer) lives in mesh-controller
internal/inventory, not internal/catalogue. internal/catalogue is the
right package for the zone-generation side of the fix (facts.go's
nodeZones), but a different concern from where the name field itself
would go. Split the two so a decision record doesn't get pointed at the
wrong package.
2026-09-24 13:54:20 +02:00
jschoubben 6a56738d7b Issue 112: diagnose — the predecessor's own DNS config already names the carried peers
Checked why ADR 0104's forward-to-predecessor shape doesn't transfer to the
resolver the way it did the proxy: DNS is one process on one port, and
assigning the mesh's dnsmasq module replaces it in place, so there is no
predecessor process left standing to forward to.

But /etc/dnsmasq.d/hal-dns.conf's static address= lines for ace/shanks/g14
match the mesh's own carried-peer addresses from overlay show exactly. The
name a carried peer needs isn't a guess the operator has to make under
pressure — it's a transcription of a record the predecessor already has and
has been correctly serving for six days. Located in mesh-controller (no way
to attach a name to a carried peer today) and the dnsmasq module (doesn't
emit a wildcard for a named-but-uncarried peer). Not implemented.
2026-09-24 13:25:44 +02:00
jschoubben 91a5c63d65 Merge pull request 'Name the migration repository in the map, so nobody has to be told it exists' (#99) from meta/name-the-migration-repository into main 2026-09-24 00:02:07 +00:00
jschoubben 545d038198 Name the migration repository in the map, so nobody has to be told it exists
hq cannot hold the migration's operational record — it names machines, addresses and
paths, and this repository is public — but it can say where that record is, which is what
this map is for. Asked for by the operator, who had to be told.
2026-09-24 02:01:41 +02:00
jschoubben 7f438d049f Merge pull request 'Issue 092: genesis publishes to a registry the container runtime does not yet trust' (#79) from issue/092-genesis-registry-trust into main 2026-09-23 23:38:46 +00:00
jschoubben 60e43f9446 Merge pull request 'Issue 091: a module definition carries a machine port' (#78) from issue/091-machine-ports-in-manifests into main 2026-09-23 23:38:40 +00:00
jschoubben 36aa722c2a Merge pull request 'Issues 111 and 112: the resolver was told the wrong set of names, twice over' (#98) from issues/111-112-the-resolver-was-told-the-wrong-names into main 2026-09-23 23:32:42 +00:00
jschoubben f704e2ca64 Issues 111 and 112: the resolver was told the wrong set of names, twice over
111, resolved: the map the control plane hands a resolution holds the machines and the
names the mesh merely serves, and the resolver's zones were given both — inventing names
under a suffix it answers authoritatively for. 112, open: adopting a tunnel gives the mesh
the peers' addresses and none of their names, so taking the resolver before they enrol
stops three machines resolving at all.

Both found by reading the plan before pushing it.
2026-09-24 01:32:04 +02:00
jschoubben e7a90be3ee Merge pull request 'Issue 110: on a converged node a container on the runtime's own network cannot reach the resolver' (#97) from issues/110-the-resolver-and-the-default-network into main 2026-09-23 23:13:48 +00:00
jschoubben 3a8515273d Issue 110: on a converged node a container on the runtime's own network cannot reach the resolver
Found reviewing the resolver's conversion. Nothing fails while the node is adopted; it
fails at the flip, and it is the same split that decided which container survived the
hub's address change.
2026-09-24 01:13:30 +02:00
jschoubben 0e70ca0808 Merge pull request 'Issue 109: a container keeps the address it was made with' (#96) from issues/109-a-container-keeps-the-address-it-was-made-with into main 2026-09-23 23:02:48 +00:00
jschoubben 6ccd138729 Issue 109: a container keeps the address it was made with
Found when adopting the tunnel moved the hub's address: the declaration followed, the
running container did not, and the forge lost its database. Issue 102's rule broken one
level down, and issue 103's fix stopping one input short.
2026-09-24 01:02:16 +02:00
jschoubben 07ee79199a Merge pull request 'ADR 0105: what review settled — the tunnel adoption is implemented' (#95) from decide/0105-implemented into main 2026-09-23 22:39:08 +00:00
jschoubben f660620637 ADR 0105: what review settled — carried peers, the flip, the refusals, and keeping the hub's identity
Implemented in mesh-controller #49 and mesh-host #24. One proposal was rejected on the
record's own terms: converging the hub is not made to wait on other machines' migrations.
2026-09-24 00:38:54 +02:00
jschoubben f61f047a5d Merge pull request 'Issue 102 resolved and verified on the machine; issue 097's orphan was on the host network' (#94) from issues/102-resolved-and-097-worse into main 2026-09-23 22:15:49 +00:00
jschoubben 133e10a738 Issue 102 resolved, verified on the machine with both forwarders gone; 097's orphan was on the host network
The addresses follow: the control plane holds the ports the node gave, a recorded build
holds no address at all, and the two forwarders that had been holding the control plane
together are removed. 097's stranded container turned out to be listening on every
interface and connected to the mesh's store — by its own old database, which is the only
reason nothing was at risk.
2026-09-24 00:02:13 +02:00
jschoubben c209a575e1 Merge pull request 'ADR 0106: the bus is NATS; issue 104 resolved' (#92) from decide/0106-the-bus-is-nats into main 2026-09-23 21:40:03 +00:00
jschoubben e022798858 ADR 0106: the bus is NATS — native, built beside the migration, cut over after its core; issue 104 resolved 2026-09-23 23:39:17 +02:00
jschoubben 6f173a7912 Merge pull request 'Issue 108: the registry has no garbage collection, and two doors make it harder to add' (#91) from issues/108-registry-gc into main 2026-09-23 21:32:52 +00:00
jschoubben 73091dcb4c Issue 108: the registry has no garbage collection, and two doors make it harder to add 2026-09-23 23:32:29 +02:00
jschoubben f6ec64ee4e Merge pull request 'Issue 107: a declaration carries no order; rescue on an enrolled node is reconcile, not apply FILE' (#90) from issues/107-declarations-carry-no-order into main 2026-09-23 21:27:24 +00:00
jschoubben 671c2f3881 Issue 107: a declaration carries no order; rescue on an enrolled node is reconcile, not apply FILE
Both from the review of the issue-104 fix: a hand-applied file on an enrolled node is
recorded as carried and would remove the foundation, and nothing on the wire orders one
declaration against another.
2026-09-23 23:27:08 +02:00
jschoubben 2445d80565 Merge pull request 'Research 014: the bus on NATS — decide now, build in the lab, cut over once after the core' (#89) from research/014-nats into main 2026-09-23 21:15:22 +00:00
jschoubben a548b34f5d Research 014: fix the reference to ADR 0039 2026-09-23 23:14:57 +02:00
jschoubben b59907ee08 Research 014: the bus on NATS — decide now, build in the lab, cut over once after the core
Measured: AMQP is spoken in three places of the mesh's own code and in none of the
sdk or the modules; the predecessor's world is AMQP and retiring. NATS answers every
guarantee the bus relies on, durability via JetStream. Recommended: not underneath
the migration, not after it either — in parallel, one rehearsed rollout.
2026-09-23 23:14:39 +02:00
jschoubben 8638ba3a4f Merge pull request 'Issues 102–106 and ADR 0105: what the core migration found, and the hub adopting the predecessor's tunnel' (#88) from core/issues-102-106-and-tunnel-adr into main 2026-09-23 20:54:53 +00:00
jschoubben cb2117f1c4 Issues 102–106 and ADR 0105 from the core migration
Two birth-address outages and a registry that would have been the third; a
container that keeps a stale environment after its file changes; a host command
that applied a converged declaration to an adopted node; the hub and the vault
without seats. And the decision the operator made under it all: the hub adopts
the predecessor's tunnel in place, key and peers and range and port.
2026-09-23 22:50:10 +02:00
jschoubben ee2bdf220c Merge pull request 'Issue 101: taking a service its neighbours reach by container name cuts them off' (#87) from issues/101-a-service-reached-by-name-loses-its-network into main 2026-09-23 18:14:08 +00:00
jschoubben 61e4e971a4 Issue 101: taking a service reached by container name cuts its neighbours off
Found checking the third cutover rather than running it. The first two were safe by
accident — both are reached through a host port, which survives a change of owner.
This is the first constraint found that decides the order of the migration.
2026-09-23 20:13:53 +02:00
jschoubben f4f58e7c30 Merge pull request 'Issue 100: a secret the mesh mints cannot be the one the service it takes over already uses' (#86) from issues/100-a-minted-secret-cannot-be-the-one-the-service-already-uses into main 2026-09-23 00:54:21 +00:00
jschoubben 157c6edd50 Issue 100: a minted secret cannot be the one the service already uses
Found at the second cutover. Carrying a value in works only for a module's own
secrets; a secret answered by the provision is minted, and six catalogue modules
take one that way.
2026-09-23 02:54:09 +02:00
jschoubben e9e5df36bd Merge pull request 'Issue 099: a module's image pin ages into a downgrade, and taking it over is where that is discovered' (#85) from issues/099-a-pin-ages-into-a-downgrade into main 2026-09-23 00:48:02 +00:00
jschoubben 9faa985be3 Issue 099: a module's image pin ages into a downgrade
Three modules in a row on one machine; the first was found by taking it and cost a
three-minute outage. The runbook's answer is a rule a person must remember, which is
the shape this repository says not to settle for.
2026-09-23 02:47:50 +02:00
jschoubben cda7a4e348 Merge pull request 'Issue 094 diagnosed and resolved; 096, 097 and 098 opened from what it uncovered' (#84) from issues/094-diagnosis-and-096 into main 2026-09-23 00:37:27 +00:00
jschoubben 668ce3ad62 Issue 094 resolved: a given port names either end and is answered once
The first pass answered under both ends, which review showed is the same fault seen
from the other side where two mappings share a number. Verified on the machine: the
forge is back on the port its own configuration has always advertised.
2026-09-23 02:37:07 +02:00
jschoubben 0c8615aa8f Issue 098: taking a module replaces a configuration nobody compared
Found reading the second module's cutover rather than running it: the catalogue's
config drops a rule the machine's has, and no step puts the two side by side.
2026-09-23 02:23:58 +02:00
jschoubben 235b9ea0e5 Issues 096 and 097, and 094 diagnosed: a setting stored where it cannot work, and a resource that changed target
094's cause is one blind spot read from two ends, written up in its diagnosis; the fix
answers the first open question and not the other two, which become 096. 097 was found
looking at what the forge's cutover left running.
2026-09-23 02:20:48 +02:00
jschoubben 1b8e5043ee Merge pull request 'Issues 094 and 095, both found in the first module's cutover' (#83) from issues/094-095-from-the-first-cutover into main 2026-09-22 23:57:48 +00:00
jschoubben 515cb8adc1 Issues 094 and 095, both found in the first module's cutover 2026-09-23 01:57:10 +02:00
jschoubben 0d8b683ad7 Merge pull request 'Research 013: the forge and the registries — a seat answers the wrong question' (#82) from research/013-the-forge-and-the-registries into main 2026-09-23 01:03:10 +02:00
jschoubben 43b55d6664 Research 013: the forge and the registries — a seat answers the wrong question; issues 090 and 085 corrected from the code 2026-09-23 00:38:34 +02:00
jschoubben 6c81ea2204 Merge pull request 'ADR 0104: a provision may be answered by an adapter to the predecessor' (#81) from decide/0104-route-adapter into main 2026-09-22 23:57:54 +02:00
jschoubben a23ede495e ADR 0104: a provision may be answered by an adapter to the predecessor; issue 093 located; connectivity says how the proxy hands over 2026-09-22 23:57:38 +02:00
jschoubben a2323ade2e Merge pull request 'Issue 093: the successor proxy cannot serve what the predecessor still serves' (#80) from issue/093-proxy-handover into main 2026-09-22 23:57:04 +02:00
jschoubben d697c6f776 Issue 093: the successor proxy cannot serve what the predecessor still serves, so no web module can migrate one at a time 2026-09-22 23:45:49 +02:00
jschoubben a7b3f7823b Issue 092: it happened twice more — the registry's mesh name, and the mesh's own trust naming the default port 2026-09-22 23:37:05 +02:00
jschoubben cc3af29084 Issue 092: genesis publishes to a registry the container runtime does not yet trust 2026-09-22 22:43:17 +02:00
jschoubben fb9d3035d4 Issue 091: a module definition carries a machine port, measured across the catalogue 2026-09-22 22:26:14 +02:00
jschoubben dd81523eab Merge pull request 'Issue 085 resolved' (#77) from fix/085-resolved into main 2026-09-22 22:00:49 +02:00
jschoubben 5c5821b210 Issue 085 resolved: the packages port is a node setting; two of its open questions stay open 2026-09-22 22:00:38 +02:00
jschoubben c46c508c03 Merge pull request 'Issues 088, 089 and 090, found fixing 085' (#76) from fix/issue-085-followups into main 2026-09-22 22:00:02 +02:00
jschoubben 1728765fe3 Issues 089 and 090: a contributed route does not follow a moved port; the forge module cannot take over the forge genesis raised 2026-09-22 21:54:35 +02:00
jschoubben e609ccdbfb Issue 088: the forge's own address names a port it may not have 2026-09-22 21:41:58 +02:00
jschoubben 60eae92fb1 Merge pull request 'Adoption mode as built: ADRs 0101–0103, issues 084–086' (#75) from feat/adoption-mode into main 2026-09-22 21:02:00 +02:00
jschoubben 7bbbfb158a ADR 0103: a unit is found when an administrator installed it or the machine uses it 2026-09-22 20:02:18 +02:00
jschoubben 87ae893407 Issue 087: the controller cannot tell that a node's host is too old for what it sends 2026-09-22 19:43:01 +02:00
jschoubben c74ea2a4a8 Records say what the build does: 0101 names only measured daemons; 0102 adds to lists and keeps what it writes over; 0103 names every held kind, the found-service rule, conflicting found rules, guards, and what of 0100 it replaces; designs 05, 09 and 17 in step; issue 084's diagnosis in its own file 2026-09-22 19:38:16 +02:00
jschoubben ca1f648973 Issue 086: taking a module narrows a port the predecessor served, without saying so 2026-09-22 19:00:59 +02:00
jschoubben 019184ec0f Issue 085: the packages port given at genesis is not a setting, and a later module can undo it 2026-09-22 18:17:21 +02:00
jschoubben 213ab898d6 ADR 0103: what an adopted node holds and what its guard refuses; the node host and connectivity designs name 0102 and 0103 2026-09-22 17:52:58 +02:00
jschoubben 84761f0600 ADR 0102: the mesh writes into a shared file, never over it; issue 084 located 2026-09-22 17:46:06 +02:00
jschoubben 1901a90d68 ADR 0101 accepted; raising a mesh names it 2026-09-22 17:43:13 +02:00
jschoubben 347bbce633 ADR 0101 proposed: a machine's own resolver does not make it in use, as measured on a fresh machine 2026-09-22 17:38:44 +02:00
jschoubben 6165a7ae02 Issue 084: taking networking on an adopted node restarts every container, and the held runtime file blocks pulling 2026-09-22 17:09:21 +02:00
jschoubben dfadfd23c0 Merge pull request 'ADR 0100 (proposed): a node in use is adopted before it is converged' (#74) from feat/adoption-mode into main 2026-09-22 16:35:32 +02:00
jschoubben 3d4ab23830 ADR 0100: the machine's own traffic is known by its interface, not its source address 2026-09-22 16:35:25 +02:00
jschoubben 37252f9c3e ADR 0100: the guard lets the machine itself through; in use is a non-loopback listener; openings say from where; 09 in step with the flip 2026-09-22 16:34:53 +02:00
jschoubben f3152d827f ADR 0100 after re-review: the bus and registry stay reachable for enrolment; the mesh guards the store in a table that only refuses; a machine in use defined; the flip refuses while a found container is held; held containers and returning to adopted spelled out 2026-09-22 16:32:37 +02:00
jschoubben 02c40bcab4 ADR 0100 after review: found means unrecorded; assigning prepares, taking cuts over; openings through the found firewall on both paths; the mesh guards its own ports; ports kept as node settings; a converged genesis refuses a machine in use; designs 05, 07, 08, 09 and 17 in step 2026-09-22 16:28:31 +02:00
jschoubben 111456abb5 ADR 0100 accepted; the node host, connectivity, the node lifecycle and raising a mesh amended for a node adopted before it is converged 2026-09-22 16:20:27 +02:00
jschoubben 5fc3cbde4c Research 012: migrating a node that is in use, measured on the control-node; ADR 0100 proposed — a node in use is adopted before it is converged 2026-09-22 16:16:28 +02:00
jschoubben 21baf397a8 Merge pull request 'Issue 083 resolved: nothing the control queue carries is lost while the store restarts' (#73) from multiple-fixes into main 2026-09-22 14:43:00 +02:00
jschoubben 4de74880cb Issue 083: the proof, the replay caveat, and what is not closed, as three reviews found them 2026-09-22 14:39:29 +02:00
jschoubben 266ee34b5c Issue 083: the diagnosis describes held messages, the enrolment's order, and what is not closed 2026-09-22 14:24:31 +02:00
jschoubben 2d45aa5f42 Issue 083 resolved: nothing the control queue carries is lost while the store restarts; an enrolment claims its token and spends it last 2026-09-22 14:11:44 +02:00
jschoubben 4e74cf29e4 Merge pull request 'Issues 081 and 082 resolved; 083 opened' (#72) from multiple-fixes into main 2026-09-22 13:53:16 +02:00
jschoubben 0d40918e92 Issue 081: proven by the two-node bed, and what it found about baserow's data directory 2026-09-22 13:53:00 +02:00
jschoubben f6ded3102d Issue 083 opened (other control messages lost while the store restarts); 082's diagnosis carries its review 2026-09-22 13:34:05 +02:00
jschoubben 9a75f2b6a9 Issue 082: a report that arrives while the store restarts was lost; 081's diagnosis corrected on review 2026-09-22 13:25:33 +02:00
jschoubben 5d1a0372cb Issue 081 resolved: neither cache consumer can keep its keys under its login, so neither takes the shared cache 2026-09-22 12:27:12 +02:00
jschoubben f1e3925009 Merge pull request 'Issues 079 and 080, found by running the large mesh bed; 074's addendum' (#71) from multiple-fixes into main 2026-09-22 02:19:17 +02:00
jschoubben 35e44c3e5e Issue 081 opened (a cache consumer does not use its login); 079 and 080 diagnoses carry the review's consequences 2026-09-22 02:04:29 +02:00
jschoubben f760bf632c Issue 080: a cache grant let the consumer flush the server; 079: the names follow the resolver's rule for the private network 2026-09-22 01:39:29 +02:00
jschoubben e31e077fc8 Issue 079: the suffix is handed down, not written twice 2026-09-22 01:27:53 +02:00
jschoubben 17cc36069f Issue 079: every machine named twice over, found by the large mesh bed; what running that bed cost, on 074 2026-09-22 01:13:29 +02:00
jschoubben 446553dd1f Merge pull request 'ADR 0099; issues 077, 078 and 074 resolved; designs 08 and 20 amended' (#70) from multiple-fixes into main 2026-09-21 23:58:50 +02:00
jschoubben 3ea5e47c21 Review: 077 says what closed it and what did not; 078 names module issue; 074's retired fixture; ADR 0099's scope 2026-09-21 23:58:15 +02:00
jschoubben ead8913a74 Issue 078: module issue's orphan account, refused before it is made 2026-09-21 23:46:51 +02:00
jschoubben faa7196ad6 Issue 074: opened date restored 2026-09-21 23:44:22 +02:00
jschoubben 8d159cee84 Issue 074 resolved: the declared list is empty of WEARING; what the last three cost 2026-09-21 23:43:31 +02:00
jschoubben 0e0f0298c6 ADR 0099: a step that runs once names what it reads; issues 077 and 078 resolved; designs 08 and 20 amended 2026-09-21 23:33:47 +02:00
jschoubben c2a81cbb20 Merge pull request 'ADR 0098; issue 076 opened and resolved; ADR 0097's base refusal live; ADR 0096 proven against the public hub; 074 down to one bed' (#69) from multiple-fixes into main 2026-09-21 22:58:26 +02:00
jschoubben 6bc9df4b49 Review corrections: 076 and ADR 0098 say what the authority could and could not do; issues 077 (a fetched fact is fetched once) and 078 (secret accept takes any name) opened 2026-09-21 22:55:11 +02:00
jschoubben 6d3cb60949 Issue 076: the route-forwarding bed proves ADR 0098; what the run taught about the overlay 2026-09-21 22:44:24 +02:00
jschoubben 75af72ca27 ADR 0096: the copy is proven against the public hub by the genesis bed 2026-09-21 22:32:03 +02:00
jschoubben 252c6042e8 ADR 0098: a fact a provider makes at first start is fetched from it; issue 076 resolved; design 08 amended; 074 down to one bed 2026-09-21 22:27:32 +02:00
jschoubben bc373797c7 Issue 076 opened: a served fact made at first start cannot be served; ADR 0097's refusal of an undeclared base is live 2026-09-21 22:16:15 +02:00
jschoubben 559683318b Merge pull request 'Multiple fixes: issues 064, 066 and 020 resolved; 074 narrowed to two beds' (#68) from multiple-fixes into main 2026-09-21 22:13:20 +02:00
jschoubben ac11bea77f Issue 020 resolved: the symptom was the bed's, proven against Pebble and step-ca alike 2026-09-21 22:12:51 +02:00
jschoubben c1a10dc3f8 Issue 066 resolved: a file and its reader are guarded by a gate, proven by the coupled-pair spike; design 20 says so 2026-09-21 22:04:14 +02:00
jschoubben e993233004 Issue 074: six of the ten beds retired rather than converted 2026-09-21 21:53:11 +02:00
jschoubben 2d77911511 Issue 064 resolved: the package half placed by the genesis run, the image half by ADR 0097 2026-09-21 21:52:32 +02:00
jschoubben f11824d299 Merge pull request 'Multiple fixes: ADRs 0094–0097; issues 069, 049, 046 resolved; 064's image half decided' (#67) from multiple-fixes into main 2026-09-21 21:05:13 +02:00
jschoubben 43ca9ce0ae ADR 0097: an undeclared base is said, not yet refused 2026-09-21 20:48:19 +02:00
jschoubben 71072e240d ADR 0097: a vendor image is a declared build input; issue 064's image half decided; design 18 amended 2026-09-21 20:45:36 +02:00
jschoubben 8d981e21b1 ADR 0096: an upstream image is copied between registries; issue 046 resolved; design 18 amended 2026-09-21 20:43:07 +02:00
jschoubben 39a01b7f7e ADR 0095: the control plane is the way to ask a module; issue 049 resolved; design 19 amended 2026-09-21 20:38:23 +02:00
jschoubben acc9824949 ADR 0094: a module may hold several secrets from one provider; issue 069 resolved; design 24 amended 2026-09-21 20:29:29 +02:00
jschoubben a259d1292f Merge pull request 'ADR 0093: a fixture that runs a module's runtime carries its name; 074 diagnosed; 075 resolved' (#66) from feat/mesh-tests-and-runtimes into main 2026-09-21 19:36:44 +02:00
jschoubben 76cd91413e ADR 0093: a fixture that runs a module's runtime carries its name; 074 diagnosed, two beds converted; 075 resolved 2026-09-21 19:30:35 +02:00
jschoubben 6427afe6ae Merge pull request 'Issues 072 and 073 name their merges, the branches being gone' (#65) from chore/merges-named into main 2026-09-21 19:25:06 +02:00
jschoubben d92068be7b Issues 072 and 073 name their merges, the branches being gone 2026-09-21 19:24:28 +02:00
jschoubben 6a35e5a58c Merge pull request 'Multiple fixes: status vocabulary aligned, 065/026/070/007 resolved, 066/069/049/046 located, 064 diagnosed' (#64) from feat/multiple-fixes into main 2026-09-21 19:23:34 +02:00
jschoubben bf4d99843f Merge pull request 'Issue 072 diagnosed: genesis registers the manifest its build produced; design 17 says so' (#63) from feat/one-controller-manifest into main 2026-09-21 19:23:31 +02:00
jschoubben 702289ad7d Merge pull request 'ADR 0089: a bed reads the catalogue it proves; issue 073 diagnosed; issues 074 and 075 opened' (#62) from feat/beds-read-the-catalogue into main 2026-09-21 19:23:16 +02:00
jschoubben 6ee95a3f31 ADR 0090: a failure is the same by resource id, not by the host's words 2026-09-21 19:23:04 +02:00
jschoubben 159a583cf0 Issue 069: the count is the catalogue's, not the report's 2026-09-21 17:51:40 +02:00
jschoubben 62ff9a5bd3 Issues 046, 049 and 069 located, each with the decision its fix needs named 2026-09-21 17:50:42 +02:00
jschoubben 16442ecd23 ADRs 0091 and 0092; issues 026 and 070 resolved; designs 18 and 24 amended 2026-09-21 17:50:41 +02:00
jschoubben ffb7fa52ec Issues 007 resolved, 066 located, 064 diagnosed 2026-09-21 17:48:38 +02:00
jschoubben fcf34727fe ADR 0090: a failure that repeats is said to be stuck; issue 065 resolved
The controller kept one report per machine, replaced, so a resource nothing can
ever apply looked like a failure that had just happened, every few minutes, for
ever. It now counts identical reports and status says stuck after three.
2026-09-21 17:43:02 +02:00
jschoubben a5e351f4cb Nine resolved issues name where their fix landed
The cycle check now asks a resolved issue for its owner, and these had none.
2026-09-21 17:39:34 +02:00
jschoubben e3279f4b5b Issue 054 names where its fix landed 2026-09-21 17:39:09 +02:00
jschoubben d52687d0c8 Issues 067 and 068 name where their fix landed 2026-09-21 17:38:48 +02:00
jschoubben 36d9b38a0d An issue is open, diagnosing, located, resolved or wontfix — nothing else
The playbook, the README and the status skill knew five statuses; the cycle check
knew a sixth, 'fixed', and not 'wontfix'. Eleven issues sat in the sixth for weeks
with their fixes shipped, one step short of closed. They are resolved; the check
refuses the word from now on and accepts the one the playbook allows.
2026-09-21 17:38:17 +02:00
jschoubben da6414c27c Issue 072 diagnosed: genesis registers the manifest its build produced; design 17 says so
ADR 0069 had already placed the controller's manifest in its own repository, and the
raising design called the catalogue copy a thing to remove. The installer's step 3 was
handed that manifest by the build and step 9 read a second copy anyway.
2026-09-21 15:18:32 +02:00
jschoubben 5dcb9dfbf6 ADR 0089: a bed reads the catalogue it proves; issue 073 diagnosed; issues 074 and 075 opened
The end-to-end design held 'the run rebuilds what it tests' for binaries and images
and not for manifests. The beds' inline copies fell into three kinds; only the first
is a stale copy. The other two are named: a mesh test wearing a catalogue module's
name (074) and a stocked runtime image the run never rebuilds (075).
2026-09-21 14:35:08 +02:00
jschoubben 1c90777124 Merge pull request 'Issues 031, 035 and 054 name their merges, the branch being gone' (#61) from chore/migration-blockers-merges into main 2026-09-21 13:50:06 +02:00
jschoubben ba07b43b18 Issues 031, 035 and 054 name their merges, the branch being gone 2026-09-21 13:49:56 +02:00
jschoubben 1db93ffc64 Merge pull request 'ADRs 0087 and 0088; issues 031, 035 and 054 resolved; issue 041 surveyed; issue 073 opened' (#60) from feat/migration-blockers into main 2026-09-21 13:49:09 +02:00
jschoubben ea42b00d0a Issue 041: mongodb names its secrets' owner; the three conversions are proven 2026-09-21 13:47:39 +02:00
jschoubben 198e4db0a3 Design 07: the base ruleset closes the hub's port until the filter module derives one 2026-09-21 13:30:26 +02:00
jschoubben a438cbdac3 Issue 041: the declared exceptions, surveyed and partly converted 2026-09-21 12:53:14 +02:00
jschoubben 5fa35362ac Issue 073: beds carry copies of catalogue manifests 2026-09-21 12:31:04 +02:00
jschoubben ee67f69ef4 ADRs 0087 and 0088; issues 031, 035 and 054 resolved; designs 05, 07 and 18 amended
A seeded file is created once (0087); the foundation filters before anything
listens (0088); a machine becomes the last thing it was told (design 05).
Each says how it is checked.
2026-09-21 12:11:50 +02:00
jschoubben a211aecdfa Merge pull request 'ADR 0086 (a secret reaches a process as a file), the vault shipped, issues 041 and 072' (#59) from feat/secret-not-in-environment into main 2026-09-21 11:59:21 +02:00
jschoubben 5fb8f06a91 Issue 072: the controller's manifest exists twice; decisions index regenerated for ADR 0086 2026-09-21 10:34:17 +02:00
jschoubben ec7b6e0748 ADR 0086: a secret reaches a process as a file, and an exception is declared
Closes issue 041 by decision and by code on the same branch: the catalogue
engine refuses a secret in a container's env, and a secret-carrying env-file
unless the container declares its reason; the controller reads all six of
its credentials from files; design 13 states the rule and how it is checked.
2026-09-21 10:10:33 +02:00
jschoubben 7015bf58ac The vault shipped: as-is 06 describes it, design 24 is implemented, 071 names its merges 2026-09-21 10:10:33 +02:00
jschoubben aa9b5b99e4 Merge pull request 'The secrets vault: handoff, the amended decision, issues 069–071, and the reconciliation of open issues' (#58) from feat/secrets-vault into main 2026-09-21 10:03:28 +02:00
jschoubben 61f61b5d06 Reconcile the open issues against main
An audit of the six code repositories found eleven open issues fixed on main
with commits and beds to show (025, 027, 033, 036, 037, 040, 045, 047, 050,
052, 053), three partly (007, 026, 035), nine not (020, 031, 041, 046, 049,
054, 064, 065, 066) and one whose fix would live outside those repos (006).
Resolved ones name their evidence; partly ones say what remains; 041 records
that the exposure has widened since it was reported.
2026-09-21 01:52:53 +02:00
jschoubben ffd7592687 Design 24: the export says what an earlier key opens 2026-09-21 01:16:54 +02:00
jschoubben 1bce3f33a8 Designs 07 and 21 and issue 071: genesis now makes the root secrets
The installer half of the amended ADR 0085 is built and proven by the
genesis bed's root-secrets step; the foundation design closes its open item
and the installation design says what the installer does and what it still
cannot check.
2026-09-21 00:50:56 +02:00
jschoubben 8607d21110 Design 24: pair credentials are sealed to the operator too 2026-09-21 00:36:30 +02:00
jschoubben 050a2d08e8 Playbook 07: a feature worktree needs the siblings the lab reads
A bed run from .work/<slug>/mesh-lab derives mesh-tools and mesh-sdk by
sibling path and fails at once when the directory holds only the touched
repos. Detached worktrees on main, never symlinks.
2026-09-21 00:27:01 +02:00
jschoubben baa3351552 Amend ADR 0085: the vault is a foundation module and holds the root secrets
Recorded on the record, dated, before anything shipped against the sentences
that change. The vault is installed at genesis like the store and broker, one
per mesh, and holds every secret a module has for itself sealed a second time
to an operator key whose private half never enters the mesh — the break-glass
path the first version left open, without a key one place holds.

Design 24 says how; 07 and 21 say what genesis does not yet do; issue 071
names the fixed credentials the foundation is raised with today.
2026-09-20 23:55:01 +02:00
jschoubben 187389b7d1 Hand off the secrets vault to mesh-catalog; open issues 069 and 070
Design 24 flips to in-progress with mesh-catalog as its owner (playbook 04).
Starting the build surfaced two gaps the decision did not settle: a module
requiring `secret` receives exactly one value (069), and no command can
accept an operator's value into a consumer↔vault pair (070). Both opened as
issues rather than improvised around.

Also fills fixed-by on 067 and 068, which the cycle check refused as resolved
with no reference.
2026-09-20 23:08:28 +02:00
jschoubben 731027c009 Merge pull request 'Graduate issues 067 and 068 — provider scoping and the secrets vault' (#57) from multi-node/harden-and-prove into main 2026-09-20 21:30:12 +02:00
jschoubben fc4ab370d6 Graduate issues 067 and 068 to decisions and to-be designs
ADR 0084 (extends 0027) — a provision is served by a node-scoped provider the
consumer selects, defaulting to co-location; a module may instead carry a private
embedded instance that is not a provision. Design: 01-to-be/23-choosing-a-provider.

ADR 0085 (extends 0031) — a secret is a provision and the vault is the module that
provides it; a module's own local secret becomes an ordinary pair credential that
rotates through the existing machinery, while the controller's provisioning-credential
mint (0048) is unchanged. Design: 01-to-be/24-the-secrets-vault; doc 13 amended to
cross-link the non-pair secret.

Issues 067/068 marked resolved with amended-design set. ADR index regenerated;
records and index checks pass.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-20 21:27:29 +02:00
jschoubben d66579c8f6 Issue 068 — secrets have no owning module
Store and broker seats were made ordinary modules; secret-minting is still a
privileged property of the controller that no module owns. Propose the vault
become a module that provides a secret provision (generate/hold/rotate/backup/
audit), node-scoped like every other provider (issue 067), subsuming the three
secret paths and giving local-secret rotation a home.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-20 21:24:42 +02:00
jschoubben c4ff0478b7 Issue 067 — fold in the embedded-vs-provisioned axis
Naming which provider is only half of how a module gets a database. The other
half: a module may carry its own version/fork-pinned instance, module-network
only, no published port, not a provision — and the model has no word for it.
Add it as a second axis with its own open questions.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-20 21:24:42 +02:00
jschoubben 1ce9ffe79f Issue 067 — a provision cannot name which provider serves it
The mesh models provisions as mesh-scoped (one provider of a kind, a single
mesh-store). But node-specific services delivered to the mesh was the plan from
the start: both nodes already run their own postgres, SQL server, redis and
object store, and identity — currently single — already serves apps on a second
node. The model cannot express which provider serves a consumer, so it collapses
a deliberately per-node fleet to one. Provider scoping is a whole-mesh decision
across postgres/s3-bucket/oidc, not an SSO patch.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-20 21:24:42 +02:00
jschoubben df2c3f63a5 Merge pull request 'Issues 065 and 066 — the apply failure model's two gaps' (#56) from issue/065-066-apply-failure-model into main 2026-09-20 14:28:18 +02:00
jschoubben a9bc0c1cc5 Open issues 065 and 066 — the apply failure model's two gaps
Surfaced while reviewing the host apply loop for the controller/host
split discussion. 065: a permanently-failing resource retries for ever
with no escalation — reported, but never raises its hand as stuck.
066: a partly-applied declaration leaves a mixed state with no rollback,
which is harmless for independent resources and unexamined for pairs
that are only correct together. Both are questions HQ must answer, not
incidents — status open, no fix proposed.
2026-09-20 14:28:04 +02:00
jschoubben 6d28683898 Merge pull request '057 and 058 resolved — fixes merged and proven' (#55) from issue/057-058-resolve into main 2026-09-20 12:55:16 +02:00
jschoubben 82c03a8d10 057 and 058 resolved — their fixes merged and proven
The P1-sweep PR left both at 'located'; the fixes have since merged
(mesh-controller #31, mesh-tools #10) and the bed proves them. Flip to
resolved with their fixed-by.
2026-09-20 12:55:00 +02:00
jschoubben 31c5aee72d Merge pull request 'ADR 0083 and issues 057/058/060/063/064 — the P1 sweep' (#54) from issue/057-058-one-push-patient-runtime into main 2026-09-20 12:53:24 +02:00
jschoubben 05857f71f0 Accept ADR 0083 — one push leaves the mesh consistent 2026-09-20 12:52:55 +02:00
jschoubben 98e9e62fce Issues 060 resolved, 064 opened; 057/058 on this branch too
060: 44 of 70 modules now mesh-buildable (was 8) — the mechanical
majority, proven by direct builds and the bed. The structural
remainders: external-dependency fetch (new issue 064) and route-proxy's
cross-repo build context. 064 records the build environment's isolation
from public npm and Docker Hub.
2026-09-18 02:52:36 +02:00
jschoubben 0d5693cc2c ADR 0083: the cascade compares against what was last sent, not a snapshot 2026-09-18 02:38:00 +02:00
jschoubben 1572d74a18 Issue 063 — the foundation's ports are forwarded, resolved
The broker's port was accepted on input but not forwarded, so a joined
node reached a DNAT'd broker only until it restarted. Fixed in
mesh-controller (25e42b3), proven by the built-store-cross-node bed
adopting the foundation's broker (which restarts it) and the joined
node still receiving declarations.
2026-09-18 02:20:11 +02:00
jschoubben 9fd7e6c458 ADR 0083 proposed; 057/058 diagnosed and located
One push leaves the mesh consistent (the 057 decision, proposed for
acceptance); the shared runtime waits for its broker (058). Fixes on
mesh-control fix/one-push-is-enough and mesh-tools
fix/the-runtime-waits-for-its-broker; the built-store-cross-node bed
enforces both.
2026-09-18 02:10:01 +02:00
jschoubben 5703c86cf6 Merge pull request 'Issue 059 resolved — the broker credential resolves the seat' (#53) from issue/059-broker-address-seat into main 2026-09-18 01:13:13 +02:00
jschoubben d85d74d607 059 resolved — the broker credential resolves the seat
Fixed by mesh-controller PR 30 (79a9e17), rebased onto the merged
registry-reach train and proven by a green fresh run of the no-fake
two-node bed built from that commit.
2026-09-18 01:13:03 +02:00
jschoubben 46efb7ae98 Merge pull request 'ADR 0082 and issues 042/048/061/062 — the registry is reached by name and trusted by the overlay' (#52) from issue/042-048-registry-reach into main 2026-09-18 01:00:55 +02:00
jschoubben 4279e4914b 042/048/061/062 resolved — the network carries the registry trust
One green fresh run of the no-fake two-node bed is the proof: node2's
consumers open the store and broker the mesh built and adopted, over
the overlay. Fixed by mesh-controller PR 29, mesh-catalog PR 26, gated
by mesh-lab PR 34.
2026-09-18 00:26:50 +02:00
jschoubben a0acaad86d Issues 061/062 — two silent-success defects the no-fake bed surfaced
061: the broker module's provisioner never ran; its runtime container
named no command and the image default is the tool host. 062: a failed
artifact-store lookup composed the network without the registry trust,
turning a transient error into permanent silent state. Both located,
fixes on the 042/048 train branches.
2026-09-18 00:23:52 +02:00
jschoubben 08c814aa0a 0082 takes its homes: 042/048 located against it, the delivery design cites it
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 23:02:32 +02:00
jschoubben 32fa6b40f5 ADR 0082: the registry is reached by name, and the overlay is its security
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 23:02:18 +02:00
jschoubben 8df3469813 042 and 048 move to diagnosing — the registry-reach work begins
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:44:39 +02:00
jschoubben e7869d27bd Merge pull request 'Issue 060 — the mesh cannot build most of its own catalogue' (#51) from issue/060-catalog-not-mesh-buildable into main 2026-09-17 22:39:01 +02:00
jschoubben 0b002ef39c Issue 060 — the mesh cannot build most of its own catalogue
Only 8 of 70 modules carry a build section; the other 62 exist through an out-of-band
build script plus placeholder rewriting, so the mesh's own build-and-deliver path has
never run for them. Found by the no-fake bed; blocks it and the migration's delivery
assumption.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:38:48 +02:00
jschoubben 101848242f Merge pull request 'Issue 059 — the broker credential names the hub, not the broker's node' (#50) from issue/059-broker-address-review into main 2026-09-17 22:37:26 +02:00
jschoubben 5395b6a47f Merge pull request 'ADR 0081: decisions are in the chain — no orphan records' (#49) from process/decisions-in-the-chain into main 2026-09-17 22:36:48 +02:00
jschoubben 90e4a368dc ADR 0081: a decision nothing cites is not yet in the chain
Decisions were the one link the cycle checks skipped, and measuring found 19 of 70 records
orphaned — the credential flow and the module-runtime cluster among them, which is how a
stale premise about a settled decision survived in working memory. cycle.py now refuses an
accepted record nothing cites; the 19 got true homes (design frontmatter, the playbook that
implements 0021, META for the process records). The overview names the practice: spec-driven
development with provenance.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:36:33 +02:00
jschoubben d1b5210d83 Issue 059 — the broker credential names the hub, not the broker's node
Findings of an adversarial review of the 055 fix: hub!=broker conflated, membership tested
as has-address, silent staleness, portless silent fallback, two hubs unrefused. One review
claim recorded as disputed against live wire measurements.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:29:25 +02:00
jschoubben 1162f7fe25 Merge pull request 'ADR 0080: the development cycle is checked, not trusted — plus review fixes' (#48) from process/the-development-cycle into main 2026-09-17 22:28:22 +02:00
jschoubben becae7ba51 ADR 0080: the development cycle is checked, not trusted
The flow the process overview draws — idea/symptom -> decision -> to-be design -> code ->
as-is — was enforced by nothing. cycle.py now refuses a to-be design naming no decision, an
in-progress/implemented design naming no owning code, a located/fixed issue with no owner,
a fixed/resolved issue with no fix, and a graduated research overview that does not say what
it became. AGENTS.md carries the cycle and a where-to-look table so a fresh session (or a
cleared context) finds the chain in frontmatter instead of assuming it. Grounding the check
surfaced two real gaps, fixed here: the work-ahead design named no owning code, and research
003 listed one became target twice.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:28:03 +02:00
jschoubben 94fee5d849 Fix what the records review found
The glossary's authority page still named the controller's seat the-controller in two
entries, contradicting its own seat section after ADR 0079; issue 058's heading kept the
pre-renumber 059; 055's fixed-by named branches that stop existing after merge (now merge
commits/PRs) and its located-in listed file paths where the convention wants repos; 056's
located-in named mesh-host, which received no fix, instead of mesh-catalog; and the design
layer never said the one-store/one-broker property is enforced — 07-the-foundation and the
installation table now state the seats.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:24:57 +02:00
jschoubben 5856555390 Merge pull request 'Issue 058 — a provisioner runtime crash-loops until the overlay is up' (#47) from issue/059-provisioner-startup-resilience into main 2026-09-17 21:57:03 +02:00
jschoubben 0ffdf4e581 The issue takes the next free number, 058
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 21:56:49 +02:00
jschoubben dd5bf9cd8c Issue 059 — a provisioner runtime crash-loops until the overlay is up
Observed in the two-node bed: a cross-node runtime exits on 'timed out fetching the
broker's certificate' until the tunnel forms, then settles. The retry belongs in-process.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 21:56:40 +02:00
jschoubben 0e9d4ab86a Merge pull request 'ADR 0079: foundation seats are named after their servers; issue 056 resolved' (#46) from multi-node/foundation-seats into main 2026-09-17 02:16:11 +02:00
jschoubben ee7abe0a8e ADR 0079: foundation seats are named after their servers; issue 056 resolved
The store and broker were "one per mesh" by convention only. Each foundation module now
claims a mesh-scoped seat named after the server it guards — postgres/mesh-store,
lavinmq/mesh-broker — and the controller's seat is renamed the-controller -> mesh-controller
so all three follow one rule. The resolver refuses a second holder, closing 056. Glossary,
the foundation and installation docs, and the decisions index follow the new name.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 02:15:08 +02:00
jschoubben f70973e222 Merge pull request 'Issue 055 resolved and proven; 056 and 057 opened' (#45) from multi-node/harden-and-prove into main 2026-09-17 01:44:59 +02:00
jschoubben 812dc3b303 Issue 055 resolved, and 057 opened for its silent operational edge
055 is proven end-to-end by the two-node bed: a consumer on a joined node reaches the
adopted broker over the overlay, its binding names the control-node, and its vhost is
minted. Three fixes — the broker's amqps port in the firewall, the broker credential
naming the overlay not the public address, and pushing the provider node after the remote
consumer arrives. The last is an operator ordering, not a code fix, and its silent-failure
edge (a cross-node consumer that never provisions and only crash-loops) is opened as 057.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 01:35:47 +02:00
jschoubben 33fc322f4d Issue 055 diagnosed — the module broker URL uses the public address, not the overlay
A two-node bed pins it: the firewall gap (broker's 5671 not in listens) is fixed
in mesh-catalog; the real bug is modules.go:294/build.go:231 building the broker
URL with the genesis public MESH_BROKER_ADDRESS, which a joined node cannot reach.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 01:02:07 +02:00
jschoubben bec1dd0082 Issue 056 — an adopted module assigned to a second node raises a second server
"One postgres, one lavinmq" (ADR 0078) holds only by genesis assigning the
adopted modules to the control-node alone; nothing enforces it. A second assign
raises a divergent second server, silently. Open questions: a mesh-scoped
exclusive seat, or explicit adoption the controller refuses to place elsewhere.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 00:22:18 +02:00
jschoubben a32f7e50af Merge pull request 'Establish the repo for the completed Phase 0-3 build' (#44) from establish/phase-3-close into main 2026-09-17 00:05:20 +02:00
jschoubben 1111bd84d7 Establish the repo for the completed Phase 0-3 build
Settles the design repository now that the self-upgrade build is on main:
- Records the two decisions that shipped without a record — ADR 0077 (the
  controller/foundation/node vocabulary) and ADR 0078 (the store and broker are
  ordinary modules); accepts ADR 0075 and 0076, which shipped work rests on.
- Fills issue 051's amended-design and wires ADR 0078 into 07-the-foundation.
- Sweeps the repo rename (mesh-control -> mesh-controller) into the mutable docs
  now that the forge repo is renamed; updates the glossary note and repos.md.
- Fixes the six broken links from the design-doc renames, indexes the glossary,
  regenerates the decisions reading order.

Both checks (records.py, index.py) are green. Statuses stay honest: the build is
on main and lab-proven but not deployed as the production mesh, so the to-be docs
remain in-progress and the as-is layer (the hal mesh) is unchanged — graduation
to implemented + as-is belongs to deployment, not merge.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 00:04:58 +02:00
jschoubben 873cc4007b Merge pull request 'Glossary, the mesh-controller/foundation vocabulary, ADR 0076, Phase 3 closed' (#43) from issue/047-the-other-half into main 2026-09-16 23:25:51 +02:00
jschoubben 0073e52881 Phase 3 closed — the store and broker are ordinary modules, issue 051 fixed
Marks WBS 3.2/3.3/3.4 done and resolves issue 051: the foundation's store and
broker are adopted in place as the postgres and lavinmq modules, upgradeable
through their stated windows, source-tracked by status. A bare-metal mesh runs
one postgres and one lavinmq, proven 22/22 in the one-node lab. The two follow-up
gaps are tracked as issues 054 and 055.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 23:07:03 +02:00
jschoubben fcf3b6f4d5 Issues 054, 055 — the debt adopting the store and broker leaves
054: the adopted servers bind 0.0.0.0 from genesis but the packet filter is
installed later, so there is a window where they are open with only bootstrap
credentials. 055: the servers bind on the control-node and the one-node bed
cannot prove a consumer on another machine can reach them over the overlay.

Both are follow-ups to issue 051's adoption (WBS Phase 3), tracked rather than
rushed.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 22:04:28 +02:00
jschoubben 43f3565617 WBS: correct 3.1 phrasing — a module gets a database only if it asks
Not "every module's database"; modules request one via requires postgres-database.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 21:20:00 +02:00
jschoubben addfdd6940 WBS: Phase 3.1 done — the store is adopted as the postgres module
One postgres, adopted in place, proven 22/22 in the one-node lab. Notes the
pre-filter exposure window as a follow-up.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 21:14:13 +02:00
jschoubben 33a00d5656 Adopt the glossary's vocabulary in the mutable design docs
"control plane" -> controller and "substrate" -> foundation throughout
03-DESIGN, 00-META and the README, with 06-the-control-plane.md and
07-the-substrate.md renamed to 06-the-controller.md and 07-the-foundation.md.
The immutable 02-DECISIONS records keep their original wording (and links to
them are unchanged) — a term retired here may still appear there, which the
glossary explains how to read.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:48:52 +02:00
jschoubben f9f48fbbf7 Glossary: one name per thing, and the words we retired
Locks the vocabulary that kept drifting in conversation — controller (not
"control plane"), foundation (not "substrate"), node and control-node, seat /
bench / claim, package vs artifact. AGENTS.md points at it as the authority.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:22:31 +02:00
jschoubben b43b60b183 ADR 0076: the SDK is a published package the toolchain resolves by version
Records the decision the package-registry work turns on — the SDK is built on a
public base and published before the toolchain that consumes it, so nothing is
circular; mesh-tools stays the thin toolchain base but resolves the SDK by
version. Reconciles docs 12/17/22 and indexes the record.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 10:28:08 +02:00
jschoubben 17c2e061df Phase 1 was mostly a phantom: correct doc 19 and the WBS
The protocol spec claimed the envelope and grant drifted across implementations.
Inspection showed the wire agrees — envelope required headers match, the two
optional ones are legitimately optional, and the grant wire (the contributions
file) is identical on both sides. The disagreement was in dead types, now removed.

So phase 1's 'make them agree' work is done by deletion and correction. What
remains is a conformance fixture as prevention — pinning the envelope and the
contributions file so a future change that breaks agreement fails a test — and a
full per-capability suite is deferred until a third language actually needs it.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 00:07:26 +02:00
jschoubben a40595fa08 ADR 0074: correct the evidence — the live wire agrees
The record claimed the two implementations already disagreed. Inspection showed
the live wire agrees: the disagreeing grant types were dead (removed), and the
envelope's two extra headers are optional and set when relevant, not missing.

The danger was dead types contradicting the wire, not live disagreement — which
is a sharper reason for specifying the wire and checking against it, not a weaker
one. The model stands; the conformance suite's job is prevention rather than
repairing a present break.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 00:06:49 +02:00
jschoubben 6e1099a905 Reorder the work: implement and unit-test first, run the lab once at the end
The correction the operator pushed: stop running a 20-minute lab against a mesh
mid-transformation, debugging paths the next phase deletes. The base-build hang
is almost certainly the SDK resolving from a git URL inside a docker build (issue
053), which Phase 2 removes — so debugging it on the current shape is debugging
deprecated code.

Phase 0 folds in: the installer's own regressions are fixed and committed;
whether it runs green is the final acceptance test, after the phases that change
its build path are in. Faults that can be reasoned out of the code path are, by
reading rather than running.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 00:02:51 +02:00
jschoubben 2e1ae061b5 The work ahead: four phases, dependency-ordered, each ending at a run
Everything decided this cycle and not yet built. Phase 0 gets the installer
green, because nothing else is testable end to end without it. Phase 1 makes the
protocol one thing and fixes the Go/TS drift the installer's own provisioning
exercises. Phase 2 stands up the private package registry ADR 0014 assumes and
publishes the SDK into it. Phase 3 adopts the substrate so one postgres and one
lavinmq serve everything, which is the hardest and needs all three above.

Order is dependency, not preference. Each phase ends at a run rather than a
paragraph, because a phase that ends at a claim is how things went missing this
cycle without anything complaining.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 23:41:42 +02:00
jschoubben 03fc13b84b 051: one lavinmq, not two — the duplication is running, not latent
The lavinmq module raises its own server container and provisions vhosts on it,
separate from the substrate's mesh-broker. A mesh with the module assigned runs
two LavinMQ servers where one belongs — the exact AMQP twin of the two postgres
containers.

The module's own provisioner already assumes one server: it creates a vhost per
consumer, named for the login, isolated by the vhost boundary — the analog of
postgres's database-per-login. So the mesh's own control traffic is the / vhost
and every consumer's broker is a vhost beside it, all on one server.

Adoption therefore means the module does not run a server of its own: its server
is the substrate broker, adopted, and the module contributes the provisioner,
tools, events and bootstrap against it. The isolation model is already built;
what remains to design is only raising the one broker at genesis and then holding
it as a module, over the broker it is.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 22:31:17 +02:00
jschoubben a062f181ce The twelve-module table names distribution, as the catalogue does
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 22:03:39 +02:00
jschoubben 87f464cc5f A finished mesh holds twelve, and two rows were one module each
The substrate's store and the postgres module are the same thing. They were two
rows only while the substrate was a different KIND of thing — a store raised from
a bundle cannot provide postgres-database, so anything wanting a database needed
a second server. It is visible on any mesh built today: mesh-store and postgres,
two containers, the same image.

Which name survives is settled by the naming rule. Where a consumer speaks a
protocol the interface is the protocol, and "database" is not a capability. The
control plane's own queries use distinct on and on conflict, so the coupling is
to postgres and a "store" module would advertise a swap that fails the first time
anyone tries it.

The broker collapses the same way with a different outcome: amqp IS a protocol
that several implementations speak, so amqp is a legitimate provision and lavinmq
is one provider of it.

So adopting the substrate is not only an upgrade path — it is two rows of a
mesh's module list becoming one, twice. And it leaves nothing that is a specialty
after the pivot, which is the claim the whole design rests on and is not true
today for exactly those two.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 20:26:42 +02:00
jschoubben 5a71e8659f The host's unit exists; the installer just does not place it
Reported as 'the host does not survive a reboot', which reads as a mesh that
cannot come back. mesh-host/packaging/ ships nox-mesh-host.service and two
companions. The installer declines to place them because a unit file is a
packaging decision, and the lab starts the host with --host-in-background, which
says in its own help that it does not survive a reboot.

So the lab run failing this was the lab being honest, and the gap is the step
that puts a shipped unit on a machine — narrower and more fixable than what I
wrote.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 20:15:48 +02:00
jschoubben 1e8273eb19 Correct 0075: co-residence is not the exclusivity that matters
The first draft implied gitea and the registry could not share a machine, because
registry claims the-artifact-store at node scope and I carried that across to
gitea without asking what the claim is for.

A machine running gitea for git and packages alongside a registry serving
artifacts is an ordinary arrangement. They are different ports doing different
jobs, and nothing about one being the mesh's artifact store requires the other
not to exist.

The exclusivity that matters is mesh-wide and already expressed: provides at mesh
scope means two providers are two answers, and the resolver refuses until one is
assigned. Forbidding co-residence adds nothing and forbids something reasonable.

Whether registry should still hold that claim is left open rather than answered
from outside its manifest — it may be protecting something about its port or its
data directory that nobody wrote down.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 19:35:05 +02:00
jschoubben 199c2f20bb ADR 0075 — an artifact store is a provision, and a package registry is another
Two questions circling turned out to be one asked twice: should gitea be the
mesh's registry, and where does the SDK come from. The framing that dissolves
both is that artifact-store is already a provision and the registry already
provides it — so this was never about replacing a component. It is a second
provider of an existing provision, which this mesh has a mechanism for and uses
for certificate authorities already.

So: two provisions, because they are two jobs. artifact-store is content
addressed, pinned by digest, no versions and no ranges — what the mesh delivers
to machines. package-registry is an ecosystem's own, addressed by name and
version — what code resolves when compiled. Conflating them is how a mesh that
pins everything ends up rebuilding one commit into two different things.

The small registry stays the provider genesis installs, not because it is better
but because of what it is: a directory and one container, installable where there
is no database and no control plane. Gitea needs both, and the pivot needs
somewhere to publish before either exists.

Gitea also provides artifact-store, so a mesh may choose it — and choosing it
answers issues 042 and 048 by adopting something that already has accounts and
TLS, rather than reimplementing them in a registry that has neither.

Moving between providers is a designed act with a verification step that is easy
to skip and is the only thing between it and a mesh that cannot restart its own
control plane.

And the bootstrap still has no package registry when the first build needs one.
Named rather than solved, so the next person does not discover it.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 18:06:03 +02:00
jschoubben 94229d95eb The installation, written out in full
Every step from a bare machine to a mesh that maintains itself, in three phases,
with each step named as the installer prints it.

The point of writing it out is the shape it exposes. The installer owns twelve
steps and ends at a mesh that RUNS. Seven more turn that into a mesh that WORKS —
the shared base, a store that is a provider rather than the control plane's own
memory, the catalogue, the replay of what was built before the catalogue existed,
the control plane rebuilt through the module path, the private network with the
node actually placed on it, and the packet filter. None of those seven is the
installer's. They are things somebody types, which is why a test had to be
written to discover they were missing.

Machines arrive last, in phase three, because a machine joining a mesh that
cannot build anything proves enrolment works and nothing else.

And five things that are not yet true are named rather than implied: phase two is
manual, a second machine cannot pull what the mesh built, ADR 0014 assumes a
private package registry that genesis has not installed when the first build
needs it, the host agent does not survive a reboot, and nothing can contradict a
claim that a machine was installed this way.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 16:51:52 +02:00
jschoubben 27b2161379 The SDK's delivery is decided; the git URL violates it
ADR 0014 is accepted and unambiguous: each module consumes its dependencies from
the private registry, the mesh's own shared library included, and a cross-package
change is publish then consume.

So the git dependency at a pinned commit is not a mechanism under consideration.
It is the shared library being consumed a way the record rules out, and the lock
naming a sibling directory is what that looks like when nobody publishes. Issue
053 is reclassified from a question about mechanism to a violation with a
direction.

What stays open is narrower and real: which software serves the private registry,
and that a fresh mesh has none when the first SDK is built — ADR 0014 assumes one
exists, and at genesis nothing has installed it.

Recorded after arguing at length for a bespoke content-addressed alternative,
against a decision that was already made and that I had not read.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 16:50:40 +02:00
jschoubben bc271d4de0 A worked guide: one module, four capabilities, four languages
What the SDK contains, answered by exclusion as much as by inclusion. It is the
protocol and nothing else — no configuration loader, because configuration
arrives as files the mesh wrote; no API clients, because a Plex client changes
when Plex changes and that has nothing to do with any other module; no storage,
HTTP or logging, because the language has those. The test for anything proposed
is ADR 0039's: does editing it recompile unrelated modules, and does it change
often. Both, and it stays out.

Then the worked module: events in TypeScript, tools in Go, a provisioner in Rust,
a scheduled job in Python. Four artifacts, four toolchains, four processes, one
module — and each part is an ordinary project in its language depending on the
mesh SDK the ordinary way, so a laptop resolves what a build resolves.

And publishing a package as a module capability, which makes the SDK unspecial:
it is simply the first module that published a library. A Plex client belongs to
the Plex module because that is the only thing that knows when Plex changed.

Three things left open rather than papered over: which registry (the catalogue
holds verdaccio and a forge usually serves one too, and nothing says which is
ours), who may publish (a credential that does not exist), and what a range means
in a mesh where everything else is pinned by digest — a mesh that can rebuild a
commit and get a different library is a real change, and should be decided rather
than arrived at.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 14:12:01 +02:00
jschoubben b637fe106f The module protocol, specified per capability
A floor every implementation needs and three capabilities independent of each
other, so an SDK can implement the floor and events and be a real thing rather
than an unfinished one.

Written as a specification, which means it says what is required rather than how
anything is arranged — and says plainly where it describes behaviour that is not
yet true. Three places it does:

x-causation-id and x-schema are specified and emitted by nothing; the Go side
writes four headers and the TypeScript side declares six. A module may serve
tools and may not call them, because a caller needs a reply queue its account may
not declare. And the two implementations disagree about what a grant carries — in
TypeScript consumer is the module, in Go it is the node and the module is From.
One word, two meanings, in two halves of one mesh.

Naming those in the specification rather than leaving them for conformance to
discover, because a specification that only described what already works would
have nothing to say about the things most likely to break.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 13:45:25 +02:00
jschoubben 033a5d384c Issue 053 — the SDK is pinned twice and the two disagree, and ADR 0074 accepted
The tool runtime's manifest names the SDK as a git dependency at a pinned commit;
its lock file names a sibling directory that exists on one workstation. It builds
only because the recipe runs npm install, which tolerates a lock disagreeing with
its manifest and re-resolves from the manifest — the one command that hides this.

It matters because a lock exists to make a build reproducible and this one
describes one machine, and because it is the first thing a fresh mesh builds: the
toolchain carries the SDK and everything with code of its own compiles inside it,
so a dependency resolved differently on the build machine than on a workstation
is a difference in every module the mesh will ever build.

And it is about to be copied. Each language's toolchain will carry that
language's SDK the same way, so the shape is worth settling before there are four
of them.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 13:44:02 +02:00
jschoubben 89302aa3e0 The protocol is split per capability, and an SDK implements it
Two corrections, both from the operator and both better than what was written.

An SDK is an implementation of the mesh's module protocol in one language, and
nothing more. The first draft defined it by the test it passes, which describes
how you check one rather than what one is — and leaves it sounding like a library
that helpers could accumulate in.

And the protocol is split per capability, which was missing entirely. A module
that only consumes events uses the event capability; one that serves tools uses
the tool capability; a provider uses provisioning. Nothing about consuming an
event requires knowing how a grant is answered, so an SDK need not implement all
of it to be real.

That has a precedent here: a host declares which resource kinds it can apply, and
a partial host is a real thing rather than a broken one (ADR 0005). An SDK
implementing the floor and events is exactly as legitimate, and a module written
against it is a module that does events.

Which changes what adding a language costs. A Rust SDK doing connection and
events is useful the day it exists, with tools and provisioning following when
something needs them — rather than a language being unsupported until it is
entirely supported, which is what makes adding one a project instead of a
contribution.

Conformance is therefore per capability too: a monolithic pass or fail would make
a partial implementation indistinguishable from a broken one, which is the
distinction the whole thing rests on.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 13:40:15 +02:00
jschoubben 3d693a8735 ADR 0074 — the wire is specified; an SDK is what passes the suite
ADR 0039 settles what belongs in an SDK. It does not say what happens when there
is more than one, and there already is: the contracts are expressed as Go types
in the control plane and host and as TypeScript types in the SDK, and nobody has
felt it because both live in one repository.

They already disagree. The provision's field is "resource" in one and "Provision"
in the other; "consumer" means the module in one and the node in the other; the
envelope declares six headers on one side and emits four on both — the missing
two being x-causation-id and x-schema, the second of which is exactly what a body
needs in order to change shape without silent misreads.

That class of failure does not announce itself. Two implementations disagreeing
about an envelope do not fail to compile — they ignore each other's messages, and
a mesh where a module stops reacting looks like a mesh where nothing happened.

So the decision is to specify the wire rather than share the types, because the
shapes are the easy half. What two implementations actually disagree about is
behaviour: queue naming and durability, which headers are required and what an
unknown one means, taking identity from the sealed credential rather than the
environment, dedup on an id only the emitter can make, pinning a fingerprint
rather than trusting an authority.

And the suite is executable rather than prose, because a specification nobody can
run is a document two implementations drift from while both believe they conform.
The two existing implementations are the first made to pass it — a suite only new
SDKs must satisfy would certify every future language against a disagreement that
is already here.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 13:26:57 +02:00
jschoubben 411f0680b8 Document the whole module surface, and keep it true with a test
A reference table goes stale the day somebody adds a field, so this one points at
mesh-catalog/modules/showcase — a module that uses all of it — and a test that
fails when it stops doing so. Read the module when the table disagrees with it.

Two rows in the coverage survey were stale because of this week's work: systemd
units were a file plus a service, which made every author write unit syntax and
is why "process" exists; and building from source was images only, where a
bundle now names a language and lets the mesh choose the toolchain.

And the two rows at the bottom of the resource table are the interesting ones.
"action" is refused to modules outright — the link may not carry a command, so a
module needing something done ships a program that reconciles. "service"
installs no unit by design, right for software shipping one and wrong for code
the mesh built, which has none until the mesh writes it.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 13:06:16 +02:00
jschoubben 8a7328c282 Correct the design: archives already work, compiling is what is missing
The first draft said the builder refused archives. It does not. An archive is
packed deterministically, hashed, published by digest, fetched by the machine and
unpacked — the whole path exists. Only the local builder used at genesis refuses
one, and deliberately: an archive is bytes that mean nothing until something
serves them, and at genesis nothing does.

What an archive cannot do is compile. Its source is a directory packed as it
stands, so shipping compiled output means compiling somewhere first, which means
a Dockerfile — the burden this document is about. The gap is not the artifact
kind. It is that no recipe both builds and packs.

Found by reading the builder rather than the manifest schema, which is where the
first draft's claim came from.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 01:54:43 +02:00
jschoubben 5d1e6d0b09 Design: building a module, and why the recipe cannot always be a Dockerfile
12-a-module-repository says what a module may build and where it goes. Nothing
said how a build is MODELLED, and the model is the problem: a recipe is implicit,
singular and always a Dockerfile; a toolchain is not modelled at all, arriving as
two build arguments the module hand-writes; a language is not a concept; and an
archive is declared in the manifest and refused by the builder.

The cost is measurable rather than theoretical. Adding a module with its own code
means repeating an incantation - two ARG bases, a specific working directory so
the SDK resolves upward, the compiler invoked by absolute path because the usual
symlink is resolved away when the base is assembled, a second stage, an env var
naming the entrypoints. Most of the catalogue is unconverted, and two conversions
done in one session were each wrong twice with a working example open.

So: recipe becomes explicit with three kinds, and toolchain becomes derived from
a declared language rather than written by every author. A Dockerfile stays, and
stops being compulsory - it is right for software needing a particular base and
wrong for "compile my module's code", which is the same operation every time.

The cost is stated before it is chosen: every language is permanent, and the
contracts are already expressed twice - Go structs and TypeScript types kept in
step by hand. A second language makes that drift. So language-neutral contracts
come first, or the drift gets worse while hiding.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 01:51:42 +02:00
jschoubben b163ed1fcc Issue 052 — the firewall closes the port the mesh runs on
The packet filter generates its rules from what modules declare they listen on.
The broker is not a module, so it declares nothing, so its port is not opened.
Every machine dials that port to enrol and to receive every declaration it is
ever sent.

Invisible where it is assembled and fatal on the next machine: a mesh of one
never dials its own broker across the network, so the ruleset looks right. The
first machine to join is refused at the packet filter during enrolment, several
steps from anything that reports it — and assigning the firewall before joining
machines is both the natural order and the one that breaks.

This is 051 in a second place. That issue says the substrate cannot be updated
because the mesh holds no record of it; the same absence means the firewall
cannot know it exists. Ssh is already a floor for the same reason — a machine
nobody can reach is a machine nobody can repair — and the broker may be the same
class of fact.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:59:33 +02:00
jschoubben f52895a646 Issue 051 — the mesh can update everything except what it depends on
The store and broker come from a bundle the installer writes once, with images
pinned in it, and nothing can change them afterwards: no build, no version to be
behind, no roll-out, and no way to report being out of date, because the mesh
holds no record of them as modules at all.

That is backwards. They are what everything else depends on, so their updates
matter most, and they are the only things with no mechanism to deliver one. A
mesh with a year-old broker reports itself entirely current.

The fix probably already exists: the control plane is carried, raised and then
adopted as an ordinary module pinned to what is running. Nothing in that pattern
is specific to the control plane. It would also remove a duplication visible on
any one-node mesh — the same postgres image running twice, because a store that
cannot be a module cannot provide a database to anything.

Recorded with the two hard parts stated rather than waved at: upgrading a store
the control plane is reading from, and upgrading a broker over the broker.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:44:10 +02:00
jschoubben 5d13f9c83d Issue 050 — the catalogue knows nothing built before it started
A mesh raised from bare metal built six modules and its catalogue reported three:
exactly those built after it began running. Missing were the shared base, the
store, and the catalogue itself.

The hole is never random. On a fresh mesh the modules built before the catalogue
are by necessity the ones it needed in order to exist, so the foundation is
always what is absent, on every mesh, at the moment the graph is first populated.

It breaks the question the catalogue is for: build edges hang off the base, so a
catalogue with no record of it answers 'what must be rebuilt' confidently and
wrongly. And nothing reports the gap, because a catalogue cannot know what it was
never told — this was found by comparing its answer against what the mesh had
just been watched doing.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:16:23 +02:00
jschoubben 0aa8f62e0c Issue 049 — a module serves tools and nothing may call them
A module's broker account is scoped to what it declares it emits and consumes. A
tool call needs a reply queue, which that scope does not cover and should not. So
the account is right, the request is reasonable, and no account exists that can
make it — asking a module its own question, from its own container, with its own
credential, is refused.

It matters because a module's tools are its operator-facing surface: the
catalogue serves the five questions it exists to answer and nothing can reach
them. It is also why a running module keeps being mistaken for a working one — a
test that cannot ask anything checks a container is up, and that substitution has
hidden two faults this week.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:08:40 +02:00
jschoubben ee70d5b451 Issue 048 — a stated rule about the registry is enforced by nothing
The registry says every machine pulls from it and opens its port to the mesh for
that reason. A machine that tries is refused by its own container runtime: the
registry serves plain HTTP and anything but loopback is treated as HTTPS.

It has never failed, and that is the finding. Every proof that a machine can
fetch a mesh-built artifact was a proof about the machine that built it, where
the reference was loopback. The bed that uses a routable address gets away with
it because the harness writes the runtime's configuration before the mesh exists.

Kept separate from 042, which they are easy to confuse: that one is about not
being allowed to pull, this one about not being able to whatever the credential
says. Fixing 042 alone leaves a machine with a valid account for a registry its
runtime will not talk to.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 21:01:32 +02:00
jschoubben 3121e47245 Merge pull request 'Issue 047 — and the half that runs the other way' (#42) from issue/047-the-other-half into main 2026-09-14 15:09:36 +02:00
jschoubben 261064ec63 Issue 047 — and the half that runs the other way
The report covered ports the firewall cannot close. The matching fault is ports
it does close and should not: rules come from what modules declare, a migrating
mesh knows about almost nothing, and anything listening on the host is dropped.

No module declares an ssh port and the generator has no allowance for one. The
session that loads the rules survives on conntrack until it drops, and then the
machine is reached from a rescue console.
2026-09-14 15:09:20 +02:00
jschoubben 3f35153dd5 Merge pull request 'Issue 047 — the firewall does not cover published container ports' (#41) from issue/the-firewall-does-not-filter-published-ports into main 2026-09-14 14:59:31 +02:00
jschoubben 027c41d125 Issue 047 — the firewall leaves published container ports open
Rehearsed on three lab machines before doing it on an anchor: loading the mesh's
rules refuses an ordinary host port and leaves a published container port
reachable, same prober, same second.

It is deliberate — no forward chain, because dropping there would stop every
container — but traffic to a published port never reaches the input chain, so
the firewall is silent about it. The anchor publishes 38 such ports today, all
filtered by the system being replaced, so the cutover would open every one of
them while reporting a firewall that is up.
2026-09-14 14:59:11 +02:00
jschoubben 9359542990 Merge pull request 'Issue 046 — an upstream image cannot be mirrored into the mesh's registry' (#40) from issue/mirroring-an-upstream-image into main 2026-09-14 13:27:59 +02:00
jschoubben 24d52a8380 Issue 046 — an upstream image cannot be mirrored into the mesh's registry
Found while migrating the first module. Five approaches were tried and observed
to fail, including resolving the index and naming the platform at both ends; a
speculative fix was written and reverted rather than shipped, because it did not
make the mirror work.

One module uses this today, which is why it went unnoticed — and it is the shape
most of a migration wants, because the services being moved are third-party.
2026-09-14 13:27:42 +02:00
jschoubben e773bf4c34 Merge pull request 'Installing ends with a builder, and the record says so' (#39) from feat/installing-ends-with-a-builder into main 2026-09-14 12:34:43 +02:00
jschoubben 256884e671 Installing ends with a builder, and the record says so
The design said how the builder arrives was unsettled and that nothing
installed it — the one gap stopping a fresh mesh from producing anything. Both
are now false. What is still true is narrower and worth keeping separate:
nothing asks a raised mesh for the rest of the catalogue, and no bed asserts
that it could.
2026-09-14 12:34:26 +02:00
jschoubben aa35ba1594 Merge pull request 'A module names its base, so the mesh can act on what it notices' (#38) from feat/a-module-names-its-base into main 2026-09-13 23:58:23 +02:00
jschoubben 2eb65a17a3 A module names its base, so the mesh can act on what it notices
The mesh could already say which modules a base change invalidated, and could
not do anything about it: each recipe named one particular copy of the base by
fingerprint, and rebuilding produced the old one. Worse, the copy each named
existed only inside a throwaway lab, so those three modules could not be built
anywhere at all — and the line each replaced had the same fault.
2026-09-13 23:58:06 +02:00
jschoubben 68d83c5b24 Merge pull request 'Two graphs, the builder's arrival, and two findings from building on a live mesh' (#37) from feat/two-graphs-and-the-build-chain into main 2026-09-13 11:17:48 +02:00
jschoubben 8e7fac8bf1 Record what genesis now does, and how each rule is checked
The builder's arrival was the one rule the document said nothing checked. It is
checked now, by both genesis beds — and so is the thing that distinguishes a
built control plane from a carried one, which every earlier assertion accepted
either way.
2026-09-13 06:09:31 +02:00
jschoubben fdc61054e3 Genesis carries a builder, and the document says so
The section saying the change was decided and had not happened now contradicted
the section below it. It also records the argument that failed, because a reader
will otherwise ask the same question and reach the same wrong answer.
2026-09-13 04:26:01 +02:00
jschoubben aabaaf2bb4 Rewrite 0073: the registry does not move, and the argument fails
Written an hour ago claiming a produced image must be published before anything
can fetch it, so the registry had to precede the control plane. The premise is
false: the machine that builds the image is the machine that runs it, and the
temporary control plane names a built image exactly as it names a carried one —
by the digest of its own configuration, which requires nothing to have served
it. Building changes where the bytes came from, not where they are.

Rewritten rather than superseded because nothing has been built on it and
nobody has read it: a record that contradicts itself is a draft, not a decision.
The argument is kept, because it was asked for and a negative answer is the
result.
2026-09-13 03:22:05 +02:00
jschoubben 9f5ac38662 Settle how the builder arrives, and what it publishes into
The two questions the design record named as the one gap stopping a fresh mesh
from producing anything. They cannot be answered apart: a builder with nowhere
to publish has made a file on a disk.

The registry's role did not change — the answer to 'must it precede the control
plane' did, because the control plane's image is now produced rather than
carried, and a produced image must be put somewhere before it can be fetched.
2026-09-13 03:20:07 +02:00
jschoubben 3983188b8a Issue 044 resolved — the mesh builds its own floor
Cloned alone, built by the mesh, every module moved onto it, and the base then
changed for real: all three went stale naming what moved. A comment-only change
correctly makes nothing stale, which found a second bug — staleness compared
commits where it should compare artifacts.
2026-09-13 02:51:08 +02:00
jschoubben 15ac7fd8dc The base has no circularity — the builder does not stand on it
Written as an open question; it has an answer, and leaving it open would have
made the fix look harder than it is.
2026-09-13 02:25:07 +02:00
jschoubben 3a9ed9d3fd Two findings from building the catalogue on a live mesh
The runtime every module compiles against cannot be built by the mesh, so the
one rule that would catch it moving can never fire. And a container keeps the
values it was created with, so two good applies can leave it running on neither.
2026-09-13 01:49:02 +02:00
jschoubben a4866ca899 Merge pull request 'Two graphs, and a build chain that orders itself' (#36) from feat/two-graphs-and-the-build-chain into main 2026-09-12 23:09:42 +02:00
jschoubben 58ad0742d8 Two graphs, and a build chain that orders itself
Corrects ADR 0070, written an hour earlier, which had the control plane consuming
the catalogue in order to compose a declaration. That was written before the two
graphs had been told apart and creates a dependency that need not exist: a
catalogue that is down would leave the control plane unable to compose the thing
that would repair it.

The catalogue links module-versions to each other and does not know nodes exist.
The control plane links module-versions to nodes, and holds capabilities and
claims. They meet only when something is installed, and everything between them
travels as events over the broker, on durable queues, so nothing is lost when a
receiver is away.

Build order is not computed anywhere. The builder never consults the graph and
builds what it is asked for; the catalogue asks for the next build after the
previous registration, so ordering holds by construction. Its rule is a condition
rather than a schedule — rebuild once everything a module was built against is
current — which covers a chain and a diamond alike.

Left open: whether an upgrade is applied or merely noticed, and whether a module
on several machines upgrades on all at once.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 23:09:23 +02:00
jschoubben 90f7bcb209 Merge pull request 'Genesis clones from a mesh, and checks what it got' (#35) from feat/where-genesis-gets-its-source into main 2026-09-12 22:20:28 +02:00
jschoubben 72b0810c53 Genesis clones from a mesh, and checks what it got
ADR 0070 has the init builder clone the source and does not say from where, and
ADR 0067 had rejected building at genesis partly because the forge runs on the
mesh being rebuilt. That objection binds only when those are the same mesh, which
is true exactly once.

So genesis clones from a mesh by name, and if that mesh is lost the name moves to
another that holds a copy — recovery is a name pointing elsewhere rather than a
backup being restored, and every installation adds somewhere it could point.

It names a commit and checks what it got, because the forge a mesh installs from
is the trust anchor for everything that mesh will run. On 2026-09-11 that forge
was running a cryptominer and tampering with git operations in flight; nothing was
altered, but a mesh installing during those hours could not have known that.

What relationship a mesh keeps afterwards is left open and named, so that whoever
writes the init builder does not settle it by accident: a snapshot and then
independence, or a continuing upstream for core modules.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 22:19:41 +02:00
jschoubben 03c73e7698 Merge pull request 'The catalogue owns the module graph, and genesis builds rather than carries' (#34) from feat/the-catalogue-owns-the-module-graph into main 2026-09-12 22:09:38 +02:00
jschoubben 64ae75a480 The catalogue owns the module graph, and genesis builds rather than carries
The graph had no owner: what modules are, what they require, what they claim and
what they are built against all sat in the control plane because that is where it
was written first. The control plane's own test says otherwise — anything a single
machine could answer alone is not its work, and what a module needs requires no
knowledge of any node.

So the catalogue becomes a core module beside the control plane and the builder,
owning the graph and serving tools over it. The control plane consumes it, which
is the opposite of what the tiers suggest and is therefore stated rather than
inferred.

That makes the catalogue a fourth thing that cannot arrive through the ordinary
path, so the claim written this morning that the list was closed at three is
corrected. All four are answered by one mechanism instead: the installer carries
an init builder and the core modules are built on the machine, so what is carried
is a builder rather than a result and nothing is left without a route.

Two things are open and named rather than assumed: where the init builder clones
from, given the forge normally runs on the mesh it would be rebuilding, and what
it publishes into, given the registry is installed later in the order today.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 22:03:05 +02:00
jschoubben 0fa4d14f54 Merge pull request 'A module is a repository and a path, and installing is described to its end' (#33) from feat/a-module-is-a-repository-and-a-path into main 2026-09-12 16:54:22 +02:00
jschoubben 837b5df2f7 Withdraw 043: the capability existed and the wrong verb was used
A build machine was refused the build queue, and this was raised as a gap in what
a manifest can express. It is not: `builder issue` creates exactly that account,
three lines from the code being read at the time.

Kept rather than deleted, for the one real thing in it — the wrong verb succeeds
and reports success, producing an account that authenticates and can do nothing,
so the failure surfaces a layer away as a permissions error that reads like a
missing feature.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 16:51:37 +02:00
jschoubben 9da45c68d6 Propose that the lab takes requests, one at a time, from its own copy
Raising a scenario occupies the machine and the person who started it, and
running in the background against a working copy is worse than waiting: a run
reads that copy as it goes, so editing while it runs yields a result describing a
state that never existed.

Proposed rather than accepted. The load-bearing part is the restriction — a
request names a bed and a commit and nothing else — because a request that could
say what to install and where would make the lab a second way of installing a
mesh, which is the arrangement that just cost a year of late-found faults.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 16:51:37 +02:00
jschoubben ddf104f8aa A module is a repository and a path within it
The design said a module's manifest sits at a repository's root, full stop, which
means one repository per module. Nothing that exists is shaped that way: the
catalogue holds sixty-seven modules one to a directory, no code repository has a
manifest at its root, and the system being replaced has always built a module
from a repository and a path.

So the builder could be asked to build nothing that exists — pointed at the
catalogue it finds no manifest, pointed at a module's source it finds none
either. Recorded as a decision because it moves the core modules' manifests
beside their source, and corrects the design that said otherwise.

Also corrects, in the same document, how the three things the build loop cannot
produce actually arrive. They were written as though all three were carried in.
Only the control plane is: the registry is pulled from the public internet, and
the builder has no route at all — which is now stated as the open one rather than
implied to be solved.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 16:51:37 +02:00
jschoubben 3d939b5c77 Describe how a mesh is raised, because only a test did
The one complete account of standing a mesh up was an integration test, and a
fixture is free to invent what it needs — which is how a registry that exists in
no production hid two faults for as long as the lab existed.

Written from what the installer does, not what it should do: genesis and joining
are separate moments, the lab runs the installer rather than describing
installing, and three things that are not true yet are named rather than glossed,
including one rule nothing checks.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 16:51:37 +02:00
jschoubben c567410687 Merge pull request 'Genesis is a pivot, public routing is name-agnostic, and five issues the fake registry was hiding' (#32) from design/bootstrap-is-a-pivot into main 2026-09-11 21:56:49 +02:00
jschoubben 26bec28c60 Issue 042 — nothing gives a node an account for a registry
Found by deleting the lab's registry and giving the machines a real path out:
public images fetched, the operator's own could not be fetched at all. There is
no provision for a registry credential, no manifest field, and no step in
enrolment that establishes one.

It applies to the mesh's own store too, which today asks for nothing — a
decision that has never been written down as one, and so reads as an absence.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 01:36:00 +02:00
jschoubben cde00e1d5f 0054 and 0055 join the topic their subject already had
Both carried 'model access', which is not one of the six the reading order
knows, so neither had a place in it. 0050 — the record they extend, on the same
subject — is 'what runs on it'.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 01:06:59 +02:00
jschoubben c29accd1be Issue 039 resolved — the operator's images are pinned, and why that is a stopgap
Nine references now name the digest their tag resolved to. What closed is the
immediate fault; the open questions stand, because a digest in a repository is
wrong the moment anybody rebuilds — which is the reason the design wants the
repository to name artifacts and the mesh to hold digests.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 01:06:35 +02:00
jschoubben 47dc3bc8e5 Accept ADR 0066 and ADR 0067
0066 was proven on a four-node bed before it was ratified: one node setting
moved an entire domain, a routed name resolved inside the mesh and was issued a
certificate by the internal authority, and TLS verified against that authority
with no override. 08-connectivity rests on it and could not while it was
proposed.

0067 records what deleting the lab's registry exposed — that pinning quietly
required a registry before the thing that lets a mesh have a registry could
start — and the pivot that resolves it.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 01:04:21 +02:00
jschoubben 423fde8534 Place the two new records in the reading order, and fix a link the merge moved
Both carried a topic outside the six the index knows, so neither had a place to
be read in — 0067 had none at all. Both are 'the tiers', beside 0036 (bootstrap
ends at a usable mesh) and 0007 (connectivity), which is what they extend.

0067 cited 0041 for tier 0's property; on this trunk 0041 is events, and the
record it meant is 0005. A citation that resolves to the wrong record reads as
corroboration, which is worse than a dead link.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:48:20 +02:00
jschoubben 56d0dcc4ff Merge branch 'feat/routing-is-name-agnostic' into design/bootstrap-is-a-pivot 2026-09-11 00:46:04 +02:00
jschoubben 84c95c53af Merge branch 'issue/021-provider-port-published-on-loopback' into design/bootstrap-is-a-pivot
# Conflicts:
#	03-DESIGN/01-to-be/04-lab-installation.md
2026-09-11 00:45:59 +02:00
jschoubben e4a5d9b1d0 Lab installation: the reachability check joins the assertions it belongs with
The doc's own rule is that the lab verifies capability by outcome, never by
reading a setting. A path out is exactly that kind of claim — a route and a
policy can both read correctly while nothing gets through — so it is asserted by
fetching something, in the table with the rest.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:11:39 +02:00
jschoubben 1ce1c3d4e0 Lab installation: a path out, and why the daemon having one is not enough
A container runtime on the same workstation sets the kernel's forwarding policy
to drop, and the virtualisation daemon's own accept rules do not override it —
both are consulted and a drop anywhere is the answer. The machines then get
addresses and resolve names, because the daemon's resolver is on the bridge, and
discard every packet to anything real.

The sharpest 'available is not adequate' yet: nothing is misconfigured, nothing
logs, and it presents as every image pull hanging. A workstation that runs
containers is the ordinary case, so it is a prerequisite — verified by reaching
something, never by reading a setting.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:11:22 +02:00
jschoubben 8f23d4b114 0067: the artifact store may require nothing, not merely build nothing
029 says a module providing the artifact store may not build artifacts. The
pivot shows that is the narrow case: it may not require anything the store is
needed to deliver. A route-label migration gave the registry a public name and
a route requirement, and at genesis nothing provides a route — nor can anything,
since the routing stack needs images and images need the store.

The same cycle through a door the existing wording did not cover, so the rule is
widened where the bootstrap decision states it, with a check that would catch the
next one where it is written.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:10:39 +02:00
jschoubben f9e129f734 Issue 040 — the only description of how a mesh is stood up is a test
There is no installer. The complete account of standing up a mesh is an
integration test in the lab, and a fixture may invent what it needs — this one
raised a registry no production has and rewrote every image reference through
it, concealing both 039 and the fact that a first node outside the lab had no
bootstrap path at all. The bed being green said nothing about whether a mesh
could be installed.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:34:59 +02:00
jschoubben e228355a52 Issue 039 — the lab's registry was pinning what the catalogue left unpinned
Nine images across seven modules name a tag, not a digest. ADR 0006 forbids it
and the host refuses it by name — and the refusal has never fired in a bed,
because the lab pushed every image into its own registry and rewrote every
reference to the digest it had just assigned. The harness was supplying the
property under test. Found by deleting the harness.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:20:10 +02:00
jschoubben cd1653bfe9 ADR 0067 — genesis is a pivot through a temporary control plane
The control plane's image is built from source and pushed nowhere, so it has no
manifest digest; a registry assigns those. Pinning therefore required a registry
before the thing that lets a mesh have a registry could start — a dependency the
rule created by accident. The lab hid it by raising a disposable registry no
production has.

Proposed, not accepted.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:19:17 +02:00
jschoubben 5852b35ab9 Renumber the routing record to 0066 — 0056 was already taken
0056 is 'the authority is the control plane, not a database', drafted on the
in-progress record chain this branch was cut from before those three records
landed. Two files would have collided at merge, which is the kind of thing that
is cheap now and confusing later. The code written against it still says 0056
and is corrected separately.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:17:20 +02:00
jschoubben a4440c9acf Issue 038 resolved — corrected root cause (same-node served-port mismatch, not loopback) + fix reference
The real defect is same-node consumers announced the declared port instead of
the assigned/published one; the loopback observation was a stale pre-0038 build.
Fixed in mesh-control fix/same-node-provider-announced-port (c147a26) with a
regression test.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 23:06:58 +02:00
jschoubben 9ffe9e97c9 Renumber to issue 038 — 021 is already taken (same-node credentials, closed) and issues run to 037 across branches
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 22:46:11 +02:00
jschoubben d16b5294f6 Issue 021: narrow to mesh-assigned ports (bare decl binds loopback, explicit host mapping binds 0.0.0.0)
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 22:35:06 +02:00
jschoubben 315bd7118a Issue 021 — a provider is announced at a name its port is not bound to
A from:mesh provider is announced (per 018's fix) at the node's private-network
name, but its port is published bound to loopback, so consumers dialing the
announced <node>.internal:port reach nothing. Diagnosed from the lab: DB
consumers that require the database at startup crash-loop; the provider is
healthy on loopback only.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 22:29:04 +02:00
jschoubben e2afb3e144 ADR 0056 + connectivity: public routing is name-agnostic, resolved in-mesh, internally certifiable
A route contribution carries a label; the node carries its public domain; the
mesh composes <label>.<public-domain> and holds no name map. A granted route is
published into internal resolution so anything in-mesh (notably an internal ACME
authority) can resolve and reach it. That authority certifies routed names by the
same path a public one would, differing only in issuer and trusted root.

Records the decision as proposed and amends connectivity SS2/SS3/SS5 plus its Open
list with the lab findings behind it.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 22:24:43 +02:00
jschoubben 450a4d8605 Merge pull request 'ADR 0055 — model access is answered by a licence or a node that hosts the model' (#31) from feat/adr-0055-node-model into main 2026-09-07 04:39:25 +02:00
jschoubben 9dfa5ba2ac ADR 0055 — model access is answered by a licence or a node that hosts the model
Extends ADR 0050: model-access, a provision, may be answered by a NODE that
hosts a model (Ollama/vLLM serving an OpenAI-compatible endpoint) as well as by
a licence record. A node-answer delivers an endpoint (base URL + model, and a
key only if the server wants one), rides the ordinary serves/binds path, and
uses no adapter; the resolver already prefers a local model over a licence. One
provision, two answers — the mesh's own model sits behind the same interface as
a vendor's. Names the mesh-scope-alongside-licence limitation and its fix.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 04:39:10 +02:00
jschoubben b7a33a1e26 Merge pull request 'ADR 0054 — model usage is a vendor-neutral record at two grains' (#29) from feat/adr-0054-model-usage into main 2026-09-06 23:39:11 +02:00
jschoubben ec2d73bc06 ADR 0054 — model usage is a vendor-neutral record at two grains (licence + session)
Closes the usage-tracking half ADR 0050 left open. One row shape (0050's), recorded at two
consumer grains: the holding module (licence-level, e.g. Anthropic utilization%) and the agent
session (per-session tokens/cost, since a session IS a consumer per ADR 0026). Produced by the
vendor adapter; read on a schedule (0053); recorded as events the audit-logger keeps (0041/0042)
plus a queryable usage context store (0008); usage is not a credential and is recorded in the
clear. Reuses the session, schedule, event, and store the mesh already has. Accepted per the
user's choice to build the full feature incl. per-session token/cost.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 23:38:54 +02:00
jschoubben a361c86c3c Merge pull request 'ADR 0053 — a scheduled step is a container run on a recurring schedule' (#28) from feat/adr-0053-scheduled-task into main 2026-09-06 13:53:11 +02:00
jschoubben 12a311ba4c ADR 0053 — a scheduled step is a container run on a recurring schedule
The recurring twin of run-once (0052): a schedule modifier on the container shape, reusing
its security bound (no new host action/shape, strictly less than an action) and reversing its
gating rule — a scheduled step runs after convergence, does not gate the apply, and a failed
run is logged, not fatal. Unblocks kometa's sync and pollers. Accepted per direction to build
the primitive now.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 13:52:55 +02:00
jschoubben 897626a9f6 Merge pull request 'ADR 0052 — a step that runs once before a container (issue 037 / run-once primitive)' (#27) from feat/adr-0052-run-once-lifecycle into main 2026-09-06 00:05:04 +02:00
jschoubben a3b0e68ec5 Accept ADR 0052 — a step that runs once before a container
The run-once lifecycle primitive: run-once:true on the existing container shape,
gated by declaration order + exit 0, idempotent by digest. No new host shape, no
arbitrary command — strictly less powerful than an action. Resolves issue 037.
Verified sound and faithful (ADR 0005/0047); status proposed->accepted.
2026-09-06 00:04:21 +02:00
jschoubben 2405d72fb0 ADR 0052 (proposed) — an init step is a container run once to completion
A module can declare state but not a step that runs. mosquitto must seed its
dynsec admin into dynamic-security.json before the broker starts, or the plugin
aborts; the database providers need the same for first-boot migrations and
health gates (04-ISSUES/037). The old event-hook engine that did this was
powerful and flaky; this is the narrowest sound mechanism instead.

A run-once step is an ordinary container marked `run-once: true`: the host runs
it to completion, requires exit 0, and gates the apply on it — so what the
declaration places after it (the broker) starts only once it has finished.
Gating is by declaration order, not a resolved dependency (ADR 0005); the
completion marker is the recorded declaration digest (ADR 0018), so a re-apply
does not re-run it unless the declaration changed. No new host shape and no
arbitrary host command: strictly less powerful than an `action`.

Points 04-ISSUES/037 fixed-by/amended-design at the record; index regenerated;
records.py and index.py pass.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 23:52:26 +02:00
jschoubben c41deb55d2 Merge pull request 'ADR 0051 — shared data is the operator's; a module is granted access (fixes issue 036)' (#26) from feat/adr-0051-shared-data-operator-owned into main 2026-09-05 22:29:44 +02:00
jschoubben d5e12c82aa Accept ADR 0051 — shared data is the operator's; a module is granted access
Your decision, ratified: the media library (and shared/pre-existing data) is
operator-owned and external; a module declares access, not ownership; the host
mounts but owns nothing (no create/chown/reconcile/remove); an absent accessed
path is refused clearly; several accessors co-resolve. Status proposed->accepted;
index regenerated (records + index checks pass). Implementation lives on the code
branches (mesh-control/catalog/host), held for merge after the convergence fix.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 22:29:24 +02:00
jschoubben 220e3c72f5 ADR 0051 (proposed) — shared data is the operator's
Resolves 04-ISSUES/036: eight media modules each declared the shared
library and download directories as their own resources, and the
resolver's duplicate-owner refusal — right in general — would refuse the
stack's only sensible assignment the first time two landed on one node.

The decision, from the operator: shared, pre-existing data is
operator-owned and external. The mesh does not create, chown, reconcile
or remove it. A module declares it needs access to such a path (read or
read-write); the host mounts it and owns nothing. Several modules
accessing one path is normal — the duplicate-path refusal is about
ownership, not use. An accessed path absent at apply is refused clearly,
not created. Extends ADR 0030: the third case the host had no word for,
what it neither made nor configured and must not touch.

Point 036's fixed-by/amended-design at the record; mark it located in
mesh-control, mesh-catalog and mesh-host. Regenerate the decision index.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 22:22:30 +02:00
jschoubben 4430cc1748 Merge pull request 'ADR 0050 (proposed) — model access is vendor-agnostic [awaiting ratification]' (#25) from feat/adr-0050-vendor-agnostic-model-access into main 2026-09-05 21:42:28 +02:00
jschoubben 860d512e91 Accept ADR 0050 — model access is vendor-agnostic
Verified and ratified: model-access stays one vendor-blind provision; per-vendor
adapter keyed by licence.vendor (mirrors public-dns registrar providers); the
sealing-vs-central-rotation carve-out bounded to refreshable-grant vendors /
refresh token / manager node only. Status proposed -> accepted; index regenerated
(records + index checks pass); design doc note updated.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 21:42:20 +02:00
jschoubben 3a9b47918d ADR 0050 (proposed) — model access is vendor-agnostic; amend 14-model-access
Turn the completed vendor-agnostic analysis into HQ design. The model-access
provision stays one vendor-blind interface (extends 0024/0027); the
vendor-specific lifecycle moves into a per-vendor adapter keyed by the licence's
`vendor` field, mirroring registrar-scoped public-dns providers (0044), named at
the consumer's real coupling per 0040.

The crux is the sealing-vs-central-rotation carve-out: for refreshable-grant
vendors only, the manager node holds the refresh token encrypted at rest (a
bounded, declared exception), access tokens sealed per holder, refresh stripped
on delivery. Static-key vendors keep full sealing.

Amend 03-DESIGN/01-to-be/14-model-access.md with the adapter generalisation as a
proposed section (prose + diagram, no code); regenerate the decision index.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 13:43:01 +02:00
jschoubben 3b0aff7585 Merge pull request 'Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work' (#24) from reconcile-init-into-main into main 2026-09-05 12:27:11 +02:00
jschoubben 495db6dc89 Restore initialization's issue 003 — the status flip was based on a wrong premise
Init's 003 was already resolved with its own consolidated attribution
(03-DESIGN/01-to-be/08-connectivity.md); the re-homing overwrote it with the
session's ADR-0045 firewall attribution. Init is canonical and the firewall
decision is recorded in the ported ADR 0045 regardless, so 003 is restored
untouched. No initialization record is modified by this reconciliation.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 12:26:23 +02:00
jschoubben e269f9a185 Re-home this session's new ADRs (0039-0049) and issues (032-037) onto the consolidated scheme; flip issue 003; port repos.md sdk line + feature-branches playbook (07); regenerate index
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 12:24:07 +02:00
jschoubben 546caa31c4 Merge initialization into main — adopt its consolidated structure as canonical
The real work lived on initialization (consolidated decisions 0001-0038, the
fuller issue set 001-031, the control-plane/substrate/node-lifecycle/delivery
design, research 011/012, the checks tooling). main had diverged onto a stale
base and only carried this session's genuinely-new work. This merge makes
initialization's tree canonical on main; this session's 11 new ADRs and 6 new
issues are re-homed on top in the following commits. initialization is recorded
as a parent so its history is preserved in main's ancestry.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 11:59:01 +02:00
jschoubben db8be15a15 Merge pull request 'Issue 013 — a module cannot run its own code at a lifecycle phase' (#23) from feat/hq-issue-013 into main 2026-09-05 11:16:27 +02:00
jschoubben 9c8c4a3ee5 Issue 013 — a module cannot run its own code at a lifecycle phase
Surfaced converting the catalogue: a module can declare things that exist
(dir/file/network/container) but not a step that runs at a point in its
lifecycle. mosquitto's dynsec admin client must be seeded before first start;
the DB providers have nowhere for a migration or health-gate; it is the timing
face of issue 011. Framed as a missing module capability, not a defect. Records
the prior-art event hooks and their real warning — powerful but complex and
flaky — so the resolution avoids rebuilding that. Ends in open questions
(run-once resource vs general lifecycle hook, where the code runs, idempotency).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 11:15:53 +02:00
jschoubben 600a57aafb Merge pull request 'Issues 011 and 012 — what the manifest review could not fix' (#12) from issues/manifest-review into main 2026-09-05 03:52:32 +02:00
jschoubben 2c8616b650 Renumber manifest-review issues 008/009 -> 011/012
Numbers 008 and 009 were taken on main after this branch was opened
(008-provider-runtime-has-no-seal-key, 009-runtime-config-change-does-not-restart,
both merged). Renumber the seed-file-wipe and shared-directory issues to the next
free numbers so merging records two more issues rather than duplicating two.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 03:52:02 +02:00
jschoubben 69fa79c585 Merge pull request 'Playbook 06 — one feature, one branch, one MR per repo' (#22) from feat/feature-branch-workflow into main 2026-09-05 03:47:25 +02:00
jschoubben cabe873503 Playbook 06 — one feature, one branch, one MR per repo
Records the branching-and-merging workflow for code changes across the mesh
repos, written against a failure it names: branches and MRs opened per unit of
thought, treated as done when opened not merged, and named differently per repo,
so they pile up unmerged — one session left sixteen to consolidate by hand. The
rule is one feat/<slug> shared across every repo a feature touches, isolated in
.work/<slug>/<repo> worktrees off main, pushed and opened as one MR per repo only
when the whole feature is done, then merged promptly. Adds the ground-rule
pointer in AGENTS.md and the row in the process overview.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 03:40:39 +02:00
jschoubben 36936dc7a3 Merge pull request 'Issue 009 (fixed) + Issue 008 (resolved via ADR 0053) — module-runtime config & provider contract' (#21) from worktree-issue-provider-seal-key into main 2026-09-05 03:02:34 +02:00
jschoubben bfa08742f8 Consolidate hq: ADRs 0044-0052 merged in, and 0017/0049-0052 accepted
Brings the independent ADR branches (0044-0052) onto one branch so hq lands as a
single MR, and ratifies the five that were still proposed — 0017, and 0049-0052,
which are implemented and green in the lab. With 0053/0054 already accepted here, the
whole ADR chain 0044-0054 is accepted on this branch.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 03:01:06 +02:00
jschoubben b5a74178c3 Merge remote-tracking branches 'origin/design/adr-0044-sdk-boundary', 'origin/design/adr-0045-what-a-module-is', 'origin/design/adr-0046-events', 'origin/design/adr-0047-event-wire-shape', 'origin/worktree-adr-0048-module-broker-account', 'origin/worktree-adr-reachability-dns-firewall', 'origin/worktree-adr-config-is-the-assignments' and 'origin/worktree-adr-module-runtime' into worktree-issue-provider-seal-key 2026-09-05 03:00:24 +02:00
jschoubben 1a83ed28fa Merge pull request '011: measure the graph against every facet a module carries' (#11) from research/011-facet-coverage into main 2026-09-05 02:53:20 +02:00
jschoubben 984194e765 Accept ADR 0054 (slug) and resolve issue 010
ADR 0054 accepted with option E (a declared slug). Issue 010 resolved: the login fits
via the slug, and the minted secret shrinks to 40 chars for S3's secret-key limit —
both halves of an S3 credential now fit the tightest backend.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 02:52:03 +02:00
jschoubben af5c939d16 ADR 0054: add option E (a declared slug) and recommend it over A+B
Implementing A (bound the identity at 20) revealed the readable budget is node+module
<= 14 chars — so tight that the catalogue's own test names (workstation+keycloak, 25)
compact to an opaque hash. B's fallback would fire for the common case, not the rare
overflow, inverting A+B into mostly-opaque identities. Option E — an optional short
slug a module/node declares, preferred over the cleaned name — is the escape hatch B
wanted to be without the opacity: legible because a person chose it, and it makes an
early refusal palatable (refuse on the slug field, not the machine's name). B dropped;
E recommended over a bound of 20, composing with C later if needed.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 02:30:11 +02:00
jschoubben 89b65dd4c0 ADR 0054 (proposed) — a consumer's identity is bounded by the tightest backend
Sketches the options for issue 010: the mesh's identityLimit (63, postgres's) is not
the shortest among the backends the derived name reaches — S3's is 20 — so CheckIdentity
lets an over-long access key through and minio fails at provision time. Options: bound
by the true minimum and refuse at assignment (recommended, with a compact fallback held
in reserve), per-interface bounds, or a provider-generated identity (rejected — breaks
"the mesh says the identity once"). Links issue 010 to it.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 02:17:46 +02:00
jschoubben 75887d0bc5 Issue 010 — the mesh's derived login does not fit every backend's identity rules
Found doing the per-backend provider e2e: redis and postgres accept the mesh's `as`
(mesh_<node>_<module>) verbatim, but minio's S3 access key is capped at 20 chars and
`as` is 22, so the provisioner cannot create the service account. `as` is doing two
jobs — a stable identity the two ends agree on, and a literal identifier a backend
must accept — and those are not always the same string.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 01:58:11 +02:00
jschoubben e5f4af8cf2 Accept ADR 0053 and resolve issue 008 — provider contract implemented and proven
ADR 0053 accepted; adds the scope boundary the umami rework surfaced (credential
provisions vs data provisions — analytics' generated siteId return is left to a
separate decision) and records the lab proof. Issue 008 marked resolved: the sdk
harness and the four adapters are reworked, the symmetric seal removed, and
provider-uses-mesh-credential is green.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 00:27:56 +02:00
jschoubben e27f8b1b7d ADR 0053 (proposed) — a provider creates the credential the mesh minted, and seals nothing
Issue 008's trace confirmed the premise in control-plane code: the mesh already
mints one password per consumer/provider pair and delivers the provider its copy
(SecretFor/SecretsFrom/grantsFor -> Grant.Sealed; the receives contribution carries
As + Secret). The provisioner's symmetric seal is an orphaned, contradictory second
model. ADR 0053 corrects the provider contract in one place (the sdk harness):
providers create the resource with the mesh-supplied login and password and drop
seal/key/return entirely. Reframe 008 as contract-first (every provider, not four).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 00:06:50 +02:00
jschoubben 332c334767 Issue 008 — sharpen: the provisioner's seal model is orphaned, not just undelivered
A cross-repo trace showed nothing writes the provisioner's grant-request files,
nothing reads its sealed credentials, and no consumer unseals — while the mesh
already mints and delivers provider/consumer credentials asymmetrically with no
shared key. The fix is to drop the symmetric seal and have providers consume the
mesh-minted password, a breaking provider-contract change that wants an ADR.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 23:40:20 +02:00
jschoubben 47b0ca5c90 Issue 009 — settings change does not restart a container-hosted runtime
Found rolling the runtime out to the catalogue: a module's runtime reads its
settings-merged config file once at start, but a container is only recreated on a
spec change, and file content is not part of the spec. So updating settings
re-renders the file and nothing re-reads it — ADR 0051's "on the fly" holds only
for config set before first start. Services have restart-on; containers do not.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 23:06:48 +02:00
jschoubben d1aa4254c9 Issue 008 — a provider runtime has no seal key the mesh can deliver
Found building the module-runtime vertical slice: a provider's provisioner
requires a seal key it has no way to receive, and the consumer no way to obtain
the matching one. The runtime cannot come up as delivered.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 22:23:24 +02:00
jschoubben f63eca13b3 ADR 0052 — a module runs its code as its own process, with its own account
The runtime-model gap the review found. A module with tools or events runs one
container — the tool runtime carrying its code — holding the one scoped account
ADR 0048 gave it. A node-wide runtime can't: it would hold the union of every
module's permissions, the isolation 0048 draws. So per-module: one module, one
process, one account. Tools served per key (serve.<tool>) so a caller names a
tool and only its module answers (superseding a shared tools.invoke); events in
the same process under the same account; the runtime image is the tool runtime
plus the module's code (the audit-logger's shape, made the rule). A plain
service module runs no such process. A provider's provisioner is a runtime too —
which is why a provisioner that emits must carry a broker credential or not emit.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 21:24:46 +02:00
jschoubben e898a3ec44 ADR 0051 — a module's configuration is its assignment's, not its manifest
A module is assigned to a node (there is no mesh assignment; 'mesh' is a scope).
The manifest is what the module IS, plus defaults; the configurable values are
settings, carried by the assignment — per-node or mesh-wide, applied at
resolution, changeable live (what a meshboard edits). Extends settings from a
config file's content to the manifest fields marked settable: foremost
listens.from (postgres from:mesh by default, from:anywhere per node — the
firewall follows), and a provider's own config (a registrar's zone/domain/
ingress). Static config in a manifest is config in the wrong place: it cannot
vary per node and cannot change without a rebuild.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 21:05:55 +02:00
jschoubben 9eb5576682 04-ISSUES/003 — fixed: the firewall scope is enforced now
The manifest refuses unknown keys (DisallowUnknownFields), 'from' is the field
that scopes a port and it is rendered to nftables (AsNftables), and the firewall
module applies the rule set. The chain from a declared scope to a dropped packet
is closed. Amended-design: ADR 0050.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 20:56:44 +02:00
jschoubben 5118258ee3 ADR 0049 and 0050 — public DNS and the firewall, two reachability decisions
0049: a public name is provisioned like any capability — a module requires
public-dns and contributes its host; a neutral interface answered by
registrar-scoped providers (cloudflare-dns, route53-dns) that create/remove
the record pointing the name at the mesh's public ingress. Pairs with route
(the proxy) and a public cert (the proxy's ACME).

0050: answers the firewall question. The firewall is NOT a provider like the
proxy — it is a machine's own filter, derived by the host as the sum of what
its modules declare they listen on, with 'from' the whole of public-vs-internal.
Enforced both ways, unknown keys refused — closing 04-ISSUES/003. A public
service is exposed through the proxy (listens from:mesh + requires route), not
by opening its own port.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 20:45:58 +02:00
jschoubben cfa3ad1930 ADR 0048 — ratified: status accepted 2026-09-04 01:22:47 +02:00
jschoubben b622de4fe6 ADR 0048 — a module's broker account is scoped by its emits and consumes
Events (0046) and their wire (0047) left open how a module reaches the
broker. The code has no generic module broker-account: only node and
builder scopes exist, so emits/consumes are enforced by nothing — a
manifest declaring a scope the broker does not draw (04-ISSUES/003).

Decides: on assign, a module gets a broker account whose permissions ARE
the manifest — read on mesh.events + its own queue bound to consumes;
write to mesh.events under module.<self>.* only; nothing else. Consuming
'#' is a deliberate, auditable grant. The account is what makes the
declaration a rule the broker enforces, not a comment.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 01:19:15 +02:00
jschoubben 820ce8fc8c ADR 0047 — the shape of an event on the wire
The wire contract ADR 0046 left open: two topic exchanges (mesh.events,
mesh.rpc, kept apart so # is a clean audit); the routing key as the event
type namespaced by origin (module.*, mesh.*, node.*); metadata in AMQP
headers (required x-event-id/x-source/x-node/x-time/content-type; optional
x-causation-id/x-schema; unknown x- headers ignored) with the body only the
payload; persistent messages; per-consumer durable dead-lettered queues
with prefetch; at-least-once with idempotent consumers (no false exactly-
once). The precedent is ADR 0043 for declarations.

Supersedes the sdk's first cut (metadata in body -> headers); that and the
queue config are code to align in mesh-sdk and mesh-tools.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-03 23:47:58 +02:00
jschoubben 4bb6dd26ed ADR 0046 — events are a relationship, provisioning's lighter sibling
A module emits and consumes events, both declared (emits/consumes),
parallel to provides/requires. Events are 1:many, broadcast, credential-
free — no provisioner, just the broker's topic routing — so most inter-
module reaction should be an event, not a provision. Every event carries
source/node/time so it is auditable; the audit logger is just a module
consuming '#', no privilege. A consumes for an event nothing emits is a
dangling edge and refused, like requires. One per-node runtime serves
tools, provisioning and events alike.

Extends ADR 0045; builds on ADR 0001 (the broker) and 0044 (emit/on are
stable sdk surface; the binding and runtime are not).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-03 23:42:25 +02:00
jschoubben b2cb481681 ADR 0045 — what a module is
A module is one self-contained piece of software the mesh installs and
manages; the software is its identity, and capabilities/seats/provisions
are the relationships between modules, not what a module is. Records the
three relationships (shared seat, exclusive seat, provide/require), that
interfaces are mesh-owned and providers adapt to them, and the naming
rule: draw the interface at the consumer's real coupling — neutral where
the coupling is thin (analytics), protocol-scoped where the consumer
speaks a protocol (postgres/mssql/mongodb), never false genericity.

Supersedes 0017 (domain grouping — wrong axis), refines 0002, generalises
0027's protocol-not-product rule.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-03 22:50:52 +02:00
jschoubben 50d61398b6 ADR 0044 — what the SDK holds, and what it refuses
Supersedes ADR 0030's 'types, not behaviour' line for mesh-sdk. The
boundary is change-frequency, not kind: the SDK holds the stable spine
(tool-serving harness, messaging/event framework, contracts, core
primitives) and refuses per-module clients, per-module tool code, and
anything volatile — because those are what turned hal/sdk into constant
maintenance and made every edit rebuild every module.

States the rule (frequent AND cascading is the disease), why the root
cause was intra-module feature-sharing leaking into inter-module
coupling, and where per-module shared code lives instead (in the
module — a shared file, or a module-local sdk for the few large ones).
Updates repos.md's canonical mesh-sdk description to match; leaves 0030
untouched (immutable).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-03 21:41:08 +02:00
jschoubben 7b4a664c50 Issues 008 and 009 — what the manifest review could not fix
Both from the 2026-09-02 review of the catalogue examples, and both
design gaps rather than defects in a file: a seed file the host
reconciles back to empty over the grants that grew in it, and a module
stack refused co-assignment because six manifests each own the
directories they exist to share. Fixing either in place would have
been picking an answer the records do not yet hold.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-02 01:07:05 +02:00
jschoubben 9638aa02c5 031 — a machine becomes each thing it was told, in turn
Declarations queue and the host applies all of them, oldest first, so a
machine pushed five things in a minute spends five applies becoming the
last one. Correct every step — each declaration is the whole machine —
and wasted in all but the final step.

Only visible since a report names its declaration: the reports arriving
were about ever-older ones, while timestamp comparisons used to happen
to pass. The fix is consumption order, not the queue; the open question
is what a superseded declaration's report should say, because silence
reads as disobedience and "applied" would be a lie.
2026-09-02 01:06:12 +02:00
jschoubben b245f5e254 011: measure the graph against every facet a module carries
The five declarations cover relations; a module is more than its
relations. Add the facet-by-facet coverage table so the effort cannot
conclude while tools, verification and contributions are unplaced —
and weigh each candidate gap rather than adopting it: contributions
probably dissolve into declared resources, mandatory verifiers risk
trivial ones, and the tool surface is the one facet with no home.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-01 23:14:03 +02:00
jschoubben 6bdd3f2bfa 030 — asking what a machine should be re-signed its certificate
Composing a declaration signed the machine's certificate anew each
time, and a signing carries a fresh random serial — so what the mesh
would send differed from what it had sent by one byte, for ever, and
every machine carrying a certificate stood eternally waiting.

Third find of the same rule: issued once and kept. The port had it, the
secret had it, the certificate composed fresh on every asking — and the
keeping column had existed since the serving-key migration, written by
nothing, the same shape ReleasePorts was found in.

Found by keeping the scenario standing and diffing two plans seconds
apart: one line, where four theories had none.
2026-09-01 22:28:26 +02:00
jschoubben 6385d52bc0 029 fixed — the store is added, not built, and the cycle is refused
The registry module names its image by digest, the way the bundle names
the three a first node starts from, and goes in as a manifest. The lab
now walks that path and passes.

A manifest that provides the store and also builds artifacts is refused
at parse. The lab had been passing only because its Docker Hub stand-in
quietly received the push — a prop covering for the thing under test.
2026-09-01 21:43:42 +02:00
jschoubben 013dbce4a9 029 — the artifact store cannot be delivered by the artifact store
Installing the module that provides `artifact-store` requires something
that provides `artifact-store`: its image is mirrored in, mirroring
publishes to the store, and the builder refuses to run without one.

Never seen, because the lab always has a registry standing before the
mesh asks for one, and so does any mesh built on a machine that already
had one.

What it blocks is larger than a registry. The store holds images and
packed archives both — it is the module catalogue in artefact form — so
until it exists a mesh can run only what its bundle already carries.

The substrate record already answers it. It asks of each candidate
whether it can grant itself the thing it provides: the store cannot
create its own database, the broker cannot create its own virtual host,
and the registry cannot grant itself a repository. So the registry
module names its image and is never built.

With a limit worth saying out loud rather than discovering: a module
providing the store may not build artifacts of its own, a UI or a tool
server included, because there is nowhere to put them until it runs.
Such a registry is two modules.
2026-09-01 21:15:05 +02:00
jschoubben e75683c01f 026 reopened — the rule is right about data, wrong about facilities
Enforcing "declare what you mount" refuses the builder, which mounts the
container runtime's socket. That socket is not the builder's data: it
exists already, the machine owns it, and declaring it as one of the
module's directories would be a lie the host would act on.

Two kinds of mount are spelled identically today — the directory my data
lives in, and a machine facility I was granted. Until a manifest can say
which, the fourteen declared mounts are right by coincidence, which is
what this issue was opened about.

Recorded rather than decided: separating them is new vocabulary, and
inventing it to turn a check green is how a mechanism nobody chose ends
up load-bearing.
2026-09-01 19:45:55 +02:00
jschoubben bd8f09d647 026 fixed — and the coincidence turned into a rule
The fourteen undeclared mounts are declared. More to the point, a
manifest that does not declare one is now refused: they were right by
coincidence, and a checklist nothing enforces is a checklist that is
true until the next commit.

Refused in the control plane, because the machine cannot tell the
difference — asked to mount a path that does not exist, it makes the
directory, which is a thing it is perfectly able to do.
2026-09-01 19:36:15 +02:00
jschoubben 0bcfeb4e80 028 fixed — the mesh assigns the port, and knows what it cannot move
A module now says its port once, in `listens`, and the container's
mapping, the rule set and what a consumer is told are all derived from
one assignment. The three hand-written copies that agreed only because
one person wrote them are gone.

The half that made this an issue rather than an inconvenience was that
the substrate is not a module: nothing in the mesh had heard of its own
store, so it handed a database module the port the store already had.
The machine now says what it carries, and the mesh assigns around it.

Ports the protocol fixes became claims, which needed no new mechanism —
the mesh already had one for what is singular on a machine.

Left open, and unchanged by any of this: whether a module should publish
to the machine at all. Assignment makes publishing safe without making
it necessary.
2026-09-01 18:32:58 +02:00
jschoubben c2a37ab7f8 38 — the mesh assigns the port, and a module does not care
Jochen's call, and the right one: a module cannot choose a port well,
because it is written once and assigned anywhere. Any number it picks is
a guess about a machine it has never seen, and two modules guessing the
same number is not a mistake either of them made.

Writing it up turned up something the issue had missed. The same number
appears three times in every module — the rule set, what a consumer is
told, and what the runtime publishes — and nothing checks that they
agree. They agree today because one person wrote all three. A module
whose `serves` said one thing and whose container published another
would resolve, compose, apply, and hand every consumer a port that
answers nothing.

So the decision is one source with the other two derived, and an
assignment made once and kept, as a credential is.

The part that needed thought is ports that cannot move — mail on 25,
submission on 587. Those become claims, which is what the mesh already
has for what is singular on a machine. Two modules wanting 25 is the
same shape as two wanting the seat, and gets refused by name at
assignment rather than by a container runtime at apply. That makes this
mostly a matter of pointing an existing mechanism at ports.

Left open: whether a module should publish to the machine at all.
Assignment makes publishing safe without making it necessary.
2026-09-01 17:42:27 +02:00
jschoubben 6330abce5f 028 — two things want one port, and nothing says so until the machine
Found by fixing 027 and pushing again. The declaration is now accepted
and the database container still cannot start: the mesh's own store
holds 5432 on that machine, and the module publishes 5432.

Nothing catches it because the substrate is not a module. It arrives
from the bundle before there is a mesh to ask, so the control plane has
never heard of the store and does not know it holds a port. Resolution
can compare modules with each other and cannot compare one against what
the mesh is built on.

Nor does it compare modules with each other. A port is exclusive on a
machine in exactly the way a claim is, and the mesh has a mechanism for
that which ports do not use.

It has been met before: the end-to-end test that exercises a real
database publishes 5433 rather than 5432, inline, with nothing saying
why. That is how a constraint becomes folklore.

The open question is bigger than the bug. Whether a module should
publish to the machine at all decides how a consumer reaches it, and
changes what `serves` means.
2026-09-01 17:21:31 +02:00
jschoubben 03266fd4a2 027 — a container cannot follow a file, and a rotated credential is the case
Found by the forge failing to start. I had put `restart-on` on nine
containers so they would pick up a rotated credential; it belongs to a
service, and the host refused the whole declaration.

Removing it fixes the modules and leaves the reason I reached for it.
The mechanism is written against exactly this, in the host's own words:
a running service does not re-read its configuration, so replace the
file, find it already running, do nothing, and the machine keeps
behaving as before while every check passes. Every word of that applies
to a container, and nearly everything the mesh runs is one.

The cost is concrete. Rotation replaces the file and tells the provider
to accept the new credential. A provider reconciles, so it takes it. A
consumer is usually a container, so it does not — and the two ends hold
different passwords, which is the fault ADR 0001 records costing two
days. The test that proves rotation works uses a consumer that reads the
file on each attempt, so it does not meet this.

Two things a fix has to keep: it stays declared state rather than a
command, because the link may not carry an action; and where an env-file
changed, the honest verb is recreate rather than restart, because a
container's environment is fixed at creation.
2026-09-01 17:19:43 +02:00
jschoubben 3f00d3b413 026 filed; 025 corrected; the image store is a module
Three corrections, two of them to things I wrote today.

026 is the serious one. Four modules mounted fourteen host paths nothing
declared — the mail spool, the databases, the object store's data. The
runtime creates those as root, so owner and mode go unapplied, and the
rule that keeps a directory holding data the mesh did not put there is
written in terms of declared directories. It reached the configuration
and missed the data. The cause was carrying compose files across: a
container shape that can express one gets filled in like one.

025 claimed nothing turns a tag into a digest. That is false, and the
answer was designed and built before I wrote it. A module names an
artifact, not an image, and `kind: upstream` mirrors somebody else's
image into the mesh's own registry, pinned by the digest it lands with.
The two-document split the issue described as the shape of a fix is the
design. Pinning twelve images by hand was treating the symptom, and left
them pointing at a public registry rather than the mesh's.

And the image store was written up as something the mesh does. It is an
ordinary module — considered for the substrate and removed, because the
test is whether the control plane needs it before its first instruction,
not whether it can grant itself one. So somebody's own registry is the
same module as the mesh's.
2026-09-01 16:10:13 +02:00
jschoubben befbb1d978 The five modules could not have run, and now the forge does
Recorded against phase 3, because the phase note said running them
needed images stocked and provisioners built — and missed that not one
of them named an image that exists. Sixty-four zeros where a digest
belongs, eighteen times, parsing and resolving perfectly.

The forge now runs: on a database another module provides, with a
password it did not choose and a connection string it could not have
written itself. First of these descriptions to be started rather than
planned, and it exercises the whole of the credential work.

What remains is the mechanism rather than the data. Nothing turns a tag
into a digest as part of the mesh's own work, so it was done by hand —
which is what the issue says a person should not be asked to do. Asking
a registry takes a second and pulls nothing, so the main argument for
leaving it undone is gone.
2026-09-01 15:40:52 +02:00
jschoubben 2bbdc52440 025 — half done: nothing unrunnable reaches a machine now
The refusal landed, and the examples pin images that exist. What is
still missing is the part that makes it unnecessary: nothing in the mesh
turns a tag into a digest, so it was done by hand — which is precisely
what the issue says a person should not be asked to do.

Recorded because the resolution mechanism turned out to be trivial:
asking a registry what a tag points at takes about a second and pulls
nothing. That removes the main argument for leaving this open.

Also records the two faults that fell out of pinning for real. The mail
system named seven repositories that do not exist, because it publishes
to a different registry than the manifest assumed, and one of the seven
had been renamed upstream. Nothing checking only the shape of a
reference could have found either.
2026-09-01 15:13:52 +02:00
jschoubben fc7717f6b1 021 closed — the record was left open after the fix landed
The code and its tests went in hours before the record was touched, so
an issue that read `located` had been fixed all along. That is the exact
failure the frontmatter exists to prevent: status is meant to be
answerable from the record rather than by reading the code.

Closed with the commit that did it, and cross-referenced to 022 and 023,
which came out of the same mistaken instinct — treating the machine as
a boundary, then as an identity, then finding a consumer had a password
and no name to present with it.
2026-09-01 15:04:36 +02:00
jschoubben 0871e6ec11 37 — where a module lives, proposed
The module descriptions sit in `examples/` inside the control plane, and
that name has been doing harm: everything there reads as a sketch, and
one shipped naming a container image nothing builds. A directory called
the catalogue would have made "does this work" the obvious question.

The shape of the answer turns on one measurement. Of the 126 modules in
the system being replaced, 47 are software in their own right — the
largest is 182 source files, and a speech-capture module carries a whole
daemon. Another 44 ship helper scripts. Only 35 are a description and
nothing else.

So a catalogue cannot be a folder of manifests, because two thirds of
modules are programs. That splits them four ways, and only two of the
four belong in a catalogue: things the world made that we describe, and
packages with some files. What the mesh is made of stays in the
repositories that build it. What we wrote keeps its description beside
its code, in the same commit, because nothing else can stop the two
drifting.

The mesh's list of modules is a table, not a repository, and it already
records where each module came from and at which commit. Nothing needs
inventing for modules from anywhere; a repository of ours is just the
source we curate.

The check that a description is valid should move to a command on the
control plane's binary. Today a test reaches into the control plane's
internals to parse manifests, and another reads its build file to check
images exist — two jobs tangled. A command would also give the same
check to somebody describing their own application, which is the case
that matters most and has none.

Left open: how a provisioner's image gets published and pinned, and
whether thirty-five install-a-package modules deserve to be modules at
all.
2026-09-01 14:24:53 +02:00
jschoubben f6b1834ea9 024 fixed — the registry was addressed by hand and nothing else was
The cause was one line. Machines get a systemd-networkd unit with a
static address, so networkd finishes and reports the link configured.
The registry ran `ip addr add` inline, which leaves networkd waiting to
configure something it was never told about — and
systemd-networkd-wait-online has an infinite timeout.

So network-online.target was never reached and everything ordered after
it never started. On these machines that is Docker, so `docker load`
blocked on a socket whose daemon was queued behind a target that would
never come, and three bounded timeouts stacked to thirty-five minutes.
These machines have no DHCP by design, so that wait was never going to
end.

The hypothesis in this record was wrong and the record now says so.
Stocking had just been changed, so stocking looked guilty; stocking
takes 34 seconds and always did, timed directly before changing
anything.

Fixed with two things that made it cost hours instead of minutes: an
image placement now waits for the runtime and refuses after 120s naming
what systemd is waiting on, and the end-to-end test passes onProgress —
the raise reported every step and the test discarded it, which is why
thirty-five minutes and four minutes of silence looked the same.

The suite then ran to completion, 23 of 24, the one failure a check of
its own flagging a path as a credential because `/` is in the base64
alphabet.

Also recorded: a redirected log lags, because Node block-buffers stdout
to a file. Read as a stall twice, the second time right after the real
fix — where a buffering artifact argues the fix did not work.
2026-09-01 09:53:08 +02:00
jschoubben 407416e6d0 024 — a run stalls before the host is placed, and says nothing while it does
Seen twice today. Once mid-run: thirteen passes, then the process ended
with no summary, no failure and no receipt. Once from the start: the
first test ran 35 minutes against a measured 4.5 and was still running
when it was stopped.

Ruled out rather than assumed: not memory (84 GiB free, no OOM), not the
daemon (the stalled machine answered `incus exec` immediately), and not
the changes under test — the anchor VM had no host log and no
containers, so the run never reached placing the host.

What changed just before is that the rebuild went from two artifacts to
six, and every one of them is pushed into the scenario's registry, which
is the step the second stall sat in. Recorded as what changed, not as
the diagnosis.

The reason this is an issue and not a slow test: the suite prints
nothing between starting a scenario and finishing its first test, so
four minutes and thirty-five look identical from outside, and the only
recourse is to guess. That is how a workstation was left unbootable in
August. And a run that ends silently after thirteen passes is a run
somebody may believe.
2026-09-01 03:20:40 +02:00
jschoubben 39916b26e9 What stands between 3.1 and a running identity provider is a program
023 is fixed, so the design faults are gone and one concrete thing is
left: the realm provisioner does not exist. Its manifest named an image
nothing builds and no program backs, which has been removed — a manifest
describing a program nobody wrote is the same mistake as the credential
files that could never be read.

Keycloak's manifest now says what is true today, and the gap is loud: it
no longer claims to provide oidc-client, so a consumer asking for one is
refused by name at plan time instead of resolving cleanly and waiting
for a client nothing will create.

The provisioner should be written against a real Keycloak in the lab
rather than from the API documentation. The object store's took three
corrections that only a running server produced.
2026-09-01 03:13:11 +02:00
jschoubben 0760280bc6 The largest gap in the coverage list is not a gap
Tool servers — 56 modules, over half — were written up as the biggest
missing thing. They are expressible with what exists, and the first
framing was wrong in a way worth keeping: a module provides `tools` and
the session requires them does not work, because a requirement has one
answer and 56 modules offering tools would be 56 answers.

Turned around it fits exactly. The session provides `tool-host`; every
module offering tools requires it and contributes where its tools are.
Many-to-one is what `contributes` has always been, and the session
receives all of them in one file. Verified by resolving it rather than
by reading the code.

It only became possible today: until 022, several modules on one node
requiring the same thing was refused outright. Worth noting because it
means the credential fix bought more than credentials.

What remains is a decision about what a tool server is, which is work
rather than a missing shape.

The entry stays in the list rather than being deleted — a checklist that
quietly loses its biggest item reads as though nobody looked.
2026-09-01 03:09:11 +02:00
jschoubben f583502fc9 023 fixed — the mesh says who a consumer is, and what it is bound to
Both halves had one cause: the mesh knew something and did not say it.

Who a consumer is now comes from one derivation, sent to the provider in
its grant and to the consumer in its binding, so the two agree by
construction. The provisioners use the name they are given and refuse to
invent one, because a name of their own would create a login the
consumer could never guess while everything reported success.

Bound values reach the file that needs them through the symmetric twin
of the sealed placeholder — simpler, because they are not secret, so the
control plane fills them in and the host gains nothing.

The lab run meant to prove this failed in a way that looked like the fix
being wrong: rotation could not authenticate against a real database.
The cause was the suite rebuilding the control plane's image and not the
provisioner's, so an image built that minute ran against a provisioner
built the day before. That is 005's family and is recorded with the
issue, because the misleading part is worth more than the fix.
2026-09-01 03:05:43 +02:00
jschoubben 4adc656b54 Where Phase 3 actually stands
All three modules have manifests, all three parse, resolve and plan, and
none of them can start. Worth writing down before it reads as progress
or as failure, because it is neither.

The vocabulary held. Nothing in 3.1–3.3 needed a new shape — including
the mail system's several containers on a private network, which was the
one expected to break it. That was the question this phase was designed
to answer.

What did not hold was underneath: 022, now fixed, and 023, open. Both
are about credentials rather than about what a module can say.

The third fault was in the manifests, not the design: a secret declared
at a path named .env and read as one, when a sealed file holds a
password and nothing else. That is what a manifest checked only by a
parser buys, and it is why there are now two tests reading the manifests
on disk.

023 is the whole of what remains before the identity provider runs.
2026-09-01 02:56:54 +02:00
jschoubben dea46c3c6c Correct what the coverage document said about actions
I wrote that a module cannot declare an action. It could — the parser
accepted one, and the refusal only came on the machine. The claim was
wrong in the direction that matters: it read as "the design prevents
this", when what prevented it was a check at the far end that nobody
would connect back to the manifest.

Health checks are still the gap most worth closing, but the shape of the
answer is different from what I wrote. An action is not available to a
module at all, so a health check needs a way to say ask this and expect
that without saying run this — closer to a listens entry than to an
action.

Also records the finding itself, because it is a recurring shape here
and not a one-off: a rule enforced only at the far end is enforced and
unusable.
2026-09-01 02:56:05 +02:00
jschoubben 9073d3f2df Say plainly that env-file never points at a sealed secret
The playbook offered `env-file` and `${secret:name}` as alternatives,
and that reading is what produced the bug every example module shipped
with: own-secrets pointing at a path named `.env`, mounted as env-file,
holding a bare password. The container starts with no password set —
which is a service running on the wrong credential, not a failure.

They are not alternatives. A sealed file holds a password and nothing
else, so env-file points at a file the module declares whose content
leaves a hole, and the host fills it on the machine. A provisioner is
the exception, because it reads a password file.

Written out as the three lines a module needs, with the failure it
prevents named, since the abstract version was already there and was
read the other way.
2026-09-01 02:52:55 +02:00
jschoubben 2baf22ac43 A pair is a module and a provider, not two machines
Amends the credentials page, which said "every pair has its own
credential" and meant two machines. Built that way, it was wrong in a
way that only shows on a real node: a machine running several services
against one database server had one credential between them, so the
provider refused to plan at all and the consuming node quietly gave the
first module a credential and the rest nothing.

The page already argues the case against itself — one credential with
many holders is the first of the three faults it was written to remove.
It just drew the boundary at the machine.

Two modules on one node are as separate as two on different nodes, and
one login opening both is what this page exists to prevent. It is also
what makes withdrawal possible: one role per machine cannot say that
this module has lost its login and the others still have theirs.
2026-09-01 02:46:37 +02:00
jschoubben 80f18caf03 022 fixed; 023 filed — a password is not a connection
022 turned out to have a silent half worth recording: the provider
refuses loudly and names the modules, which reads as a decision, while
the consuming node does not refuse at all. Three modules wanting one
database produce one need, so two of them get no credential file and
each starts and fails to authenticate with nothing saying why.

023 is what remained after fixing it. A consumer now gets its own
password, in whatever shape its configuration wants, and still cannot
connect: the user name is invented by the provisioner and recorded
nowhere in the mesh, and the host and port sit in a JSON binding that an
application reading KEY=value cannot use.

The asymmetry is backwards and the coverage document now says so. The
secret is the hard case, because the mesh must not be able to read it,
and the secret is the part that arrives. The host and port are ordinary
facts the mesh holds in the clear, and they are the ones stuck.

Keycloak, Gitea, Mailu and MinIO all parse and resolve and none of them
can start. This is what stands between the module set and a running one.
2026-09-01 02:40:40 +02:00
jschoubben c86adbe3cc 022 — a credential belongs to a node, so a second consumer refuses
Found while checking whether the module vocabulary covers real use
cases. A node running three modules that all want a database cannot be
planned at all:

  anchor has 3 modules asking for "postgres-database" and they would
  share one credential: gitea, keycloak, umami

The refusal is right about what it says and wrong about what it implies.
They would share one credential, and sharing is worse than refusing —
but the arrangement being refused is the ordinary one, and the node this
mesh exists to take over runs eight modules against one database server.

The cause is the key: a credential is keyed by provision, consumer node
and provider node, so `consumer` is a machine. The provisioner inherits
it and names the role `mesh_<node>`. The refusal is not a check that
caught something; it is the only honest thing that function can do with
a key that cannot tell two consumers apart.

It is the same mistake as 021 with a different face. There the machine
was treated as a trust boundary; here it is treated as an identity, as
though "who is asking" is answered by naming a host. Two modules on one
node are as separate as two on different nodes.

Worth stating plainly: without the refusal, gitea's login would have
opened keycloak's database, and nothing would have said so — from the
provisioner's side it created exactly what it was asked to create.

Not a local fix. It crosses the control plane, the grant file naming and
every provisioner that names something after a consumer.
2026-09-01 02:28:39 +02:00
jschoubben 36d342f176 What a module must be able to say, measured against 127 that exist
Every manifest in the system being replaced was read and every key
counted, then set against what the new one can express. Three findings
worth more than the table.

**The most-used key was already covered and I expected a gap.**
Depending on another module — 65 manifests, the commonest thing any of
them says — is a requirement naming a module, which already means that
module rather than anything providing the name.

**The largest real gap is tool servers: 56 modules, over half.** A
module can already run one; what is missing is anything saying it offers
tools. That is plausibly a provision rather than new vocabulary, which
would need nothing added — not yet decided, and recorded as undecided.

**The gap most worth closing is health, at seven modules.** The mesh
knows a container is running, which is not whether it answers, and this
project has paid for that distinction twice. An action with a verify is
exactly the right shape and may not arrive over the link, so a module
cannot declare one.

Two things are missing deliberately and say so: stage hooks, because the
link may not carry an action and a module needing setup ships a program;
and flavours, retired in favour of claims.

Config merging is missing and should stay missing. A mechanism that
understands TOML gets asked for YAML, then INI, which is how the thing
being replaced became unholdable.

Also records what the survey found that is not about coverage: manifests
that had stopped matching what was actually brokered, one fact derived
in two places giving two answers, and a live listing returning
credentials in plaintext.
2026-09-01 02:20:34 +02:00
jschoubben 1ee62392b9 Playbook 06 — writing a module, from doing it once
Written after porting the first real workload end to end. Every step
exists because skipping it cost something, and the ratio is recorded
because it is the lesson: six attempts, one real bug, and the mesh was
right every time.

The rule worth carrying out of it: read the host's log before
theorising. A declaration that was sent and not applied says so there
and nowhere else — it took an hour to look, and the answer was one line.
2026-09-01 01:46:46 +02:00
jschoubben ecfc2c215e Point the plan at the survey, and name the likelier failure
The conversion's detail is operational and names machines, so it lives
in the mesh's knowledge base rather than in this repository:
`migration/where-service-data-lives` for where every service's data
actually sits, and `troubleshooting/db-password-frozen-at-first-init`
for the lockout. This document says the rule; those say the specifics.

The lockout is the finding worth carrying here, because it is worse than
the one this plan was already guarding against and it is likelier. A
database image consumes its password variable only when its data
directory is empty. Everything keeps data on a persistent directory, so
the role holds whatever password it was created with for ever;
regenerate the variable and the application moves on while the database
does not, permanently, because nothing reconciles it.

Eight modules are in that state today and work only because nobody has
regenerated their credential since their data directory was created.
It was already documented in the knowledge base and my survey had missed
it — found by searching, which is the argument for the knowledge base
existing.
2026-08-31 21:46:53 +02:00
jschoubben 122405df8e 020: the server version is not it either
Pinned 2.5.0 rather than latest, on the suspicion that its draft
profiles extension was involved. Identical failure, so that is ruled out
and recorded — two of the three guesses in this issue have now been
tested and both were wrong, which is the useful half.

The scenario keeps the pin regardless; it should have had one from the
start.
2026-08-31 21:42:13 +02:00
jschoubben acd5a0d80b File 020 — a certificate is issued and never collected; close Phase 1
Against a real ACME server the proxy orders, the challenge is answered
at the name on port 80 through the proxy itself, the authorisation goes
valid, finalisation is accepted, and the authority issues a certificate.
The client then posts to an empty URL to collect it, and never does.

Read from the authority's own log rather than inferred. Across one run
it issued two certificates and accepted finalise three times: the client
reaches issuance every attempt and fails at the same step after it.

Ruled out and recorded, so nobody repeats it: the directory is complete;
the authority's API certificate covers the address; the challenge path
works. A hand-written server config was suspected and was wrong —
replacing it with the server's own default, changing only the challenge
port, gives the identical error.

Filed rather than pursued because what remains is interop between two
libraries against a server that exists to be a test server, and may say
nothing about a real authority. What the mesh needed to show, it showed:
a routed name gets a certificate ordered from a configured authority,
and an unrouted one gets nothing — that second assertion passes.

Phase 1 closes with this one item partly open. Two of its four tasks
needed no code at all, the network shape was built, and the next thing
to learn comes from moving a module rather than a fourth lab run.
2026-08-31 21:40:03 +02:00
jschoubben f4347e2f14 Correct an overstatement: a sealed secret is readable on its node
An earlier paragraph implied a secret becomes unrecoverable once
accepted. It does not. It is sealed to the node, which holds the private
half and writes the plaintext into the module's own file at 0600 — the
value is there, on the machine, as an ordinary file.

What does not exist is a way to ask the mesh what a secret is. That is
the property worth having and it is narrower than what was written.

The reason to capture the old system's environment first is simply that
adoption means supplying those values, not that they become
unrecoverable.
2026-08-31 21:34:50 +02:00
jschoubben b68103d198 A module is adopted with the credentials it already has
Nothing is rotated during the conversion. A service keeps the password
it is already using, because minting a new one is how a running service
stops being able to reach its own database mid-migration.

The mesh has both paths already: generate-and-seal for a new module,
accept-and-seal for an adopted one. Adoption needs the second, and it is
built.

Rotation becomes a separate act afterwards, once everything works — the
machinery is proven, and it is a thing to do deliberately rather than as
a side effect of moving a service between systems.

Records the step that has to come first and is easy to miss: read the
current environment out of the old system while it can still be read.
Once accepted, the mesh cannot show a secret back, and once the old
system is gone neither can that. A password nobody wrote down is a
service nobody can adopt.
2026-08-31 21:31:26 +02:00
jschoubben f801b4b3e2 The board is published publicly, which decides everything else about it
Assumed throughout and stated nowhere — the wrong way round for the most
consequential fact about this component.

A board reachable only over the private network would sit inside the
boundary 0004 already calls the security boundary, and a login there
would guard a room whose door is inside the building. This one faces the
internet, so its login is a perimeter rather than defence in depth.

Which makes the identity provider the mesh's outermost gate. The board
presents the control plane, and the control plane's networked surfaces
can change the mesh (0035) — so whoever that provider admits can assign
modules, from anywhere. Written flatly because it is easy to arrive at
one reasonable step at a time and then be surprised by.

What follows is not the board's own design: who may log in is a decision
about the mesh rather than about an application; a public name needs a
certificate from an authority the world trusts, which is why that work
exists; and the provider going wrong in the permissive direction is a
mesh-wide exposure with no local symptom.

The command line is unaffected and is why this is tolerable — it
authenticates through nothing and answers to the machine's own login, so
the mesh stays operable by somebody standing at it whatever happens to
the gate. That is the property to protect if the rest is ever traded
away.
2026-08-31 21:26:06 +02:00
jschoubben a1a10e9ed2 Bootstrap ends at a usable mesh, and the first credential comes from a person
Bootstrap stopped when the control plane started — a mesh that runs and
cannot be used by anybody not standing at the machine, since the
networked surfaces need an identity provider and no module has been
assigned yet. It now runs through the provider and the first login.

The obstacle was not incidental. The mesh has never held a readable
secret: Make generates and seals, keeping no readable copy. An initial
administrator's credential is the first value a person must read.

Generating it and printing it once was the convenient option and is
refused. It would give the control plane a plaintext secret for the
first time — briefly, and to one terminal, but the capability would then
exist, and an exception made for one case does not stay one. The next
awkward credential gets printed too, and "a copy of the database is a
copy of nothing" stops being checkable by reading the code.

So the operator supplies it, on standard input, not echoed — the path
that already exists for a model-access key. What is created is an
account in the identity provider, not a user of the mesh; there is still
no user model.

Unattended bootstrap remains possible and the value still comes from
outside: automation supplying it is the operator supplying it. What is
refused is the mesh inventing one, so an unattended bootstrap with
nothing provided yields a mesh with no administrator — correct rather
than broken.
2026-08-31 21:23:55 +02:00
jschoubben 0ef6d3f574 One implementation, several surfaces, and what that costs
The mesh is operated from a command line and must be operable from a
browser and from a model's tools, without becoming three systems. The
pattern is already in the code and was unnamed: `board` serves HTTP by
calling the same functions the CLI calls, holding nothing.

Takes the decision 0034 said had to be taken deliberately rather than
arrive with a feature: the HTTP surface is not read-only, so a browser
login now carries authority over the mesh.

Names the dependency by protocol — an OAuth2 identity provider — as the
mesh does for AMQP, S3 and OCI. Keycloak is what fills the role; what
the control plane knows is that it validates a token, and replacing the
provider is a migration rather than a redesign.

Says what this must not become, because it is the failure the project
was started over: a kernel every module imports, 155 files of code from
every context. Shared surfaces are not a shared library. Three adapters
calling the same functions is not the same as logic leaving the context
that owns it.

And records the loop it creates. The networked surfaces depend on a
module the control plane assigns, so when identity is down nobody can
authenticate — including whoever is trying to fix it. The way out is the
command line, which authenticates through nothing and is available to
the account that owns the machine. Hence the rule: no capability exists
only behind an authenticated surface, because that is a capability which
disappears exactly when identity does.
2026-08-31 21:17:17 +02:00
jschoubben fccac61e58 The board is a web application, not a category
Supersedes 0032, which decided the right thing and described it wrongly.
The decision is unchanged: the account that installed the host owns the
mesh, and there is no user model.

What was wrong was inventing "a surface that delegates authentication"
for the board. It is a web application with a login, in the way every
web application has a login. That is a fact about an application, not a
property of the mesh.

The cost was not cosmetic. It made the identity module look like part of
the mesh's authority — something the mesh depends on to know who anybody
is — when the mesh knows nothing about people at all and one of the
applications running on it happens to have a login.

Keeps the line that is worth writing down, and states it more plainly:
signing in to an application must not become authority over the mesh.
Today it cannot, because the board reads and does not act. The moment it
can assign a module, whoever it lets in has mesh authority — and it
would arrive as a feature rather than as a decision. So a surface that
can change the mesh is a change to who owns the mesh, and is taken as
one. Not forbidden; just not something that turns up in a pull request
titled "add assign button".
2026-08-31 21:08:39 +02:00
jschoubben 10fa7c76d7 The substrate is a store and a broker
Third correction to one table today, found the same way as the other
two: by asking whether both halves of the test were answered, or only
the easy one.

0006 admits the registry because "it cannot grant itself a repository" —
true, and the second half. Nothing established that the control plane
needs one in order to run. Counted rather than argued: the bundle raises
twelve resources and no registry is among them. The registry arrives
afterwards as an ordinary module, which is exactly what the lab asserts.

0006 half-said this already, calling it "substrate by role and ordinary
by delivery, provisioned once there is a control plane to do it". A
member provisioned by the thing it supposedly precedes is not a member;
that phrase was carrying a contradiction rather than resolving one.

The registry is a closer call than the object store and the difference
is worth keeping: the control plane never touches an object store at
all, but it genuinely uses the registry. So the registry is a real
dependency of the mesh operating and not of the control plane starting —
and it is the second that the word means.

The substrate is now exactly what the bundle raises, which is the
strongest form the list can take: checkable by counting rather than by
reading an argument, and the two cannot drift.

The finding is not about substrates. A test with two conditions is a
test only when both are asked.
2026-08-31 20:30:19 +02:00
jschoubben a028337490 The local account owns the mesh; a surface delegates to a module
Answers what 0031 left open, and a question it did not ask — who owns
the mesh at all. There was no answer, and the absence was invisible
because every operation so far has been run by the person sitting at the
machine, so nothing had to say whether that was the design or the
circumstance.

The account that installed the host owns the mesh on that node. No user
model, no roles, nothing to administer. It follows from 0004 rather than
adding to it: there is no authorisation between nodes because every node
is the operator's own, so a user model inside that boundary would guard
nothing — anyone it could stop could read the node's key off the disk.

The board is different, and the difference is the network. A surface
reachable by a browser has to know who is asking, because those people
are not by construction people with a shell on the machine. So it
delegates to an OAuth provider, which is a module.

That does not make identity substrate. A surface delegating
authentication is not the control plane delegating it: the control plane
runs, applies declarations and reaches nodes with no identity provider
in existence. Only the board needs one.

Records the cost plainly: anybody with a shell on a node has full
authority there, and there is no way to give somebody authority over one
node without giving them a login on it.
2026-08-31 20:23:42 +02:00
jschoubben e3934e4449 The control plane authenticates nobody, so identity is a module
Closes the last open question about what the substrate contains. 0006
left an identity provider conditional — substrate only if the control
plane delegated authentication — and said the decision had not been
taken. It is now: it delegates to nothing.

The conditional was never about machines. A node proves itself with a
keypair it generated over a broker account issued at enrolment, and
declarations are verified by signature; none of that involves an
identity provider. It was only ever about whether a person signing in to
a mesh surface would be authenticated by something else.

So the substrate is three — a relational store, a message bus, an image
registry — and with 0028 having removed the object store, no member is
conditional and every one is there for the same reason.

It does not settle how a person signs in to a surface, deliberately.
What is settled is that whatever answers that is not something which
must exist before the mesh does, so it can be decided late or replaced —
which being substrate would have prevented.
2026-08-31 20:18:03 +02:00
jschoubben c570c687f6 The handover: switch off the old brain, leave the services running
The conversion method, recorded because it decides everything else and
was not written down.

The old control plane is stopped — provisioning, coordinator, syncs, the
pipeline, anything that decides or writes. The workloads it was managing
keep running, because nothing is managing them. The new mesh then takes
ownership one module at a time.

Nothing is ever unassigned in the old system. Unassigning is how it
removes things and removing is how data is lost; it is asked to stop
having opinions, never to take anything away.

Disabled rather than merely stopped, which is the part easy to get
wrong: those units are enabled, so a stop lasts until the next reboot. A
reboot mid-conversion would bring the old control plane back to
regenerate managed files underneath the new one — the one situation
where two systems really would fight over a machine.

A service left running with nothing managing it is the safe state: it
has its data, its configuration is on disk, and nothing will change
either. The risk in a conversion is in the managing, not the running.

Also records why taking ownership piecemeal is safe: the new host's
orphan removal is per-origin, so it only removes what it recorded
itself. Services it was never told about are not orphans to it.
2026-08-31 19:55:35 +02:00
jschoubben d1ab2dc0b4 Data outlives the mesh that declared it, and the conversion starts where it lives
0030, found by asking what the conversion actually needs rather than by
reviewing anything. The host deleted a directory and everything under it
when it stopped being declared — which happens when a module is
unassigned, or when a manifest is edited to move a data folder, which is
the exact operation this plan needs. A database's files, a mail spool.
The report said "removed".

A directory still holding something is now kept and said so. No flag and
nothing to remember: emptiness is the test, and it works because the
removal order was already right — the mesh's own contents are gone by
the time the directory is reached, so what remains is by definition
something nobody declared.

The plan now says data outranks its own ordering: copy, read back
through the service that owns it, and only then point anything at the
new location. Never move and then check.

And it records where this starts — the node holding all the production
data — with what that costs stated rather than argued with. Everything
proven so far was proven on machines that could be destroyed and raised
again. A scenario proves the mechanism, not the state on that machine.
2026-08-31 19:54:41 +02:00
jschoubben 9ad0ec35e0 The conversion is done by hand, and that removes work from this plan
Recorded because it is load-bearing and was not written down: moving
from the current system to this one is a person at a command line, not a
migration program.

What that removes is larger than what it adds. Nothing in this plan
needs an importer, a translation layer, a compatibility shim, or a way
of keeping two systems agreeing while both are live — each of which
somebody would otherwise reasonably build, use once, and maintain for a
year.

It also settles what "safe" means for the system being retired: a fix to
it must be safe on its own, because there is no careful rollout to
sequence it into. A change needing three steps in the right order is a
change that will be half-applied. That reversed a certificate default I
had chosen this morning.
2026-08-31 19:46:21 +02:00
jschoubben e4327a3a5e Phase 1.3 done: ordering was already there, the network was not
Ordering needed no change for the third time running — resources apply
in the order declared and nothing sorts them — and is now asserted,
because sorting them for any sensible reason would have passed every
other test.

Separates ordering from readiness, which the task had run together: a
container started is not a container ready. Nothing waits, and what
needs something usable retries. That is deliberate and more robust than
start ordering, since a dependency can restart long after apply.

The network was the first thing in Phase 1 that genuinely needed
building, and the first that needed a decision: 0029 records why a shape
rather than an action, and the vocabulary is nine.
2026-08-31 18:55:22 +02:00
jschoubben ce486fd5d2 Phase 1.2 done, and it is the same surprise as 1.1
A session as a licence consumer needed no change either: the two
sessions are two modules, so the existing (node, module) binding already
names them apart. 14-model-access.md's "a step toward it and not it" is
true of a worker and not of a session, and the difference is that there
is one session per node rather than many per machine.

Records what stays open: the worker half of that gap is real and
unaffected, and belongs with 0003, which is unbuilt.

Two tasks in a row that were already possible. Both were written from
the design rather than from the code — the review's own finding arriving
in the plan it produced. The remaining Phase 1 items should be checked
against the code before being started rather than after.
2026-08-31 18:41:19 +02:00
jschoubben fdd909ec40 Phase 1.1 done, and it was not the task that was written down
An object-store provision, proven against a real store with seven
assertions.

The finding is worth more than the task: the control plane
special-cases nothing. provides, requires, contributes and grants are
name-agnostic, so asking for a bucket needed no change to the mesh at
all. What was missing was a provider and the last step on the machine —
"add an object-store provision" was never mesh work, and the breakdown
now says so rather than leaving the next person to rediscover it.

Named s3-bucket by 0027: the coupling is to the API, not the product,
because swapping one store for another does not break a consumer. A
database is the other case and names its engine.

Records the assertion a database does not need, because it is the one
that will be forgotten when somebody writes the next provider: one store
holds every bucket behind one endpoint, so isolation is a policy rather
than a property, and a policy granting everything passes every test that
only checks a consumer can reach its own bucket.
2026-08-31 17:53:02 +02:00
jschoubben cb1954e7b5 Accept 0024, and rewrite the work breakdown around what is actually being done
**0024 accepted.** Model access was decided, built, and proven in the
lab, and two design documents rest on it; only the status had never
moved. The gate is green again.

**The work breakdown rewritten.** It planned a decomposition of the
existing system in place — extract contexts, declared features, shrink
the shared library. That is not the work. A replacement is being built
beside it, and only the old Phase 0 survived contact with reality, so
the one document meant to say what happens next was describing a system
being retired.

Now ordered by what "modules move across one at a time until the old
registry is off" actually requires:

- Phase 0 is marked done against the twenty-two lab assertions, **and
  carries its own limitation**: every module exercised was written to
  test the mechanism, so the vocabulary was shaped by its own fixtures.
- Phase 1 is the vocabulary gaps found by asking what real modules
  need — an object-store provision, a session as a licence consumer, a
  network shape with ordering, public certificate issuance.
- Phase 2 is one module, then a week of running it, because the point of
  going first is to find what Phase 1 missed.
- Phase 3 picks modules that each prove something the first did not; the
  mail system is last because it is the one that may send work back into
  the declaration language.
- Phase 4 is switching the registry off, named as a phase so it is not
  mistaken for the goal.

Keeps the rules of engagement unchanged — they were about how work is
done, not what it is — with one addition: stop and ask before anything
that touches a machine outside the lab.

Adds a section on keeping the list true, since the document it replaces
was wrong for weeks and nothing said so. A claim here is counted, not
reasoned, and a phase is done when the lab says so.
2026-08-31 17:25:27 +02:00
jschoubben 1b5308c9cc Review of the to-be layer: check what the documents claim against what runs
First pass of a design review, done by reading documents against code
and against a raised mesh rather than against each other. Every error
below was invisible to a proofread.

**Statuses were stale, and nothing checked them.** Ten to-be documents
said `designed` while naming working, lab-proven code — several with a
*What was built* or *Raised, and observed* section. Added a
`status-vs-code` check: naming a file is a claim that the file
implements this, so a document that points at one has stopped being
merely designed. It failed on all ten before it passed, per the rule
this folder sets for its own checks.

**The bundle carries three images, not two.** 07 reasoned about which
substrate services go in and overlooked that the control plane is in
there too — it is what the substrate exists to start, and there is
nothing to fetch it with yet. Counted, not deduced.

**The bootstrap uses four shapes, not six.** It listed `file` and
`directory`, which substrate-first-node.lock never asks for. The claim
that mattered — nothing is blocked on the host — was true either way,
which is why the wrong count survived.

**The eight capabilities were documented nowhere.** Implemented in
internal/profile/detectors.go and enumerated in no document, including
the one about the host that detects them. A vocabulary modules write
against, readable only by reading the code. Now written down, with the
seat/graphical-session distinction that is wrong in both directions if
collapsed.

**MinIO swept out of the to-be layer** per 0028.

The gate now fails on one thing left deliberately: ADR 0024 is
`proposed` while two documents rest on it and the feature it decides is
built and lab-proven. Accepting a decision is not mine to do.
2026-08-31 17:20:47 +02:00
jschoubben cbcbba8099 A provision names its engine; the substrate supplies only the control plane
**0027 — provisions.** A module written against PostgreSQL could be
matched to a provider of SQL Server, resolve as satisfied, and fail on
its first query. The name said the role, so nothing distinguished
engines. Refusing on ambiguity could not help: with one provider of
each name nothing is ambiguous. Enforced at parse rather than
documented, because the old naming was the documentation.

**0028 — the substrate.** 0006 admits an object store on the grounds
that it cannot grant itself a bucket. That answers the second half of
the test and assumes the first: the control plane does not need one.
Verified — no S3 client in mesh-control, and internal/builder/registry.go
records the deliberate choice to put artifacts in the OCI registry as
content-addressed blobs. The row was inherited from the system being
replaced, where an object store distributed module tarballs, and was
never re-tested against the definition above it.

So an object store is an ordinary module, and a mesh with nothing
needing one runs none. Migrating it is module work, not substrate work.

0028 also states what 0006 left unsaid: a substrate service and a
module of the same product are different instances. The substrate is
raised from the bundle before any mesh exists, so it is not in the
module graph — a workload depending on it would depend on something the
graph cannot see, cannot rotate a credential for, and cannot move, and
would put workload data in the store the control plane keeps its own
state in.

Both records were found by reading code against design rather than
design against itself, which is the review that should have happened
sooner.
2026-08-31 17:13:07 +02:00
jschoubben 046a990198 Several sessions at once is a surface property, not a session one
Answers the question 15 raised: a board showing many sessions leaves
one-per-node untouched, because each is still one conversation. Only
concurrent conversations with the same session would touch 0004.

Records soulstream and herdr as the prior art to draw from, and marks
it explicitly off the provisioning path so it stays a note rather than
becoming the work.
2026-08-31 16:52:15 +02:00
jschoubben fbf5b1d04b A session's memory is its own, and it is not declared
Settles the question 15 left open: the mesh session holds its own
memory in the mesh root, rather than assembling a view over the node
sessions. Memory follows the rule the rest of the design already uses —
the context root is the whole of what makes one session a different
agent, and memory is part of what makes it that agent.

The control-plane node is what makes this load-bearing rather than
tidy. Two sessions share that machine; if memory belonged to the
machine instead of the root they would share it too, and the mesh's
recollection would be indistinguishable from that node's own — the
collision 0026 exists to avoid, arriving through the back door.

Also corrects an error made writing it up: memory is NOT declared
state. The engram and tools are — the mesh says what they are and the
host writes them (0011). Memory is written by the session itself and
declared by nobody, so a mechanism that regenerates the root wholesale
would erase it on the next heartbeat, silently, while reporting
success. The root is not uniformly managed and which parts are has to
be explicit.
2026-08-31 16:41:06 +02:00
jschoubben 3c6c16abdf The mesh has a session of its own, and it is the node session's mechanism
A session for the mesh itself, addressed as the mesh, differing from a
node's in exactly three things: the context it starts in, its engram,
and its licence binding. Not a new kind of agent — the same mechanism
pointed at a different root. Two implementations of one mechanism drift,
and the vocabulary collision 0001 exists to undo began exactly that way.

It runs on the control-plane node, and the reasoning is easy to get
backwards: not "the important agent on the important machine", but that
this node is already the one place excepted from "compromise of a node
is compromise of that node". Placed anywhere else it would create a
second such place.

It is an addition to per-node messaging and never a replacement. 0001
holds that losing the control plane costs change, not operation — and a
mesh whose only conversational surface lived there would lose the
ability to ask anything while every machine kept running perfectly.

Writing it up exposed that the node session's setup was never designed
at all. 0004 gives behaviour and stops: nothing said how a session
starts, where its context lives, or how a broker message becomes a
prompt. That gap was invisible until something had to be built *like* a
node session. 15-the-agent-session.md covers both as one mechanism.

It also makes "a consumer that is not a machine" undeferrable. The
control-plane node now hosts two sessions that must hold different
licences, and a per-machine binding cannot express that at all. Noted in
14-model-access.md against the gap it was already recorded as.

Also completes the to-be index, which stopped at 10 and omitted four
documents. Pre-existing broken ADR references in the older rows are left
alone rather than guessed at.
2026-08-31 16:32:59 +02:00
jschoubben e823cc1cc5 The design record is read where it is written, never copied to be found
Decides the question 006 narrowed to. An agent reads this repository
directly and the search consults it, so these documents surface beside
ordinary results instead of only when somebody already suspects they
exist.

A scheduled sync into the mesh's memory was the option that works with
what exists today, and lost on the ground this repository can least
afford: it makes a second copy, and the copy that is searched quietly
stops matching the copy that is edited. A design record that has
silently diverged from the reasoning it claims to carry is worse than
one that cannot be found — the first misleads, the second merely fails.

Amends what 0019 promised rather than satisfying it: these documents
will not be indexed, they will be read. The commitment that survives is
the one that mattered — that a searcher finds them without already
suspecting they exist.

Gated on an agent that does not exist yet, so 006 stays open on the
build with a decided shape. What closes it is a check that fails today
by design: search the mesh's memory for a phrase that appears only in a
design document here, and require it back.
2026-08-31 15:45:00 +02:00
jschoubben c192810fba 004 and 008 resolved; 006 narrowed to the decision it actually needs
**004 — certificate issuance.** The resolver declared no authority at
all, so the client fell to its built-in production default: there was no
setting set wrongly, there was no setting. It is now a node property
defaulting to staging, which answers the first open question. Staging by
default rather than production-with-an-override, because the alternative
leaves the safe path depending on remembering to opt out of it — 005's
lesson, in a second place. The rollout is ordered and the order is the
dangerous part; recorded, not performed.

**008 — node rescue.** Read back from running nodes as the report asked,
and one of its own claims was wrong in a way that matters: the health
timer does exist and does fire. It simply never calls the rescue script.
A trigger that exists and does not do what the script claims survives a
halfway check, which makes it worse than the absence the report
described. Resolved by making the documentation true, not by
implementing rescue — the replacement host already supervises recovery,
and wiring unattended restart into the fleet being retired is a
deliberate decision rather than a tidy-up. Two "self-healing" claims
narrowed to what they actually do.

**006 — deliberately not closed.** Re-checked today: the indexing still
does not exist. What is gone is the reason it was an issue — the claim
is no longer load-bearing, because the README names the gap and the
decision's reasoning never invoked indexing. A signpost now points here
from the knowledge base, and was measured rather than assumed: it is
reachable, it is not surfacing. Closing it while the indexing does not
exist would be this repository's own named failure, one folder from
where it names it.
2026-08-31 15:20:11 +02:00
jschoubben 345bbe0552 005 resolved: a suite that cannot run on every push says when it last ran
Retired in favour of the lab rather than repaired — that answers the
first open question. The second finding is the one that generalises:
"nothing runs it, and nothing reports that nothing runs it" is not a
fact about that harness, it is a fact about any suite too expensive to
run on every push. The replacement inherited the fault it was replacing.

Records the three rules that now hold, and what the fix taught twice:
the remedy rebuilt the symptom inside itself, and the code that counts
results passed every test while reading nothing.
2026-08-31 15:02:26 +02:00
jschoubben f3ffdae909 Record the resolver as built, and the two things it must not do
A service is reached at <service>.<node>.internal, so what resolves is anything
under a node's name. The mesh writes the data and runs no daemon; two roles,
two claims, because systemd-resolved cannot serve a wildcard at all.

Both prohibitions were found by a machine rather than by reasoning: an address
systemd already held, and reading resolv.conf for upstreams that now point at
itself.
2026-08-31 14:23:02 +02:00
jschoubben 6ecd03694b Issue 019: a comment asserting a fact about a machine, which nothing checked
Twice in one file, a statement about a machine that read as reasoned and was
wrong — and the module's unit tests all passed while the daemon could not
start. That is what a unit test is: it confirms the assertion was made, never
that it is true of any machine.

003 in prose rather than in a manifest key.
2026-08-31 14:11:34 +02:00
jschoubben a036bac47b Record why the module's own secret is named for whose it is
The field was called needs, beside secrets, and both were name-to-path holding
something secret. What separates them is whose, not how secret — so that is
what the name says now.
2026-08-31 13:47:21 +02:00
jschoubben aafeb5c9df Three issues resolved: one closed by evidence, two answered by the replacement
012 named its own closing condition — a scenario with four images coming up —
and the scenario now stocks seven and has raised cleanly many times at the
memory the wrong diagnosis had raised.

001 is answered by the host reading the package database back after installing.
002 was NOT answered and was present here too, so it is a fix rather than a
note: a stale index is now named instead of reported as a failed install.
2026-08-31 13:00:03 +02:00
jschoubben 573a94e102 Correct the record: a limitation that no longer exists, and one that was never written
The connectivity design still said a hub cannot be filtered — a gap recorded in
the morning and closed in the afternoon, left standing as though it were
current. Worse than a stale date: it would send somebody away from something
that works.

`restart-on` was described nowhere, including the part added today that lets a
service reflect a file another module put on the machine. A rule the host
enforces and no document mentions is a rule nobody can rely on.

And nine of fifteen design documents claimed an `updated:` older than their last
change, some by a week. That field is what cross-cutting views are generated
from, so it is not decoration.
2026-08-31 12:33:35 +02:00
jschoubben f4e81074d3 Which resolver is a claim, and was decided before it was asked
A resolver takes over /etc/resolv.conf, which is a singular resource — ADR 0009
lists it in the table beside the seat and pid 1. So choosing between resolved,
dnsmasq and unbound is assigning a module, per machine, and the mesh refuses
two rather than letting them fight over the file.

Recorded because it was treated as an open question two days after being
decided, which is the argument for that table being a table.
2026-08-31 12:15:12 +02:00
jschoubben a3cee17d48 Record what a container can see of the mesh's names, and what it cannot
Found by a container failing to resolve a name every machine could: a container
gets its own hosts file holding only its own hostname, and on the machine it
always worked, which is what made it easy to miss.

Declared containers are given the names. A container somebody starts by hand is
not the mesh's to configure — which is a second, different reason to want a
resolver, recorded beside the first rather than folded into it.
2026-08-31 11:29:25 +02:00
jschoubben 35ce23fb76 Record that a machine coming back is ordinary, and what waking now does
Asked whether a machine that drops off needs re-adopting: it does not, nothing
expires, and the only thing that forces re-enrolment is losing its own key.

The gap was the twenty or thirty seconds after a resume in which a node
believes it is in a mesh it has left — recovering on its own, which made it a
quality gap rather than a fault, and still a machine waiting to be told
something it already knew.
2026-08-31 10:21:50 +02:00
jschoubben 6374c1eb60 Record what "behind" means now
It meant failed-or-refused, so the question this record says must not be lost
was answerable only for the machines that broke. Out of date, never told, and
not worked out are kept apart: the remedy is the same push and they read
differently to whoever is looking.
2026-08-31 05:28:11 +02:00
jschoubben 7c0be6968e Record what a machine says about itself and what the mesh keeps
The yes gates an assignment and the detail carries a value; they are one fact
read two ways, and only one read was being kept.
2026-08-31 04:54:07 +02:00
jschoubben 7f314eb399 Point three design documents at the code that exists for them
Their subject matter has been built and proven for days and their frontmatter
still said code: [] — which is what the cross-cutting view is generated from,
so it was claiming nothing existed for the substrate, the node lifecycle and
delivery.
2026-08-31 04:51:04 +02:00
jschoubben 5ad3641cbf Record the board, which is the last designed document with no code
One reading answered three ways, holding nothing and touching no context's
store — which is the constraint the whole document is about, and the thing the
board being replaced gets wrong.
2026-08-31 04:50:45 +02:00
jschoubben 0bc4b7774f Record what four more pieces of the mesh became
Rotation and the provisioner contract; model access as a provision answered by
a record, with ADR 0024's other two gaps left as gaps; exposure, which closes
the open question about revoking a route; and the delivery loop, which closes
the gap ADR 0010 left when it replaced a pipeline with a comparison.
2026-08-31 02:56:50 +02:00
jschoubben 6b1c80b442 What a builder-as-a-module can and cannot reach
The broker's fingerprint travels with its credential, and the machine's
filesystem does not travel at all — it runs in a container, which is the
arrangement working rather than a limitation to route around.
2026-08-31 01:53:35 +02:00
jschoubben 8448219de1 Issue 018: a provider on the same machine was never announced to its consumer 2026-08-31 01:35:54 +02:00
jschoubben 00e98f1f92 The vocabulary has no word for a unit that runs and exits
Found by the firewall: every packet filtered as declared, and the machine
reported as not doing what it was told, because the unit that loaded the rules
had finished. Stated as a gap rather than worked around silently.
2026-08-31 01:20:58 +02:00
jschoubben 0ac99f0d68 An action's own idea of being finished must be its verify's
Otherwise it succeeds into a state its verify rejects, and the host's report is
accurate and names nothing. Recorded where the vocabulary is described, because
it is a rule about writing an action rather than about one action.
2026-08-31 01:03:50 +02:00
jschoubben 1ade18209d Issue 017: an action succeeded into a state its own verify rejects 2026-08-31 01:01:17 +02:00
jschoubben a9cd3de5be Issue 016: anything after the declaration in a file was ignored 2026-08-31 00:51:46 +02:00
jschoubben 6e7e77acbe Issue 015: the harness read a swallowed answer as success 2026-08-31 00:48:16 +02:00
jschoubben 890c3ee3fc State the one rule the derivation does not yet reach
A hub needs its overlay port open and a node that is not a hub does not, and
they are the same module — so listens, a static manifest field, cannot express
it while the overlay module's resources are computed per node. Written down
rather than left as an oversight for whoever first puts a firewall on a hub.
2026-08-31 00:42:39 +02:00
jschoubben a6872ac099 A key that is present and unusable, and what the certificate work became
Issue 014: the node's serving key was stored in the host's own encoding, so
every check that reads the file passed and no server could start. Same shape as
013 — two halves of one mechanism designed separately, each correct about its
own half. Where a file exists so a third party can read it, the format is the
interface.
2026-08-31 00:42:13 +02:00
jschoubben 778efaba8b The mesh runs its own registry, certifies its own names, and computes its own filtering
Issue 003 is answered in both halves: manifests are parsed strictly, and a
module says what it listens on and from where rather than carrying a key
nothing reads. The design records what was built and how each part is checked.

Issue 013 is new, found by reading while writing the first module that has
both a computed file and a service that needs it. The file arrived second.
It failed, then the next reconcile fixed it, which is why nothing caught it.
2026-08-31 00:37:34 +02:00
271 changed files with 20389 additions and 545 deletions
+1
View File
@@ -9,6 +9,7 @@ works — plus the engineering practice that holds across everything Novox build
| [`context.md`](context.md) | The environment — conditions, not aspirations |
| [`effect.md`](effect.md) | What is different when the work is done |
| [`how-we-build.md`](how-we-build.md) | The rules that hold across the mesh, each one earned. **The source of the mesh constitution** — the governed page the mesh injects into design sessions is derived from it. |
| [`glossary.md`](glossary.md) | One name per thing — the authority on vocabulary, and the words that were retired |
| [`repos.md`](repos.md) | Where implementation lives, and what each repository owns |
| [`process/`](process/) | The playbooks — how work moves through this repository, for engineers and agents alike |
+8
View File
@@ -26,6 +26,7 @@ indistinguishable from one that cannot.
| `numbering` | the number in the filename is the number in the heading | — |
| `topics` | every record names a topic the index knows | — |
| *(index.py)* | the written reading order matches what the records say | — |
| `status-vs-code` | a to-be document naming specific code is not still `designed` | **ten documents**, several with a *What was built* section, describing lab-proven code |
## What is deliberately not checked
@@ -47,3 +48,10 @@ what to do.
State what incident it would have caught, and make it fail before you make it pass. A check
whose failure has never been observed is a guess about its own correctness.
## cycle.py
The development cycle, checked ([ADR 0080](../../02-DECISIONS/0080-the-development-cycle-is-checked.md)):
a to-be design names a decision, an in-progress/implemented design names its owning code, a
located/fixed issue names its owner, a fixed/resolved issue says what fixed it, a graduated
research overview says what it became. `python3 00-META/checks/cycle.py`
Binary file not shown.
+188
View File
@@ -0,0 +1,188 @@
#!/usr/bin/env python3
"""The development cycle, checked.
The knowledge flow (00-META/process/00-overview.md) says work moves idea -> research ->
decision -> to-be design -> code, and symptom -> issue -> diagnosis -> fix. Those are rules,
and a rule states how it is checked (AGENTS.md) -- this is how. Everything here reads only
frontmatter, because status lives in frontmatter and nowhere else.
What is enforced:
design every 03-DESIGN doc parses, carries `layer:` matching its directory, and a
known `status:`. A TO-BE doc names at least one decision (`decisions:`) -- no
design without a decision -- and once `in-progress` or `implemented` it names
its owning code (`code:`) -- no development without a design that says where.
issues a known `status:`; once `located`, `located-in:` names the owner;
once `resolved`, `fixed-by:` says what fixed it (prose counts --
"nothing, the capability existed" is an answer).
research a known `status:`; a `graduated` overview says what it `became:`, and every
target it names exists.
decisions every accepted record is REACHABLE from the cycle: cited by a design doc's
frontmatter, a research overview, an issue report, a 00-META document, or another
record's extends/supersedes chain. A decision nothing points at is one nobody will
find by following pointers -- which is how records go stale in people's heads.
Deliberately NOT enforced: `resolved` issues may leave `located-in` empty (a symptom that
turned out not to be a defect has no owner), and as-is docs need no decisions (they
describe what exists, not what was decided).
python3 00-META/checks/cycle.py
"""
import glob
import os
import re
import sys
ROOT = os.path.normpath(os.path.join(os.path.dirname(__file__), "..", ".."))
DESIGN_STATUSES = {"proposed", "designed", "in-progress", "implemented", "abandoned"}
ISSUE_STATUSES = {"open", "diagnosing", "located", "resolved", "wontfix"}
RESEARCH_STATUSES = {"active", "graduated", "abandoned"}
def rel(path):
return os.path.relpath(path, ROOT)
def frontmatter(path):
"""The YAML block between the first two --- lines, as {key: raw-value-string}.
Minimal on purpose, like records.py: enough for the fields these checks read. A list
value (block or inline) is joined into its items; a scalar stays a string.
"""
text = open(path, encoding="utf-8").read()
m = re.match(r"^---\n(.*?)\n---", text, re.S)
if not m:
return None
front, out, key = m.group(1), {}, None
for line in front.split("\n"):
item = re.match(r"^\s+-\s*(.+?)\s*$", line)
if item and key:
out[key].append(item.group(1))
continue
kv = re.match(r"^([A-Za-z-]+):\s*(.*)$", line)
if not kv:
continue
key, value = kv.group(1), kv.group(2).strip()
if value.startswith("[") and value.endswith("]"):
out[key] = [v.strip() for v in value[1:-1].split(",") if v.strip()]
elif value == "":
out[key] = [] # a block list may follow; stays [] if nothing does
else:
out[key] = value
return out
def listy(front, key):
v = front.get(key)
if v is None:
return []
return v if isinstance(v, list) else ([v] if str(v).strip() else [])
def main():
failures = []
def bad(path, why):
failures.append(" %s: %s" % (rel(path), why))
# ---- design ------------------------------------------------------------------------
for layer, name in (("00-as-is", "as-is"), ("01-to-be", "to-be")):
for path in sorted(glob.glob(os.path.join(ROOT, "03-DESIGN", layer, "*.md"))):
if os.path.basename(path) == "README.md":
continue
front = frontmatter(path)
if front is None:
bad(path, "no frontmatter")
continue
if front.get("layer") != name:
bad(path, "layer is %r; this directory is %s" % (front.get("layer"), name))
status = front.get("status")
if status not in DESIGN_STATUSES:
bad(path, "status %r is not one of %s" % (status, sorted(DESIGN_STATUSES)))
if name == "to-be":
if not listy(front, "decisions"):
bad(path, "names no decisions -- no design without a decision")
if status in ("in-progress", "implemented") and not listy(front, "code"):
bad(path, "status %s but code: names no owner -- no development "
"without a design that says where" % status)
# ---- issues ------------------------------------------------------------------------
for path in sorted(glob.glob(os.path.join(ROOT, "04-ISSUES", "*", "00-report.md"))):
front = frontmatter(path)
if front is None:
bad(path, "no frontmatter")
continue
status = front.get("status")
if status not in ISSUE_STATUSES:
bad(path, "status %r is not one of %s" % (status, sorted(ISSUE_STATUSES)))
if status in ("located", "resolved") and not listy(front, "located-in"):
bad(path, "status %s but located-in is empty" % status)
if status == "resolved" and not listy(front, "fixed-by"):
bad(path, "status %s but fixed-by says nothing" % status)
# ---- research ----------------------------------------------------------------------
for path in sorted(glob.glob(os.path.join(ROOT, "01-RESEARCH", "*", "00-overview.md"))):
front = frontmatter(path)
if front is None:
bad(path, "no frontmatter")
continue
status = front.get("status")
if status not in RESEARCH_STATUSES:
bad(path, "status %r is not one of %s" % (status, sorted(RESEARCH_STATUSES)))
if status == "graduated":
became = listy(front, "became")
if not became:
bad(path, "graduated but became: names nothing")
for target in became:
if not os.path.exists(os.path.join(ROOT, target)):
bad(path, "became names %s, which does not exist" % target)
# ---- decisions -------------------------------------------------------------------
records = {}
for path in sorted(glob.glob(os.path.join(ROOT, "02-DECISIONS", "[0-9]*.md"))):
front = frontmatter(path)
records[os.path.basename(path)] = (path, (front or {}).get("status"))
cited = set()
sources = (glob.glob(os.path.join(ROOT, "03-DESIGN", "*", "*.md"))
+ glob.glob(os.path.join(ROOT, "01-RESEARCH", "*", "00-overview.md"))
+ glob.glob(os.path.join(ROOT, "04-ISSUES", "*", "00-report.md"))
+ glob.glob(os.path.join(ROOT, "00-META", "**", "*.md"), recursive=True))
for path in sources:
text = open(path, encoding="utf-8").read()
if os.sep + "03-DESIGN" + os.sep in path:
# A design doc's governing citations live in frontmatter; a prose mention is
# commentary, not a home.
m = re.match(r"^---\n(.*?)\n---", text, re.S)
text = m.group(1) if m else ""
for m in re.finditer(r"([0-9]{4}-[^\s\)\],#]+\.md)", text):
cited.add(os.path.basename(m.group(1)))
for name in records:
front = frontmatter(records[name][0]) or {}
for key in ("extends", "supersedes", "superseded-by"):
v = front.get(key)
if isinstance(v, str) and v:
cited.add(os.path.basename(v))
for name, (path, status) in records.items():
if status == "accepted" and name not in cited:
bad(path, "an accepted decision nothing in the cycle cites -- give it a home in a "
"design doc's decisions:, a research became:, an issue, or 00-META")
checked = (
len(glob.glob(os.path.join(ROOT, "03-DESIGN", "0*", "*.md")))
+ len(glob.glob(os.path.join(ROOT, "04-ISSUES", "*", "00-report.md")))
+ len(glob.glob(os.path.join(ROOT, "01-RESEARCH", "*", "00-overview.md")))
+ len(records)
)
if failures:
print("cycle: %d document(s) break the development cycle:" % len(failures))
print("\n".join(failures))
return 1
print("cycle: %d documents checked, the chain holds" % checked)
return 0
if __name__ == "__main__":
sys.exit(main())
+32
View File
@@ -284,6 +284,37 @@ def check_numbering(failures, records):
)
def check_status_against_code(failures):
"""A design document naming specific code may not still call itself `designed`.
**Naming a file is a claim that the file implements this**, so the two fields have to agree.
They drifted: ten to-be documents named working, lab-proven code — several with a *What was
built* or *Raised, and observed* section — while still saying nothing had been built.
Deliberately weak, and that is the point of it being mechanical. It cannot tell whether the
prose is true, only that a document has stopped claiming to be unbuilt once it points at
something. `code: [mesh-controller]` — a repository with no path — is a plan and stays
`designed`.
"""
for path in markdown_files():
if not rel(path).startswith("03-DESIGN/01-to-be/") or path.endswith("README.md"):
continue
front = frontmatter(read(path))
if front.get("status") != "designed":
continue
for entry in front.get("code") or []:
named = re.sub(r"\s*\(.*\)$", "", entry).strip().split(None, 1)
if len(named) > 1:
failures.add(
"status-vs-code",
rel(path),
f"`designed`, but names {named[1]!r} in {named[0]}. Naming a file claims "
f"it implements this — use `in-progress`, or `implemented` once it is "
f"defensible from that repository's main branch.",
)
break
def main():
failures = Failures()
records = load_records()
@@ -293,6 +324,7 @@ def main():
check_supersession_symmetry(failures, records)
check_numbering(failures, records)
check_topics(failures, records)
check_status_against_code(failures)
print(f"records: {len(records)} decision records checked")
return failures.report()
+1 -1
View File
@@ -30,7 +30,7 @@ named, and nothing should be designed around a particular one existing.
- **A hosted model provider** supplies the thinking for non-human agents, drawn from a
shared pool of subscriptions — which is why budget pacing is a first-class concern.
- **Long-lived user services** rather than an orchestrator. No cluster scheduler, no cloud
control plane.
controller.
Defaults, not mandates. A second model provider is anticipated by design; nothing in the
domain may assume one vendor's credential lifecycle.
+61
View File
@@ -0,0 +1,61 @@
# Glossary — the words this repository uses, and the ones it stopped using
One name per thing. This page is the authority; where an older record says something else, that
record is being superseded, not this page. It exists because the terms kept drifting in
conversation — control plane / controller / master / hub for one thing, substrate / foundation for
another — and a mesh you cannot name precisely is a mesh two people describe differently.
## The mesh and its machines
- **node** — a machine in the mesh. There are 0..n of them, and each runs the host agent. A node is
just a machine that has joined; being one implies nothing about what it runs.
- **control-node** — the one node that also holds the `mesh-controller` seat. There is exactly one
per mesh. "control-node" is not a separate kind of machine — it is a node that additionally runs
the controller (and, today, the foundation). Lose it and the other nodes keep running what they
were last told; they simply cannot be told anything new.
- ~~master / slave~~, ~~hub / peer~~ — not used. The relationship is *controller and nodes*, and no
node is subordinate: a node applies declarations on its own and survives the control-node dying.
## What runs the mesh
- **controller** — the component that decides what each node should be, holds the mesh's records,
and tells nodes over the broker. Replaces **"control plane"** (borrowed from networking's
control-plane/data-plane, and opaque here).
- **mesh-controller** — the module that runs the controller. It **claims** the `mesh-controller`
seat at mesh scope, which is what makes it singular. Replaces the module name **`mesh-control`**.
(The git repository has been renamed `mesh-control` -> `mesh-controller` on the forge; the module,
container and image it produces are `mesh-controller`.)
- **foundation** — the store and the broker, raised at genesis before any module system exists.
Replaces **"substrate"** (a biology metaphor that landed for no one). The foundation is not a
third thing beside the store and broker — it *is* those two, named together.
- **store** — the one postgres server. It holds the controller's own context databases
(`inventory`, `identity`, `licences` — a context owns its store, [ADR 0008](../02-DECISIONS/0008-a-context-owns-its-store.md))
and every module's own database. One server, many databases — never one shared "mesh database".
- **broker** — the one lavinmq message bus. It carries the mesh bus on the `/` vhost and a vhost per
consumer that requires `amqp`.
## What the mesh stores and serves
- **package** — what code resolves when it is **compiled**: an npm/cargo/pypi dependency, by
**version**. Served by the **package-registry** (gitea). Only a builder talks to it.
- **artifact** — what the mesh delivers to a machine to **install and run**: an OCI image, by
**digest**. Served by the **artifact-store** (distribution). Every node pulls from it.
- These are two protocols, not one store being weak — see [ADR 0075](../02-DECISIONS/0075-two-stores-and-which-provides-what.md).
## How modules relate to the mesh
- **seat** — a named position at a scope (node / site / mesh) with a **capacity**. A capacity-1 seat
is exclusive (one holder); a higher-capacity seat is a **bench** (several holders coexist).
- **claim** — a module taking a spot on a seat. `claims: [{name, scope}]` in a manifest. A
mesh-scoped exclusive claim is how the mesh says "there is one of me". A foundation seat is
named after the server it guards: the `mesh-controller`, `postgres` and `lavinmq` modules claim
the `mesh-controller`, `mesh-store` and `mesh-broker` seats ([ADR 0079](../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md)).
- **provision** — a service one module `provides` and others `require`; the mesh resolves a provider
and wires the two with an endpoint and a credential. This is separate from seats: a provision is
a service you offer, a seat is a slot you occupy.
## How this page is kept
A new name for an existing thing lands here first, in the same change that introduces it in code. A
record under `02-DECISIONS/` keeps whatever word it was written with — those are immutable — so a
term retired here may still appear there, and the mapping above is how to read it.
+18 -1
View File
@@ -1,6 +1,10 @@
# Process — overview
How work moves through HQ, and who may do what. Every other document in this folder is a
How work moves through HQ, and who may do what. In industry terms this is **spec-driven
development, with provenance**: the decision is the why, the design doc is the spec, `code:`
names the implementation, and the lab beds are the conformance tests — and unlike the common
form, the chain itself is checked ([ADR 0080](../../02-DECISIONS/0080-the-development-cycle-is-checked.md),
[0081](../../02-DECISIONS/0081-a-decision-nothing-cites-is-not-yet-in-the-chain.md)). Every other document in this folder is a
playbook: trigger, who runs it, steps, outputs. Engineers and agents follow the same
playbooks; agents must not act outside them.
@@ -49,6 +53,19 @@ the expensive half.
| [03](03-issues.md) | Issues | Something is wrong — often with the owner unknown |
| [04](04-build-handoff.md) | Build handoff | A design is ready to be built |
| [05](05-constitution-sync.md) | Constitution sync | `how-we-build.md` changed a rule the mesh enforces |
| [06](06-writing-a-module.md) | Writing a module | Something that runs today must run on the mesh |
| [07](07-feature-branches.md) | Feature branches across repos | Work that changes code, in one repo or several at once |
## The cycle is checked
The flow above is a rule, and a rule states how it is checked:
[`00-META/checks/cycle.py`](../checks/cycle.py) refuses a to-be design that names no
decision, an `in-progress`/`implemented` design that names no owning code, an issue marked
`located`/`fixed` with no owner or `fixed`/`resolved` with no fix, and a `graduated`
research overview that does not say what it became. Run it with `records.py` and `index.py`
before any HQ merge. What the checks cannot see — that code work actually started from a
handoff — is held by playbooks [04](04-build-handoff.md) and [07](07-feature-branches.md):
a feature branch exists because a design or an issue sent it.
## Status lives in frontmatter
+3
View File
@@ -1,5 +1,8 @@
# Playbook 05 — Constitution sync
Implements [ADR 0021](../../02-DECISIONS/0021-hq-is-the-source-of-the-constitution.md): HQ is
the source of the constitution, and the knowledge-base page is derived, never edited.
**Trigger.** [`how-we-build.md`](../how-we-build.md) changed a rule that the mesh enforces at
runtime.
+107
View File
@@ -0,0 +1,107 @@
# Playbook 06 — Writing a module
**Trigger.** Something that runs today must run on the mesh, or a new capability must be
declarable.
**Who runs it.** Whoever is porting or writing it.
*Written 2026-09-01 from doing this for the first time end to end. Every step below exists
because skipping it cost something.*
## Before anything: read what runs
**A module is written from the thing, not from memory of the thing.** For a port, that means its
current compose file, its environment, and where its data actually sits. Assumptions about any of
the three have been wrong every time they were not checked.
Three questions, answered from the machine:
| | why it decides something |
|---|---|
| **what containers, and how do they find each other?** | more than one means a `network`; names between them must match what the software is configured to dial |
| **where is its data?** | a bind mount moves with a path; a named volume does not; an anonymous volume is already losing data on every redeploy |
| **which values are secret, and which are merely settings?** | a secret goes in `own-secrets` or a grant; a setting goes in the manifest and may be overridden per node |
## The steps
1. **Name what it provides and requires**, if anything. A name is what a consumer is coupled to,
not the role it plays ([ADR 0027](../../02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md)):
`postgres-database`, not `database`. Most modules provide nothing and require nothing — an
application is usually a leaf.
2. **Declare capabilities, not dependencies, for facts about the machine.** `container-runtime`,
`package-manager`, `seat`. A capability is detected and refused against; it is not something a
module can install.
3. **Write the resources in the order they must happen.** They are applied in the order written
and orphans are removed in reverse, so a `network` is written before the containers that join
it and removed after them.
4. **Put every secret in a file, never in `env`.** A declaration travels over the broker in plain
text: a password in `env` is a password the broker sees. **The mesh delivers parts; a module
that needs them combined combines them.**
**A sealed file holds the password and nothing else** — no key, no `=`, no newline that means
anything. So `env-file` must never point at one. It points at a file the module *declares*,
whose content leaves a hole:
```
own-secrets superuser → /var/lib/postgres/superuser.secret the password, alone
a file /var/lib/postgres/superuser.env, mode 0600,
content: POSTGRES_PASSWORD=${secret:superuser}
the container env-file: [/var/lib/postgres/superuser.env]
```
The host fills the hole on the machine, which is the only place both halves exist — the mesh
discarded the value
([credentials and their rotation](../../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md)).
A **provisioner** is the exception: it reads a password file, so it mounts the `.secret`
directly.
Every example module in `mesh-controller` had this wrong and shipped: `own-secrets` pointing at a
path *named* `.env`, mounted as `env-file`, holding a bare password. Docker reads that as a
malformed line and the container starts **with no password set at all** — not a failure to
start, a service running on the wrong credential. They parsed and they resolved. Two tests in
`examples/modules` now refuse both halves of it.
Add `restart-on` naming the env file, or the container keeps the credential it started with
through every rotation.
5. **Pin every image by digest.** A tag moves. The manifest in a repository names artifacts; the
manifest the mesh holds names digests, and they are not the same document.
6. **Decide generate or accept.** A new module's credential is generated. **An adopted one keeps
the credential it already has** — `secret accept` — because minting a new password for a
database that already exists locks the application out of its own data.
7. **Add a provisioner only if the software cannot read a file.** A proxy that watches a
directory needs nothing. PostgreSQL needs `CREATE ROLE`, an object store needs a bucket and a
policy, an identity provider needs a realm and a client — those need a small program beside
them. It reads what the mesh granted and reconciles; it does not decide anything.
8. **Prove it in the lab, against the real software.** Not that a container started — that the
thing works: the credential authenticates, a wrong one is refused, the containers reach each
other, the data survives a restart.
## What the first port actually cost
Six attempts, one real bug. Recorded because the ratio is the lesson: **the mesh was right every
time and the scaffolding was not.**
- A shape existed in the language and no host implemented it, so every declaration carrying one
was refused whole — correctly, and the host said exactly that. **Nobody was reading the host's
log.** Read it first; it is the only place that says why a machine did nothing.
- A blind find-and-replace renamed a provision in quotes and missed the same word bare.
- A command was tested only for the invocations that should fail, so it rejected every real one
and the suite stayed green.
- A test asserted on a helper rather than on the code that calls it, three separate times. **A
test that cannot fail when the behaviour is deleted is not defending the behaviour.**
## Rules
- **Read the host's log before theorising.** A declaration that was sent and not applied says so
there and nowhere else.
- **A failing test is kept, not skipped.** It is the reproduction.
- **Never rotate during an adoption.** Rotation is a separate act, afterwards, deliberately.
- **A data directory is never removed by the mesh** ([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)),
and that protects against the mesh only — not against a disk or a mistaken command.
+69
View File
@@ -0,0 +1,69 @@
# Playbook 07 — Feature branches across repos
**Trigger.** Work that changes code — in one code repo or in several at once (`mesh-sdk`,
`mesh-controller`, `mesh-catalog`, `mesh-host`, `mesh-lab`, and `hq` when a decision rides along).
**Who runs it.** Anyone who writes code, engineers and agents alike. Agents follow it exactly —
it is the guard against the failure it was written for.
## The failure it prevents
A feature was worked as a branch-and-MR per *unit of thought* — one per decision, one per
stacked increment — and each MR was treated as finished when it was *opened*, not when it was
*merged*. Across repos the same feature took a different branch name in each. The MRs piled up
unmerged: one session left **sixteen** stacked intermediate MRs that had to be consolidated and
closed by hand. An MR is a review checkpoint, not a scratchpad.
## The rule
One feature is **one branch name**, **one worktree per repo**, **one MR per repo**, opened
**once, at the end**.
1. **Name the feature once.** `feat/<slug>`. The *same* branch name in every repo the feature
touches — never a different name per repo, never a fresh branch per increment within the
feature.
2. **Isolate each repo.** One git worktree per touched repo under `.work/<slug>/<repo>`, branched
off `main`:
```
git worktree add .work/<slug>/<repo> -b feat/<slug> origin/main
```
Parallel features never collide, and no shared checkout is edited.
**Add the untouched siblings the lab reads.** The lab beds find the other repositories by
sibling path from the lab checkout (`../mesh-tools/module.json`, `../mesh-sdk`, …), the way
the main layout has them. A `.work/<slug>/` directory holding only the touched repos fails a
bed at once with `no manifest for mesh-tools at …/.work/<slug>/mesh-tools/module.json`, after
genesis has already passed. Give the directory those repos as **detached worktrees on `main`**
— never symlinks:
```
git worktree add --detach .work/<slug>/mesh-tools main
```
3. **Commit as you go — locally.** Increments land on the one branch. Nothing is pushed and no
MR is opened mid-feature.
4. **Finish, then publish.** When the whole feature is done — every repo, tests green — push
every branch and open **one MR per touched repo**, together.
5. **Merge promptly, once approved.** Every merge into `main` is notified and approved
([ADR 0023](../../02-DECISIONS/0023-approval-is-the-checkpoint.md)); once it is, merge —
do not leave it sitting. The branch is deleted on merge.
6. **Leave nothing behind.** After the MRs merge, no `feat/<slug>` branch and no `.work/<slug>`
worktree survive.
## What this is not
- **Not a licence to batch unbounded work.** A feature is a *bounded* unit; if it sprawls for
days, end-of-feature bloat merely replaces per-increment bloat. Split it into features, each
its own branch and MR.
- **Not a second trunk.** Every repo branches off `main`. There is no longer an `initialization`
trunk.
## How it is checked
The end state is visible, and its absence is the smell:
- After a feature merges, `git branch -r | grep feat/<slug>` and `git worktree list` return
nothing for it. A surviving branch or worktree means step 6 was skipped.
- A bed run from `.work/<slug>/mesh-lab` that fails naming a `.work/<slug>/<repo>/…` path it
cannot find is a missing sibling worktree, not a mesh fault.
- More than one open MR in a repo that share no feature name, or a stack of MRs none of which is
merged, is the failure this playbook exists to prevent — stop and consolidate before opening
more.
+10 -7
View File
@@ -1,6 +1,6 @@
---
status: canonical
updated: 2026-08-23
updated: 2026-09-24
---
# The Novox repositories
@@ -17,21 +17,24 @@ and a forge address is an operational detail (see [`README`](../README.md)).
| `hal` | The monorepo — the node runtime, the module catalogue, the delivery machinery, and the bootstrap scripts. Every core module lives here. |
| `hq` | This repository, under the company organisation — mission, research, design, decisions, issue diagnosis. Company-scoped ([ADR 0019](../02-DECISIONS/0019-how-this-repository-works.md)); the mesh is its first product. The source of truth for *why*. Carries no implementation. |
| *(one per application)* | Every standalone application, site or side-project gets its own repository, with `module.yml` at the root. Registered with the mesh as a build source; built and deployed by the same pipeline as anything in the monorepo. |
| `migration` | **Private.** The record of one installation replacing the predecessor mesh with this one: the runbook, a dated log of every step and what it cost, the per-service data procedures, the readiness checks, and the scripts. Private because it is the opposite of this repository in every way that matters — it names machines, addresses, ports and paths, because a procedure that cannot be followed is not one. Where hq asks *what did we decide and why*, that repository answers *what happened on the machines, in what order, and what to do next*. Its `HANDOFF.md` is where somebody picking the work up starts. |
## What the mesh becomes
[ADR 0019](../02-DECISIONS/0019-how-this-repository-works.md) records the repositories the
monorepo decomposes into. **`mesh-lab`, `mesh-host` and `mesh-control` exist so far** — the lab is built first
([ADR 0016](../02-DECISIONS/0016-the-lab.md)); the rest are the
target, not the present.
monorepo decomposes into. **`mesh-host`, `mesh-controller`, `mesh-catalog`, `mesh-lab`, `mesh-sdk` and `mesh-tools` exist
so far** — the lab was built first ([ADR 0016](../02-DECISIONS/0016-the-lab.md)). The tiered
decomposition below is the planned shape; the repositories built to date do not map onto it
one-for-one — `mesh-catalog`, `mesh-sdk` and `mesh-tools` exist where the table names
`mesh-foundation` and `mesh-surfaces`, and reconciling the two is itself still ahead.
| Repository | Tier | Holds |
|---|---|---|
| `mesh-host` | 0 | **exists.** The node host — one statically linked binary, requiring nothing present ([ADR 0005](../02-DECISIONS/0005-the-node-host.md)) |
| `mesh-substrate` | 1 | the four pinned services, as declarations |
| `mesh-control` | 2 | **exists.** The control plane and its contexts — one of seven built ([ADR 0006](../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)) |
| `mesh-foundation` | 1 | the four pinned services, as declarations |
| `mesh-controller` | 2 | **exists.** The controller and its contexts — one of seven built ([ADR 0006](../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)) |
| `mesh-surfaces` | 3 | tools, web, cli |
| `mesh-sdk` | — | contracts shared across tiers |
| `mesh-sdk` | — | the stable spine modules build against — the tool-serving harness, the messaging/event framework, the contracts and core primitives. Holds nothing per-module and nothing volatile ([ADR 0039](../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md)). |
| `mesh-lab` | — | **exists.** The lab — scenario lifecycle, networking, placement. Ships to nobody; runs on a workstation. |
Tier 4's shape is open, and deliberately so: see ADR 0019 and
@@ -3,7 +3,6 @@ status: graduated
initiated: 2026-08-22
touches: [03-DESIGN/00-as-is/05-runtime-and-installation.md]
became:
- 02-DECISIONS/0005-the-node-host.md
- 02-DECISIONS/0005-the-node-host.md
- 03-DESIGN/01-to-be/05-the-node-host.md
---
@@ -90,5 +90,5 @@ the catalogue where modules genuinely change together under one intent. The skel
| One repository per tier, or per context? | Already open from ADR 0001 as "catalogue destination — one repository or many". The skeleton assumes per tier and does not settle it. |
| ~~Does an unprivileged node earn a place in the inventory, or only a presence?~~ | **Answered 2026-08-25** by the operator: a node is a *managed machine inside the mesh*, not an unprivileged something — and a disconnected node is still a node, in a different situation. The question posed a class distinction; the answer is that there is none, and what varies is **state**. Recorded as [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md). |
| ~~Does absorbing overlay, filtering, packages, supervision and the container runtime make the host too large?~~ | **Answered 2026-08-25** — [`host-size.md`](host-size.md). Measured: the absorption is smaller than the machinery that already applies state, and eight of ten adapters already carry no dependency. The risk is not size but direction, and it is two modules wide. The claim survives with its scope corrected — the host carries one concern, *apply declared state on this machine*, of which the six are instances. Recorded as [ADR 0005](../../02-DECISIONS/0005-the-node-host.md), designed in [`05-the-node-host.md`](../../03-DESIGN/01-to-be/05-the-node-host.md). |
| ~~Four substrate services or five?~~ | **Answered conditionally**, which is the honest form — [`07-the-substrate.md`](../../03-DESIGN/01-to-be/07-the-substrate.md). The substrate is *what the control plane consumes and cannot grant itself*. The identity provider qualifies only if the control plane delegates authentication; if it authenticates natively it is an ordinary hosted service. The count follows from a decision not yet taken, and asserting four was asserting that decision. |
| ~~Four substrate services or five?~~ | **Answered conditionally**, which is the honest form — [`07-the-foundation.md`](../../03-DESIGN/01-to-be/07-the-foundation.md). The substrate is *what the control plane consumes and cannot grant itself*. The identity provider qualifies only if the control plane delegates authentication; if it authenticates natively it is an ordinary hosted service. The count follows from a decision not yet taken, and asserting four was asserting that decision. |
| Does `feature` survive? | The skeleton splits it in two and argues the conflation is what makes the delivery pipeline hard to reason about. Unproven. |
@@ -4,8 +4,8 @@ initiated: 2026-08-25
became:
- 02-DECISIONS/0009-modules-and-the-graph.md
- 02-DECISIONS/0008-a-context-owns-its-store.md
- 03-DESIGN/01-to-be/06-the-control-plane.md
- 03-DESIGN/01-to-be/07-the-substrate.md
- 03-DESIGN/01-to-be/06-the-controller.md
- 03-DESIGN/01-to-be/07-the-foundation.md
touches:
- 02-DECISIONS/0009-modules-and-the-graph.md
- 02-DECISIONS/0009-modules-and-the-graph.md
@@ -7,6 +7,9 @@ touches:
- 03-DESIGN/01-to-be/05-the-node-host.md
- 03-DESIGN/00-as-is/05-runtime-and-installation.md
- 01-RESEARCH/011-the-module-graph/00-overview.md
- 02-DECISIONS/0078-the-store-and-broker-are-modules.md
- 02-DECISIONS/0088-the-foundation-filters-before-anything-listens.md
- 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
---
# 012 — The minimum viable node, and adopting what is already there
@@ -180,6 +183,16 @@ same makes a node where something the mesh needed never happened indistinguishab
where a log level differed. Whether a failed line still lets adoption complete is therefore
reopened by adding severity, and is not decided here.
## Adoption as a mode, for the migration
*2026-09-22.* The migration from the predecessor mesh gave the middle state a length. A machine
running the predecessor is **adopted** when the mesh comes up on it — the predecessor's control
stopped, the machine's firewall and files kept in force, the mesh opening what it needs through
them — and stays adopted while its modules migrate one at a time, until the operator **converges**
it. Measured on the control-node, and weighed against the alternatives, in
[*migrating a node that is in use*](migrating-a-node-in-use.md); decided in
[ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md).
## Open questions
| Question | Why it is open |
@@ -191,6 +204,7 @@ reopened by adding severity, and is not decided here.
| Can a module say which of its settings are load-bearing? | The question that dissolves the conflict rule rather than choosing a side. A setting the module *requires* cannot be kept from the machine without producing something installed and broken; a setting it merely *prefers* should always yield. Until a module can say which is which, adoption is defaulting in the dark. Belongs with the graph. |
| Does a `failed` line still let adoption complete? | *Flags inform, they do not block* was decided about conflicts, where the mesh chose and the machine works. A failure is *we could not*, which is different in kind — and treating them alike hides the worse one behind the commoner one. |
| How is a flagged conflict reconciled, and by whom? | The briefing hands it to a session. What that session is empowered to change, and whether the resolution is recorded so the next adoption does not re-raise it, is undecided. |
| ~~How long is a machine adopted?~~ | **Decided** ([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)) — for as long as it is being migrated: a mode per node, recorded, ended by an explicit and previewed flip. |
| Where does the kept original live, and for how long? | Whether it is recorded in the node's state so adoption is visibly reversible, and whether it is returned when the mesh stops managing the thing. |
| What shape is a briefing? | Structured enough to be acted on, prose enough to be read. It is the first thing a session on a new node sees, which makes it an interface rather than a log. |
| Does owning a package mean owning its version? | Owning configuration and owning the package are different scopes. The second means the mesh decides which version is installed, and that decision then has to survive the machine's own package manager updating it. |
@@ -0,0 +1,154 @@
# Migrating a node that is in use — adoption as a mode, not a moment
*2026-09-22. Measured on the machine that will be the control-node, which is running the
predecessor mesh today. Nothing below names it; the counts are its own.*
## The question
The mesh replaces a predecessor mesh that is running, on the same machines, with the services
people use. The control-node is decided: it is the machine that already carries the predecessor's
broker and build pipeline. So the question is not *where* the mesh starts but **how a machine
running the predecessor becomes a node of the mesh without its services noticing** — and then how
each of the other machines follows.
## The operator's proposal
Proposed by the operator, 2026-09-22, and the shape this document tests:
1. **Stop the predecessor's control on a machine** — its daemons that write configuration: the
network and firewall configuration above all, which decide what is reachable and what is
blocked. Its services keep running; only the control over their configuration stops.
2. **Bring the mesh up on that machine in adoption mode.** It takes custody of those files and
keeps what it finds in force.
3. **Migrate the modules one at a time**, data preserved, per the cutover procedure.
4. **Move to the next machine and repeat** — adopted first, keeping its local configuration, then
migrated.
5. **When every machine is migrated, flip adoption mode**, and the mesh takes full control of the
configuration it has been holding.
This is the research above made concrete. It keeps the conflict rule already decided here — *on
conflict, what is on the machine stays* — and gives the middle state, *adopted*, a length: not a
one-time import before generating starts, but a mode that lasts for as long as the machine is
being migrated, ended by an explicit act.
## What the machine actually looks like
Measured, read-only:
| | |
|---|---|
| Containers running | 60, all the predecessor's services and their stores |
| The predecessor's control | user-level daemons, separate from the services; none of the 20 running system services is the predecessor's control |
| Firewall | the predecessor's, active: 54 incoming rules and 52 forwarding rules, each served port allowed explicitly — how a default-deny firewall reads |
| Files the predecessor's configuration sync writes | 12, of which 3 are system files (an ssh server drop-in, the package manager's configuration, one service's configuration); the rest are the operator's shell and agent files |
| Per-service configuration | environment and composition files per service, written by the predecessor's service tooling |
**Stopping the predecessor's control stops nothing that serves.** The services are containers and
system units that run without it; what stops is the rewriting of their configuration. Nothing
changes on the machine until something else writes.
## What collides, measured rather than assumed
An earlier note assumed the mesh's foundation could not stand beside the predecessor because both
want the store's and the broker's standard ports. **The measurement says otherwise.** The
predecessor publishes its own store on a non-standard port and its broker on another; the standard
ports the foundation binds for its store and its bus are free.
What does collide:
| The mesh wants | Held by | When |
|---|---|---|
| the registry's port | the predecessor's registry | at genesis — the foundation raises a registry |
| the broker's management port, on loopback | the predecessor's broker | at genesis |
| the private network's port | the predecessor's own tunnel | at genesis on a control-node that is the private network's hub; otherwise when the node is placed on it |
| the web ports | the predecessor's reverse proxy | when the route proxy is assigned |
| the resolver's port | a resolver the predecessor runs | when a resolver module is assigned |
On the control-node, which is the private network's hub, three of these fall at genesis and cannot
be deferred; two only when a particular module is taken, the moment its predecessor stops anyway.
Today the foundation's ports are **fixed**: written in the installer's bundle and in the
catalogue's manifests, so a collision is found when a container fails to bind, not before — and a
port changed at genesis would be changed back when the foundation is adopted as modules, since the
applier recreates a container whose declared spec differs. The private network's address range
must also stay clear of the range the predecessor's tunnel uses; on the machines measured they are
distinct.
## What would break if the mesh came up as it is today
**The firewall.** Genesis loads a base ruleset — drop anything undeclared — in its own table
([ADR 0088](../../02-DECISIONS/0088-the-foundation-filters-before-anything-listens.md)). The
predecessor's firewall is a different table. The kernel runs every base chain registered at the
same hook, in priority order: an accept ends only its own chain and the packet goes on to the next,
and a drop in any is final — whether the other firewall's chains are nftables or legacy iptables.
So the
mesh's base ruleset would drop everything the predecessor's firewall allows and the mesh has not
declared — every web, mail and database port in the table above — the moment genesis ran.
**The files.** The host writes a declared file whatever it finds at the path, reporting it as
updated. A file the predecessor left — the ssh server drop-in, the resolver's configuration — is
replaced the first time a module declaring that path is assigned, before that module's service
has moved.
**The container names.** The applier keys a container on its name. A catalogue module whose
container carries the same name as the predecessor's service it replaces takes that container over
the moment it is assigned: today, *assigning a module is migrating it*, never a preparation.
**The published ports.** The foundation publishes its ports on every interface, and a published
container port reaches the container through the forwarded path, not the incoming one. A firewall
that filters only incoming traffic never sees it. The base ruleset is what keeps the store
unreachable from outside today — and it is the thing that cannot be loaded on this machine.
## The pipeline freezes while the control-node migrates
The predecessor's build pipeline and its coordinator run on the control-node. Stopping its control
there stops the predecessor's updates for **every** machine it manages, until the migration is
done. Their services keep running; they receive nothing new. That is the price of the proposal and
it is worth stating, not a reason against it: the migration is the period in which the predecessor
is being replaced, and it does not need to keep changing.
## What adoption mode has to mean
For the proposal to hold, *adopted* must be a **state the mesh records per node**, not an
intention, and each thing the mesh would otherwise take must say what it does in that state:
- **What is found is kept until its module is taken.** *Found* is precise: present at a declared
path or name with no record in the host's store. A found file or container is held, its original
recorded, until the operator **takes** the module on that node — the cutover, done when the
module's data has moved. Assigning prepares; taking migrates. Without the distinction the rule
never fires: the host only ever sees what assigned modules declare.
- **The firewall found on the machine stays in force.** The mesh loads no table on an adopted node
that drops by default or holds an accept. What it needs open it declares as openings the host converges
*through the found firewall*, on the incoming and the forwarded path, marked as the mesh's and
re-checked on every reconcile so a reload or reboot does not lose them. An accept in a table of
its own would not help: the found firewall's drop would still be final. What a table of its own
*can* do is refuse, and a refusal is final too — so the mesh guards the store and the broker's
management port from everyone but the private network and the machine itself, in a table that
only refuses, ahead of the container runtime's redirect — which the found firewall does not do
and cannot undo. The bus, the registry and the hub's port stay open to
anywhere: a node enrols before it has a private-network address.
- **The foundation's ports are the node's to give** — set at genesis, checked free, and kept as
that node's settings, read everywhere they are used, so adopting the foundation as modules does
not move them back.
- **The flip is per node and previewed** from what is actually reachable — listening sockets and
published ports, not the found firewall's allow list, which does not see what a container runtime
forwards.
## Options weighed
| Option | Verdict |
|---|---|
| Cut the machine over in one go: stop the predecessor's store, broker, registry and proxy, raise the foundation in their place | Rejected. Every predecessor service goes down until it has migrated, and the predecessor's other machines lose their broker. The rollback is restarting the predecessor, which is a recovery, not a step. |
| A separate machine as the control-node | Rejected. Contradicts the decision that the control-node is the machine that already carries the predecessor's broker and pipeline. |
| Converge on joining, as the mesh does today | Rejected. The base ruleset closes every predecessor port at genesis, and found files are replaced before their services move. |
| Make only the foundation's ports configurable and otherwise converge | Rejected as insufficient. It solves the bind collisions and none of the firewall or file ones. |
| Treat assigning a module as migrating it | Rejected on review. The rule that keeps found files would never fire — the host only sees what assigned modules declare — and an assignment on an adopted node would be an outage rather than a preparation. |
| **Adoption as a mode, per node, ended by an explicit flip** | The proposal. It is the conflict rule already decided here, given a duration. |
## What this leaves open
Answered by the decision record this feeds — [ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)
— only for the migration's needs. The general questions above stay open: whether a module can say
which settings are load-bearing, what shape a briefing takes, how a flagged conflict is reconciled.
One is newly sharp: the mesh opening ports through a firewall it did not install needs to speak that
firewall. There is one kind on the machines measured; a machine with another is not covered until
someone writes for it.
@@ -0,0 +1,54 @@
---
status: active
initiated: 2026-09-23
touches:
- 02-DECISIONS/0075-two-stores-and-which-provides-what.md
- 02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md
- 02-DECISIONS/0071-genesis-builds-from-a-mesh-that-already-exists.md
- 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md
- 04-ISSUES/085-the-packages-port-given-at-genesis-is-not-a-setting/00-report.md
- 04-ISSUES/090-the-forge-module-does-not-take-over-the-forge-genesis-raised/00-report.md
---
# 013 — The forge and the registries: what a seat is for, and what it is not
**The question, asked during the first migration:** the mesh runs a forge that serves git and
packages, an image registry, and a bootstrap forge genesis raises before any module exists. Should
there be a mesh-scoped **git seat**, the way there is one for the store and the broker?
**The short answer: no, and the question points at a different hole.** A seat answers *is there
exactly one of you*. Every symptom around the forge is about **an address and a credential nobody
resolves**, which a seat does not answer. The mechanism that would answer it — a provision, with
`serves` and a grant — already exists, is already used for the image store, and is already
declared for packages. It simply has no consumer: the one module that needs it carries a literal
instead.
See [the survey](the-survey.md) for what the code does today, with evidence.
## What was found
- **A seat does exactly two things**: it refuses a second claimant at resolution, and in one place
it answers *where is the broker*. It is a resolution-time predicate; no host ever hears of it.
- **Four mesh-scoped seats exist**, all named after a singular server the mesh runs **for its own
working** — controller, store, broker, catalogue. A forge is an application the world made, not
part of that set.
- **Nothing in the mesh requires git.** There is no provision for it, and the forge's git service —
over https and ssh — is declared nowhere. Only its npm half is a provision.
- **`package-registry` is a fully built provision with zero consumers.** The builder, its only
real consumer, bypasses it with a file naming the module, the address `127.0.0.1` and the port.
- **The builder is told where source lives, per build, by a person.** The clone URL is an opaque
string; each module remembers its own; no credential is ever attached; nothing polls a forge.
- **A seat would forbid something the mesh should allow**: a second forge — one for the mesh's own
source, one for something else — which ADR 0075 already argues for against a node-scoped claim.
## What this leaves open
- Should the forge's git service become a **provision** (`source-forge`), so the builder resolves
the address and a credential the way it already resolves the image store? That is what would let
a forge move, be renamed, or be replaced without editing every module's stored URL, and it is the
unanswered half of [issue 085](../../04-ISSUES/085-the-packages-port-given-at-genesis-is-not-a-setting/00-report.md).
- The obstacle is known and written down: such a requirement **may go unanswered during genesis**,
and the mesh has no optional requirement. Deciding that is the real work.
- Should the mesh learn when a source moves? It records the fact and never discovers it: nothing
polls, and there is no receiver for a forge's push. `build --behind` answers a question only a
person can currently make true.
@@ -0,0 +1,130 @@
# The survey — seats, the forge, the registries, as the code has them
*2026-09-23. Read from the code and the records, not from memory. Every claim here was checked
against a file.*
## What a seat is
The glossary calls a seat *a named position at a scope with a capacity*, and a claim *a module
taking a spot on it*. **Capacity does not exist in the code.** There is no capacity field and no
bench; every claim is exclusive. The "shared seat" of the design is the word `provides` wearing
that name.
Claiming does two things and no others:
- **It refuses a second claimant when a node's modules are resolved** — within one node at any
scope, and against every other node for mesh and site scope.
- **It answers one question, once:** the controller finds where the broker is by looking for the
module claiming the broker's seat.
No host is ever told about a seat. It is a resolution-time predicate.
Two mechanical asymmetries matter. A mesh-scoped seat's exclusivity is defended only by nodes
that *resolve*: a node whose modules fail to resolve contributes nothing, so a broken node does
not hold its seat against a second claimant. And a site-scoped seat does nothing at all when
either node has no site.
## The seats that exist
Twelve claims, twelve names. Four are mesh-scoped: the controller, the store, the broker, the
catalogue. All four are things the mesh runs **for its own working**, and three of them are raised
by genesis and adopted in place — which is the argument of the record that named them.
The node-scoped ones split in two: a **role on this machine** (the build machine, the packet
filter, the intrusion prevention, the private network) and **a scarce machine resource** (the
resolver's configuration file, the DNS port). The second kind is a claim for the reason a port is:
there is one of it on the machine.
Two oddities worth stating:
- The image store claims its seat at **node** scope while providing its service at **mesh** scope.
The record that discussed this left open whether it should hold that claim at all.
- **Nothing claims the two public ports.** The proxy binds them and claims nothing, which is the
hole the route handover fell into.
## The forge, and what it serves
The forge serves three things: git over https, git over ssh on its own port, and an npm registry.
**Only the npm half is in the mesh's vocabulary** — the forge provides `package-registry` and says
where it answers. Git is served and declared nowhere: no provision, no `serves`, nothing that can
require it. The only trace is the public label in its route contribution, which the mesh is
explicitly not meant to interpret.
The image store is a separate module and a separate provision, plain HTTP, trusted because it is
reachable only over the private network.
A second module also provides `package-registry`. Two providers of one provision is not a refusal
but an ambiguity, and the mesh already has the command that settles it: pin one.
## Where the builder's three addresses come from
| What it needs | How it finds it |
|---|---|
| the broker | a sealed secret of its own |
| the image store | **a real provision binding**, resolved by the mesh |
| the package registry | **a file in its own manifest**, naming the module, `127.0.0.1` and the port |
| the source repository | nothing at all — a person types a clone URL per build |
The installer says this plainly in a comment: *the one binding nothing resolves: the builder dials
the forge by a number it carries.* Genesis can override the port of that literal, and nothing
else: not the forge's identity, not its address.
`package-registry` therefore has **zero consumers**. The builder does not require it. The
provision is fully built — two providers, grants, a served description including the registry's
path — and nothing asks for it.
## How a module's source is found
It is not found; it is supplied. The clone URL arrives as an argument, travels to the build machine
and reaches `git clone` unexamined. No credential is attached, so a private forge works only if the
build machine's own git configuration already authenticates. Each module remembers the URL it was
built from, so there are as many forge addresses as modules, each frozen at whatever was typed.
At genesis there are four independent source URLs, one per flag, unrelated to each other.
If the forge moves or is renamed: every stored URL goes stale independently, each failing at its
own clone, and there is no command to re-point them. Genesis keeps working, because it takes URLs
as flags — which is the existing record's position: the forge is reached *by a name outside the
mesh*, and failover is repointing that name.
## What genesis's forge is
A bare container run before any module exists, because the toolchain resolves the mesh's own
packages by version from an npm registry. It holds **no seat, no manifest, no provision**. The
mesh is told one thing about it — a port — and only when that port is not the default. **On an
ordinary genesis the mesh never learns the bootstrap forge exists.**
Its successor module differs from it in five ways: container name, network, data directory, the
address it is told to call itself, and — contradicting a comment that says they are pinned
identically — **the image digest**. So assigning the module raises a second forge beside the first.
A seat cannot fix that. A seat is checked between manifests; the bootstrap forge has no manifest,
so no seat can see it. The check that is actually wanted — same name, same data, same network — is
between an installer constant and a manifest.
## Would a git seat be coherent?
In form, yes: add the claim and a second forge is refused. In effect it buys one refusal nobody
has hit, and forbids an arrangement the mesh should allow — a second forge for something other
than the mesh's own source.
Measured against the four mesh seats, it does not fit: those are singular servers the mesh runs
for itself, and a forge is an application. Nothing would *use* it, because the only seat consumer
answers "where is the broker", and a forge address is not one fact: a module's source is a
repository **and a path and a ref**.
## What the code already supports and nobody uses
- **`package-registry` with no consumers.** The builder's literal file is a hand-rolled copy of
the binding the mesh would have written for it.
- **Pinning a provider** already expresses *this* forge, of several, without forbidding the second.
- **Grants** on the forge are the credential mechanism that would hand the builder an account.
Today the installer creates that account directly, because nothing requires the provision and so
the grant has no consumer to write for.
## The reading
The question *should there be a git seat* is the wrong shape for the problem under it. Every
symptom — the literal address, the port that would not follow a node, the takeover that is not a
takeover — is about **an address and a credential nobody resolves**. A seat resolves neither. The
provision mechanism does, and is already built.
@@ -0,0 +1,108 @@
---
status: graduated
became: 02-DECISIONS/0106-the-bus-is-nats.md
initiated: 2026-09-23
touches:
- 02-DECISIONS/0002-nodes-communicate-over-a-broker.md
- 02-DECISIONS/0033-the-substrate-is-a-store-and-a-broker.md
- 02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md
- 02-DECISIONS/0041-events-are-a-relationship.md
- 02-DECISIONS/0042-the-shape-of-an-event-on-the-wire.md
- 02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md
- 02-DECISIONS/0078-the-store-and-broker-are-modules.md
- 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
---
# 014 — The bus on NATS: replace the broker, and when
**The question, asked mid-migration:** the operator wants the mesh's bus — today an AMQP broker —
replaced by NATS, with everything the broker does today. Do we finish the migration on AMQP and
move to NATS after, or go to NATS directly?
**The short answer: decide NATS now, build it in the lab in parallel, and cut the mesh's bus over
in one rehearsed rollout — after the migration's core is done and never underneath it.** "First
everything on AMQP" is already the state and costs nothing more: the predecessor's broker was
merged into the mesh's tonight, and every module converted from here on targets the sdk's broker
contract, which names no protocol. The only code that speaks AMQP is the mesh's own, in three
places, and it is swapped once.
## What was measured
**Where AMQP is spoken** (non-test files): the controller's `internal/link` (5 files, ~2.4k lines
with tests) and the builder's main; the host's `internal/link` (~1.5k); the tool runtime's
`broker-amqp.ts` (~600 with its main). **The sdk speaks none** — its `Broker` is `request`,
`handle`, `publish`, `subscribe`, `close` ([ADR 0039](../../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md)
put the client in the runtime for exactly this). **Every module's tools and events go through the
sdk**; one module (`amqp-ping`, a probe) talks AMQP on purpose.
**What the bus carries:** a control queue (`control`, `.upgrades`, `.catchup`), one queue per node
(`node.<name>`), `builds`, the enrolment/report/alive/built flows, a topic exchange of events
(`mesh.events.<module>.<event>`, dead-lettered to `mesh.events.dead`), an RPC exchange (`mesh.rpc`)
and per-tool service queues (`serve.<module>.<tool>`), plus an MQTT exchange.
**What it relies on:** TLS for the bus a node enrols over; prefetch with reject/nack so the control
queue **holds messages unacknowledged while the store restarts** and retries
([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)); publisher confirms and
mandatory routing; a dead-letter exchange; **per-module accounts scoped by `emits`/`consumes`**
([ADR 0043](../../02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md)),
minted through the broker's management API; the broker as a module holding the `mesh-broker` seat
([ADR 0078](../../02-DECISIONS/0078-the-store-and-broker-are-modules.md)).
**The predecessor's world**, now on the same broker: ~50 AMQP connections per machine from four
machines, 595 queues, four vhosts, six users. All of it AMQP, all of it retiring module by module.
NATS speaks no AMQP: that world cannot move; it does not need to.
## How NATS answers each of those
| The mesh relies on | NATS | Note |
|---|---|---|
| topic routing keys | subjects with wildcards | same shape (`mesh.events.>`) |
| RPC and tool invocation over an exchange + reply queue | request/reply, native | simpler than today |
| competing consumers | queue groups | same |
| durable control queue, hold-unacked-and-retry, catch-up | JetStream: streams, durable consumers, ack/nak with delay, replay | the 0083 guarantee moves to JetStream; core NATS alone is at-most-once and would not do |
| dead-letter | max-deliver + advisories, or a stream fed from them | different mechanism, same effect |
| TLS bus | TLS | same |
| accounts scoped by emits/consumes, vhosts | accounts (isolation) with users and per-subject publish/subscribe permissions | stronger than today; the controller writes an auth config the host declares and the server reloads, instead of calling a management API |
| management API | `nats` CLI and an HTTP monitoring endpoint | no vhost concept — accounts instead |
| MQTT | built in | same |
| a module seat `mesh-broker` | unchanged — the seat is the server, the module changes | ADR 0079 |
| multi-node | clusters and leaf nodes | not needed now; a leaf per node is a later question |
Nothing the mesh needs is missing. The differences are in the shape of durability (JetStream must
be declared, streams and consumers are objects) and of accounts (configuration, not API calls).
## The cost, honestly
The bus is the mesh's nervous system. Moving it means: a new `nats` module in the catalogue taking
the `mesh-broker` seat; the controller's and the host's link packages rewritten; the tool runtime's
client swapped behind the unchanged contract; enrolment, reports, builds and the guard's ports
re-derived; the lab beds that prove the bus (store window, enrolment, upgrades) re-run on the new
one; and eight decisions amended or superseded. Weeks, not days — and none of it can be done
halfway on a live mesh: the controller, every host and every tool runtime move together.
## The sequencing question, answered
1. **NATS first, directly.** Stalls the migration for the length of the build; the mesh's bus
changes under a half-migrated node; the predecessor's clients still need AMQP, so a second
broker runs anyway, plus a bridge for whatever crosses. Rejected.
2. **Finish on AMQP, NATS after.** Wastes nothing — no module written from here on speaks AMQP —
but leaves the decision unmade while modules are written, and the runtimes idle. Adequate.
3. **Decide NATS now; build it in the lab in parallel; cut the mesh's bus over in one rehearsed
rollout after the core is migrated.** The predecessor's clients never notice: their broker is
the one the mesh adopted, kept as a compatibility module with an end date — the day the last
AMQP client is gone. **Recommended.**
## What this leaves open
- The **decision itself**, as a record: the bus is NATS; the AMQP broker becomes the predecessor's
compatibility broker and retires with the last AMQP client. Written when the operator says so.
- **JetStream's shape for the control plane**: one stream per concern (control, nodes, builds,
events) or one with subjects; retention; what the store window guarantee looks like as ack-wait
and nak-delay. Measured in the lab, not designed on paper.
- **Accounts as configuration**: the controller writes users and permissions into a file the host
declares, reloaded on change — which is the [ADR 0102](../../04-ISSUES/102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md)
discipline applied from the start — or the JWT/operator model. The first is simpler and matches
how the mesh already writes everything.
- **The MCP bridge** the operator asked for the same day is written against the sdk's contract, so
it moves with the bus and is not written twice.
- Whether a **leaf node per machine** replaces the hub-and-spoke bus later — out of scope here.
@@ -0,0 +1,119 @@
---
status: active
initiated: 2026-09-24
touches:
- 02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md
- 02-DECISIONS/0033-the-substrate-is-a-store-and-a-broker.md
- 02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md
- 02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md
- 02-DECISIONS/0078-the-store-and-broker-are-modules.md
- 02-DECISIONS/0084-which-provider-serves-a-consumer.md
- 03-DESIGN/00-as-is/03-provisioning.md
- 03-DESIGN/01-to-be/07-the-foundation.md
- 04-ISSUES/113-the-object-stores-images-were-withdrawn-upstream/00-report.md
---
# 015 — The object store after MinIO: which S3 implementation, and how the data moves
**The question.** The mesh's object store is MinIO. Its community edition is archived upstream,
its server and client images have been deleted from every public registry, and the pinned release
is four and a half years old and will never be patched
([issue 113](../../04-ISSUES/113-the-object-stores-images-were-withdrawn-upstream/00-report.md)).
Which S3-compatible implementation replaces it, and what is the migration track for the data and
the provisioning model that sit on top of it?
**Why now, and why not sooner.** Nothing is on fire: nodes that already hold the images keep
running, and issue 113 establishes that the deploy path tolerates an unfetchable-but-present
image by design. The forcing function is not an outage but a one-way door — **no node that does
not already hold the images can ever provision the module again**, so the mesh's ability to stand
a node up from its declarations is already broken for this module, and silently.
**The direction is not a departure from the design; it is the design.** The foundation document
already states the commitment:
> The dependency is on the **protocol**, not the product: AMQP for the bus, S3 for the object
> store, the OCI protocol for the registry. That is what keeps the naming safe rather than a
> commitment that cannot be revisited.
The object store is also **not** a foundation service — ADR 0028 removed it, and it is an
ordinary module required through the module graph by whatever wants one. (The "exception that is
not a swap" in that passage is the relational store, whose provisioning model borrows PostgreSQL's
own meaning of databases, roles and schemas. The object store carries no such coupling: a bucket
is a bucket.) So this effort is an instantiation of an existing principle, not a redesign — which
is the cheapest kind of decision to make and the strongest kind to cite.
## What the replacement has to carry, measured
Taken from the module's manifest, its composition, its tool surface, and a search for its
consumers across the catalogue — not from assumption.
| Requirement | Evidence in the module today |
|---|---|
| S3 API | The protocol every consumer speaks; already the design's stated dependency. |
| **OIDC login against the mesh's identity provider** | Six configuration variables are wired and populated in practice — discovery URL, client id, client secret, scopes, display name, redirect — plus a dedicated entrypoint script that blocks startup until the provider answers. This is live, not aspirational. |
| Erasure-coded multi-node topology | Four server nodes with two data directories each, behind a load balancer. |
| A single-node form | Declared as a flavour, for development and small nodes. |
| Buckets as a typed provision | The module declares a provision type of `bucket` on a named network; the mesh mints the credential and the provider creates it (ADRs 0048, 0084). |
| A tool surface | Bucket create/list/delete, object list/info/delete, presigned URL, and provisioning. |
| A console | Published on its own subdomain through the reverse proxy, with an unlimited request-body middleware for uploads. |
**Consumers, counted:** one application module, one capture module that takes a private bucket per
node, one workflow module's tools, and the delivery/rescue internals of the shared library. The
surface is small — the cost is concentrated in the provisioning handler, the tool handlers and the
OIDC story, not spread across the catalogue.
## Candidates
Scoped to **SeaweedFS** as the primary, with the others recorded so the rejection is not
rediscovered.
- **SeaweedFS** — Apache-2.0, Go, twelve-plus years of development, erasure coding, and OIDC
support in its S3/STS layer. Chosen to scope because it is the only candidate that plausibly
preserves the OIDC requirement above, which is the one requirement that is live and least
substitutable.
- **Garage** — the lightest to operate and the simplest model, but **no native identity-provider
integration**. Adopting it means losing OIDC console login or fronting it with a proxy. A real
functional regression against something currently in use.
- **RustFS** — markets itself as a binary-level drop-in retaining existing data, buckets and
configuration, which would make the data migration close to trivial. Young, and that claim is
exactly the kind that must be verified on a copy before it is believed.
- **Ceph RGW** — the most capable and the most operationally expensive; disproportionate to a mesh
where the object store is an ordinary module, not a platform.
**The first thing to verify, because the choice turns on it:** how much of SeaweedFS's OIDC story
is in the freely licensed build, and whether its shape — IAM/STS token exchange — can actually
stand in for a console that redirects a human to an identity provider. If it cannot, the honest
finding may be that **no** candidate preserves the current feature set, and the decision becomes
which regression to accept. That question is worth answering before any migration work starts.
## The migration track, in outline
Data movement is the easy half, and deliberately reversible.
1. **Stand the replacement up beside the incumbent**, on its own ports and its own provision type.
No downtime, nothing removed.
2. **Copy bucket by bucket with a neutral tool.** `rclone` rather than the incumbent's own client
— the client has been withdrawn upstream too, so building the migration on it would inherit
the same dependency this effort exists to remove.
3. **Verify per bucket** — object counts and checksums, not a transfer exit code.
4. **Repoint consumers through the connection the module already publishes.** Consumers read an
API URL from the module's declared connections rather than addressing the store directly, so
the cutover surface is that value plus the provisioning and tool handlers.
5. **Freeze writes, final incremental sync, flip**, and keep the incumbent read-only as the
rollback until confidence is earned.
6. **Retire**, and only then remove the module.
The genuinely new work is not the copy. It is the **provisioning handler** and the **tool
handlers**, which are written against MinIO's admin API, and the OIDC wiring.
## Open questions
- How much of the OIDC requirement survives, and in which build? See above — this gates the
choice.
- Does the mesh's bucket provision translate to the candidate's identity model without weakening
what ADR 0049 says about a consumer's identity fitting the tightest backend?
- Should this effort also answer issue 113's general question — mirroring third-party images into
the mesh's own registry — or is that a separate decision? Replacing one withdrawn product with
another unmirrored upstream leaves the same one-way door in place, just further from the hinge.
- Is the four-node erasure-coded topology still warranted, or was it inherited? Worth re-asking
while the product is being chosen, rather than reproducing a shape by default.
@@ -1,6 +1,6 @@
---
topic: what runs on it
status: proposed
status: accepted
date: 2026-08-30
deciders: jochen
reconstructed: false
@@ -0,0 +1,105 @@
---
topic: how we work
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0019-how-this-repository-works.md
---
# 25. The design record is read where it is written, never copied to be found
## Context
**These documents cannot be found by searching the mesh's memory, and never could.** Checked on
2026-08-23 and again on 2026-08-31, against both the symptom-indexed store and the structured
archive, using a decision record's full title and a distinctive phrase from a design document: no
result, no partial match, no stale copy.
That matters because of what was promised. The objection to giving this material its own
repository was that the mesh already has a knowledge store, and a second one repeats the mistake
that store was created to fix. **The answer offered was indexing rather than location** — that
these documents would be returned beside everything else in a search, so where they were authored
became a separate question. The indexing was never built.
**The claim has since stopped being load-bearing**, which is why this is a decision rather than an
incident. [`README.md`](../README.md) names the gap in the place the claim used to sit, and
[ADR 0019](0019-how-this-repository-works.md)'s reasoning rests on cadence, reviewers and scope —
none of which depend on being searchable from elsewhere. What remained was an unbuilt capability
and an open question, recorded as
[`04-ISSUES/006`](../04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md).
**A signpost was added on 2026-08-31 and measured.** One entry in the mesh's memory naming what
lives here and when to come looking. A search for *design records, decisions, repository* returns
it; a search phrased the way somebody actually asks — *why is the mesh built this way* — returns
nothing, because the store matches terms and not meaning. **Reachable is not the same as
surfacing**, and the measurement is what established which one a signpost buys.
## Considered Options
1. **A one-way sync into the mesh's memory.** A job reads this repository on a schedule and writes
the documents into the searchable store. It works with what exists today and needs nothing
built first. **Rejected**, because it creates a second copy of every document, and the failure
mode of a derived copy is the one this repository is least able to tolerate: *the copy that is
searched quietly stops matching the copy that is edited*, and the enforced one wins. A design
record that has silently diverged from the reasoning it claims to carry is worse than one that
cannot be found — the first misleads, the second merely fails.
2. **Leave the signpost and close nothing.** Honest, free, and it keeps the gap visible.
**Rejected as an end state**, though it is what stands until the option below exists. It
answers only for a reader who already suspects these documents exist, which is precisely not
the reader the mesh's memory is designed for.
3. **An agent reads this repository directly, and the search consults it.** Nothing is copied.
**Adopted.**
## Decision
**The design record is read where it is written.** Retrieval is an agent reading this repository,
not a copy living in a second store — and a search of the mesh's memory consults that agent, so
what it knows appears beside ordinary results rather than only when it is asked.
Both halves are the decision. The first alone is merely a reader, and would leave this repository
reachable but not surfacing — the state measured above. **The second half is what discharges the
promise** that these documents are returned beside everything else.
**There is no copy, and that is the point.** No sync, no schedule, no reconciliation, and nothing
that can drift, because there is only ever one of each document. It is also always current,
including for work that is not yet committed.
**The direction of reading is one-way and stays that way.** The agent reads this repository and
answers from it. Nothing flows back: this repository is public, the mesh is not, and a return path
would be how installation-specific detail arrives into documents that must not carry it
([`README.md`](../README.md)).
## Consequences
**This repository stops being a fourth knowledge system, properly.** The original objection was
about adding a knowledge *system*. An agent with read access adds no store at all — which answers
the objection more completely than the indexing that was promised, rather than merely as well.
**ADR 0019's promise is amended, not satisfied.** It said these documents would be *indexed*. They
will not be. They will be *read*, and the search will ask. The commitment that survives is the one
that mattered — that a searcher finds them without already suspecting they exist — and the
mechanism behind it is different from the one named.
**It is gated on an agent that does not exist yet.** Until it does, the signpost is what stands,
and this repository is reachable rather than surfacing. That is a known and stated gap, not a
silent one — and the gap is now a build task with a decided shape rather than an open question.
**The search must degrade honestly.** When the agent cannot be reached, a search has to say that
this material was not consulted. A result set that silently omits it looks identical to one where
nothing matched, and *silence and success must never look alike*
([ADR 0004](0004-a-node-and-how-it-joins.md)) — the rule this repository has now paid for twice.
**A rule states how it is checked, and this one is checkable.** The check is the measurement that
produced this record: search the mesh's memory for a phrase that appears only in a design document
here, and require it back. That check fails today, deliberately, and passing it is what closes
`04-ISSUES/006`.
## References
- [`04-ISSUES/006`](../04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md) —
the gap, the two measurements, and why closing it early was refused
- [ADR 0019](0019-how-this-repository-works.md) — the promise this amends
- [`README.md`](../README.md) — the objection, and the gap named where the claim used to sit
@@ -0,0 +1,130 @@
---
topic: what runs on it
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0004-a-node-and-how-it-joins.md
---
# 26. The mesh has a session of its own, and it is the node session's mechanism
## Context
[ADR 0004](0004-a-node-and-how-it-joins.md) gives every node a session: one per node, permanent,
remembering across callers, its system prompt the node's engram, reachable over the broker like
everything else. **Any node can message any node**, and that is called the one part of the system
that is genuinely a mesh — symmetric, with no centre.
**There is no way to address the mesh itself.** A question that spans machines — *what is running
across all of this*, *which nodes are behind*, *why is it built this way* — has to be put to some
node, which then asks the others. That works, and it makes a mesh-wide question **nobody's
question**: every node answers it as a foreigner, from a position where the whole is not in view.
**Three things independently arrived at the same missing piece.**
[ADR 0025](0025-the-design-record-is-read-not-copied.md), taken hours before this one, commits to
an agent that reads the design repository directly and answers into search. That agent has to
exist, run somewhere, and be askable — and nothing in the record says what it is or where it
lives.
[`14-model-access.md`](../03-DESIGN/01-to-be/14-model-access.md) records, as a gap deliberately
not half-built: *this worker uses that licence is a binding to an agent, not to a node* — and the
provisions model has no consumer identity other than a node. A session that must be assigned a
licence is exactly that consumer, and node sessions are already one.
**And ADR 0004 never said how a session is set up.** It describes behaviour and stops: nothing
states how a session starts, where its context lives, how the engram reaches it, or how a message
off the broker becomes a prompt. There is no design document for it. That gap was invisible until
something had to be built *like* a node session, because describing a second instance of a
mechanism requires the mechanism to have been described once.
## Considered Options
1. **No mesh session; keep relaying through a node.** Costs nothing and works today. **Rejected.**
It leaves mesh-wide questions belonging to nobody, and it does not survive contact with
ADR 0025 — that agent still needs a home, so the thing gets built anyway, unnamed, as an
attachment to whichever node happened to host it.
2. **A new kind of agent, built separately.** Purpose-built for the whole mesh. **Rejected.** It
would hold a session, a memory, a licence and broker plumbing — every one of which the node
session already has. Two implementations of one mechanism drift, and the vocabulary collision
that [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) exists to undo began exactly this
way: two things that were nearly the same, built twice, until neither word meant one thing.
3. **The same mechanism, started in a different context.** **Adopted.**
## Decision
**The mesh has one session, addressed as the mesh, and it is a node session in every respect but
three.**
| | |
|---|---|
| **the context it starts in** | the mesh's, not a machine's — this is the whole of what makes it different |
| **its engram** | the mesh's system prompt, as a node's engram is that node's |
| **its licence binding** | assigned in its own right, not inherited from the machine it runs on |
Everything else is unchanged and deliberately so: it is permanent, it remembers, it is reachable
over the broker, it holds its own tools, and switched off it still answers *I am switched off*
rather than falling silent.
**It runs on the node that holds the control plane** — not for convenience, but because that node
is already the one place excepted from *compromise of a node is compromise of that node*
(ADR 0004). An agent able to reach everything, placed anywhere else, creates a **second** such
place. Putting it where the authority already sits concentrates nothing new.
**It is an addition to per-node messaging and never a replacement.** Every node remains directly
addressable. This is not a preference: ADR 0001 holds that losing the control plane costs *change,
not operation*, and a mesh whose only conversational surface lives on that node would lose the
ability to ask anything while every machine kept running perfectly. **The front door may not be
the single point.**
**It is not an employee** ([ADR 0003](0003-agents-are-persistent-employees.md)). Nobody hires it,
it holds no task queue, it is never drained or reassigned. What it does with work that belongs
somewhere else is **dispatch it** — to node sessions, or to workers — which is what a node session
already does when asked something it does not have.
**It is ADR 0025's reader.** The agent that reads the design repository and answers into search is
this session, not a second one. One agent, one memory, one place to reach; two would both need
that repository and would eventually disagree about what it says.
**"One per node" is about address, not about process count.** ADR 0004's rule — *two and nothing
decides which replies* — forbids ambiguity in who answers when a **node** is addressed. The mesh
session answers when the **mesh** is addressed. The control-plane node therefore hosts two
sessions and no ambiguity, and stating this here is what stops it reading as a contradiction
later.
## Consequences
**The node session's setup must now be designed, and it never was.** This decision is expressed as
*the same as a node session, elsewhere*, which is only meaningful once that mechanism is written
down. The design document covering both is the immediate consequence of this record, not a
follow-up to it.
**A consumer that is not a machine stops being deferrable.** The licence binding above is the gap
`14-model-access.md` names, and it now has two consumers rather than a hypothetical one. Until it
exists, a session's model access can only be expressed as *this module on this machine*, which
cannot say *this node's session uses the personal licence and the mesh's uses the company one* —
the thing the binding is for.
**Symmetry is preserved, and it is worth being precise about why.** ADR 0004's claim is about what
a node can reach, and it is untouched: node-to-node messaging is unchanged, nothing is routed
through the mesh session, and it is a participant rather than a hop. What arrives is a
participant that happens to be the one a person usually addresses.
**Availability degrades to inconvenience rather than to silence** — but only because of the
addition rule above. If that rule is ever relaxed, this consequence inverts, and it inverts
quietly: everything keeps working and nobody can ask about it.
**The surface a person uses is not decided here.** That a board is a good place to talk to it is
likely and is not this record's business; the session is reachable over the broker like everything
else, and what puts a text box in front of it is a separate choice.
## References
- [ADR 0004](0004-a-node-and-how-it-joins.md) — the node session this extends
- [ADR 0025](0025-the-design-record-is-read-not-copied.md) — the reader this session is
- [ADR 0003](0003-agents-are-persistent-employees.md) — the vocabulary this is not
- [`03-DESIGN/01-to-be/14-model-access.md`](../03-DESIGN/01-to-be/14-model-access.md) — *a
consumer that is not a machine*, the gap this makes concrete
@@ -0,0 +1,109 @@
---
topic: what runs on it
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0009-modules-and-the-graph.md
---
# 27. A provision names what the consumer is coupled to, not the role it plays
## Context
Provisions are named after roles. The catalogue and every test fixture built so far say:
```
provides: database
requires: database
```
**Nothing distinguishes one engine from another.** A module requiring `database` is satisfied by
any module providing `database`, so a module written against PostgreSQL can be matched to a
provider of Microsoft SQL Server, resolve as satisfied, deploy, and fail on its first query.
**The mesh runs several engines** — PostgreSQL, Microsoft SQL Server, MariaDB, and others behind
products that expose their own. This is not a hypothetical collision.
**The failure is in the direction that hides.** Resolution *succeeds*. Nothing is refused, nothing
is logged, and the breakage surfaces later as an error inside an application, on a machine, with
nothing connecting it back to a match made elsewhere by something that thought it had done its
job. **A wrong answer delivered confidently costs more than a refusal**, and the whole point of
refusing on ambiguity ([ADR 0009](0009-modules-and-the-graph.md)) was to not do this.
**How it got in:** every test written for the resolver had exactly one provider of each name, so no
mismatch was expressible and none was caught. The fixtures agreed with the design. That is the same
fault as [`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)'s
imagined output and [`019`](../04-ISSUES/019-a-comment-asserting-a-fact-about-a-machine/00-report.md)'s
unchecked comment, at the level of a name rather than a line.
## Considered Options
1. **Keep role names; let the operator assign correctly.** The mesh would refuse ambiguity when two
providers exist, so a person picks. **Rejected.** It makes correctness depend on somebody
knowing that the module they are assigning speaks a particular dialect — which is exactly the
knowledge the provisioning model exists to remove. And with one provider of each name, nothing
is ambiguous and nothing is asked.
2. **A role name plus a `flavour:` or `engine:` qualifier**, matched as a second field.
**Rejected.** Two fields that must agree is a constraint the resolver has to enforce and a
manifest author has to remember, to express something one field already can. The name is the
contract; splitting it invites a requirement that names a role and forgets the qualifier, which
then matches everything again.
3. **The name says what the consumer is coupled to.** **Adopted.**
## Decision
**A provision is named for the thing a consumer's code is written against.**
```
provides: postgres-database
requires: postgres-database
```
**The test is whether the consumer can tell the difference.** If swapping the provider would break
the consumer, the name must say which provider — because a match that breaks the consumer is not a
match. If the consumer genuinely cannot tell, a role name is correct and better.
| provision | | why |
|---|---|---|
| `postgres-database`, `mssql-database` | **specific** | applications are written against a dialect; a swap breaks them |
| `route` | **role** | the consumer wants its name reachable and does not care what proxies it |
| `resolver` | **role** | the consumer wants names to resolve |
| `artifact-store` | **role** | the consumer fetches by digest over a protocol, and nothing else |
**`database` is not a provision and may not be provided.** There is no context in which an
application talks to a generic database: it talks to PostgreSQL or it talks to SQL Server. A name
that cannot be true of any real consumer should not be expressible.
**This is about coupling, not about products.** Two providers of `postgres-database` — a container
on this node and a managed instance elsewhere — are interchangeable and *should* both match. What
may not be interchangeable is what the consumer's queries are written in.
## Consequences
**Every manifest that names a database changes.** Doing this now costs a rename across a handful of
examples. Doing it after modules are migrated costs it across all of them, plus every deployment
that resolved against the old name.
**Wrong requirements now fail loudly, and at the right moment.** A module requiring
`postgres-database` where only `mssql-database` is provided is unsatisfiable, so it is **refused at
resolution** with both names visible — rather than deployed and broken later. This is the property
that was lost, restored.
**Generic role names are still right, and the rule says when.** This does not push specificity
everywhere; it puts it exactly where a consumer is coupled. Naming `route` after a particular proxy
would be the same error in the other direction, and would prevent a swap that genuinely changes
nothing.
**It is checked, not merely stated** ([`00-META/how-we-build.md`](../00-META/how-we-build.md) §5).
A manifest providing a name known to be engine-generic is refused, naming what to say instead.
Without that, this record is a convention, and a convention is what the previous naming was.
## References
- [ADR 0009](0009-modules-and-the-graph.md) — provisions, and refusing on ambiguity
- [`03-DESIGN/01-to-be/07-the-foundation.md`](../03-DESIGN/01-to-be/07-the-foundation.md) — *the
provisioning model uses databases, roles and schemas as PostgreSQL means them*, which is this
record's point made about the substrate before it was made about modules
@@ -0,0 +1,118 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0006-the-substrate-and-the-control-plane.md
---
# 28. The substrate supplies the control plane and nothing else
*Corrects one row of [ADR 0006](0006-the-substrate-and-the-control-plane.md) and makes explicit
something it left unsaid. The rest of that record stands.*
## Context
ADR 0006 defines the substrate by a circularity: **what the control plane needs in order to run,
and cannot ask itself for, because it is not running yet.** Two questions, and both must be
answered *yes* for something to be substrate.
Its membership table admits the object store on this line:
| role | product | |
|---|---|---|
| object store | **MinIO** | it cannot grant itself a bucket |
**That answers the second question and assumes the first.** It is true that a control plane cannot
grant itself a bucket. Nothing establishes that it needs one.
**It does not.** Verified 2026-08-31 against `mesh-control`: no S3 client, no bucket, no object
storage of any kind outside comments. Artifacts reach nodes as content-addressed blobs in the OCI
registry, and the code records the decision and its reasoning:
> One store, and it is the registry the bootstrap already pulls from. An OCI registry is a
> content-addressed blob store that happens to also understand images… The alternative considered
> was a second store beside it — S3-shaped, buckets, signed URLs. It is the right answer for
> objects that are *mutable*, or need per-reader access, or are not build output. None of that
> describes a digest-pinned archive, and standing up a second service to hold one kind of
> immutable blob means two things to run, two things to back up and two ways for an artifact to be
> missing.
**The row is inherited from the system being replaced**, where an object store distributed module
tarballs. Here nothing does, and the row was never re-tested against the definition it sits under.
**A second thing ADR 0006 never says:** whether a substrate service and a module of the same
product are the same instance. It says the substrate is *not the control plane* and *not a place
for logic*, and stops. The question is not idle — an application wanting a database, on a mesh
whose substrate is already running PostgreSQL, has an obvious wrong answer available.
## Considered Options
1. **Leave the object store as substrate, unused.** Harmless-looking. **Rejected.** A membership
list that includes something nothing needs is a list that has stopped being derived from its
test, and the next member is admitted by precedent instead of argument. It also mandates that
every mesh run a service no mesh uses.
2. **Applications share the substrate's instances.** One PostgreSQL, one of everything.
**Rejected**, below.
3. **The substrate is exactly what the control plane consumes; everything else is a module.**
**Adopted.**
## Decision
**The object store is not substrate.** It fails the first half of the test: the control plane does
not need one. An object store is an ordinary module, required through the module graph like
anything else, and a module wanting one depends on a module providing one.
**The substrate has four members, not five**: a relational store, a message bus, an image registry,
and conditionally an identity provider. The registry stays — the control plane genuinely cannot
deliver an artifact without somewhere to put it.
**A substrate service and a module of the same product are different instances, and are not
shared.** The mesh's own PostgreSQL and a PostgreSQL a workload was given are two servers, two
containers, two lifecycles.
Three reasons, and the first is the one that matters:
**The substrate is not in the module graph.** It is raised from the pinned bundle the host carries,
before any mesh exists to declare it. A workload depending on it would depend on something the
graph cannot see, cannot rotate a credential for, and cannot move — which is every property the
provisioning model exists to provide.
**It would put workload data in the control plane's own store.** The mesh's contexts own their
stores exclusively ([ADR 0008](0008-a-context-owns-its-store.md)). An application sharing that
server can exhaust it, lock it, or fill its disk, and the failure is the control plane going down
— which is the one failure that makes every other one harder to fix.
**They are bounded differently.** The substrate is sized, backed up and upgraded as part of
bootstrapping a mesh. A workload's database follows the workload — moved with it, destroyed with
it, restored with it.
## Consequences
**Migrating an object store is ordinary module work**, not substrate work. It was previously going
to be done as part of completing the substrate, which would have been the wrong shape and would
have coupled every mesh to a service the mesh does not use.
**A mesh with no workload needing one runs no object store at all.** That is the correct outcome
and was not previously available.
**Two PostgreSQL containers on a node that hosts both is expected**, not duplication to be
optimised away. Anyone tidying them together should find this record first.
**"Substrate by role and ordinary by delivery" loses one of its two members.** ADR 0006 uses that
phrase of the object store and the registry — things that are substrate but provisioned once a
control plane exists. It now describes the registry alone.
**The definition is applied, not just stated.** Both halves of the circularity test are asked of
each member, and *cannot grant itself one* is not sufficient on its own — it is true of almost any
service, which is what made it possible to admit a member on that half alone.
## References
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — the definition, and the table this
corrects one row of
- [ADR 0008](0008-a-context-owns-its-store.md) — a context owns its store exclusively
- `mesh-control internal/builder/registry.go` — where artifacts go, and why not S3
@@ -0,0 +1,112 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0005-the-node-host.md
---
# 29. A network is a shape, because an action cannot be undone
## Context
**A module of several containers has no way to let them reach each other by name.** A container
declaration carries a `network` field, and it only ever *joins* one that already exists — it was
added so the control plane could reach the store and the broker on the machine it was raised on.
Nothing in the vocabulary **creates** a network.
Without one, containers on a machine share the runtime's default bridge, which gives addresses and
no name resolution between them. So a module that is several containers can only be written by
publishing ports onto the machine and pointing its own parts at the host — which puts a module's
private wiring on the machine's own address space, where anything else on the machine can reach it
and any other module can collide with it.
**This is the gap a mail system meets and nothing else so far does**
([`00-work-breakdown.md`](../03-DESIGN/01-to-be/00-work-breakdown.md) 3.3). It is being taken now
rather than then, because 3.3 is the task most likely to send work back into the declaration
language and the least useful place to discover it.
**Adding a shape is not a small change, and the host says so** — the vocabulary is asserted
against a stated number, with the reason written into the failure: *every addition widens what a
compromised control plane can express, so a change here is a decision.* The host applies what it
is told; the only bound on a hostile control plane is what the language can say
([ADR 0004](0004-a-node-and-how-it-joins.md)).
## Considered Options
1. **An `action` that creates the network.** The vocabulary already has one, the bundle already
uses seven of them, and `docker network create` with a `verify` is exactly the shape an action
takes. Nothing would need adding. **Rejected**, on removal:
> An action has no footprint the host can undo — it ran, and whatever it did belongs to
> whatever it acted on.
A network made this way **leaks when the module is unassigned**, and the mesh cannot tell: the
record says an action ran, and there is nothing to reverse. Unassigning a module would leave a
network behind on every machine it was ever on, and the only way to find them would be to go
and look. *A resource the mesh can create and never clean up is one it should not create.*
There is a second reason, and it is the one that generalises: an action is opaque. **The mesh
cannot tell what an action did**, so a network created by one is not a thing the mesh knows
about — it cannot be reported, counted, or reasoned about, and a module could not require one.
2. **Publish ports on the machine instead.** No new shape, and it works today. **Rejected.** It
makes a module's internal wiring part of the machine's address space: two modules that each
want a database on a fixed port collide, and anything else on the machine can reach what was
meant to be private. It also makes the module's manifest depend on what else is installed,
which is the thing provisioning exists to remove.
3. **`network` as a ninth shape.** **Adopted.**
## Decision
**`network` joins the vocabulary, and the vocabulary is nine shapes.**
```
{"id": "internal", "type": "network", "name": "mail"}
```
**A name and nothing else.** Not a driver, a subnet, an address range or a gateway: every one of
those is a thing a module would have to know about the machine it lands on, and a module that
names a subnet is a module that collides with whatever else chose the same one. The runtime picks;
the mesh names.
**It is created if absent and removed when no longer declared** — an ordinary shape, with the same
lifecycle as a directory. That is the whole reason it is a shape.
**Declared before the containers that join it.** Resources are applied in the order the module
wrote them, and orphans are removed in **reverse** — so a network written first is created first
and removed last, after the containers attached to it are gone. This is not a new rule; it is the
existing one, and it happens to be exactly right here. A network written *after* its containers
would fail to remove while they still hold it, and that failure is reported rather than silent.
**What it does not do:** it does not reach across machines. A network is one machine's, like
everything else the host applies. Modules on different machines reach each other over the private
network the mesh already provides ([ADR 0007](0007-connectivity.md)), and a shape that tried to
span machines would be a second overlay with a worse contract.
## Consequences
**The vocabulary is nine, and the count moves with a record.** The test that asserts it names this
one, so the next person to change it finds the argument rather than a number to edit.
**A compromised control plane can now create and destroy networks on a machine.** Stated plainly
because that is the cost, and the bound is the point: it can create a named network and remove
one, and it can do neither to anything it did not declare. It cannot inspect, attach to, or
reroute what is already there — those would be different shapes, and are not being added.
**A multi-container module becomes expressible**, which unblocks 3.3 and, less obviously, makes
several smaller modules simpler: anything that is a service plus a sidecar currently has to
publish a port to talk to itself.
**Nothing is required to use it.** A module of one container declares no network and joins none,
exactly as now. The substrate keeps using `host`, which is a runtime-provided network and not one
the mesh creates.
## References
- [ADR 0005](0005-the-node-host.md) — the host's vocabulary, and why each shape is a decision
- [ADR 0004](0004-a-node-and-how-it-joins.md) — what may be pushed is bounded by form, not by trust
- [`03-DESIGN/01-to-be/00-work-breakdown.md`](../03-DESIGN/01-to-be/00-work-breakdown.md) — 1.3,
and the mail system at 3.3 that this is for
@@ -0,0 +1,96 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0005-the-node-host.md
---
# 30. Data outlives the mesh that declared it
## Context
**The conversion runs on live services holding real data**, and starts on the node that holds all
of it. Identity, mail, everything. The requirement stated plainly: a data directory may be
*moved*, and may never be *lost*.
**The host deleted them.** A directory that stopped being declared was an orphan, and an orphan
directory was removed with `os.RemoveAll` — everything under it — while the report said
`removed`. A module unassigned took its database's files with it, and nothing anywhere said what
had been in there.
Reproduced before it was fixed: assign a module, let a service write into its directory, unassign
the module, and the file is gone.
**A directory stops being declared for ordinary reasons**, which is what makes this sharp rather
than theoretical. A module unassigned from a node. A manifest edited to move a data folder — the
exact operation the conversion needs. A resource renamed. A typo. **Every one of those is a normal
day's work, and every one of them was destructive.**
**The removal order was already right, and that is what makes a fix possible.** Everything the
mesh puts inside a directory is itself a declared resource, and orphans are removed in reverse
declaration order — so by the time a directory is reached, what the mesh wrote there is already
gone. Anything still present was put there by something else.
## Considered Options
1. **A `keep` flag on the directory.** A module declares which of its directories hold data, and
the host leaves those. **Rejected.** It is safe only when somebody remembered, and the failure
of forgetting is total and silent. A rule that protects data only when it was asked to is not
a rule about data, it is a rule about attentiveness — and this is the one place in the system
where being wrong does not recover.
2. **Never remove a directory.** Simple and unarguably safe. **Rejected**, narrowly: every module
ever assigned would leave its directories behind for ever, and a machine that accumulates
things nobody can account for is one where nobody can tell what is still in use. The clean-up
that is genuinely the mesh's is worth keeping.
3. **Remove a directory only when it is empty.** **Adopted.**
## Decision
**A directory that still holds anything is kept, and the mesh says so.** An empty one is removed.
**This is the host's existing line applied to the one shape where getting it wrong is
unrecoverable** — *it removes what it made and leaves what it merely configured*
([ADR 0005](0005-the-node-host.md)). An empty directory is what the host made. A full one is not.
**No flag, no declaration, nothing to remember.** Emptiness is the test, and it is derived from
the removal order rather than asserted: the mesh's own contents are gone by then, so what remains
is by definition something nobody declared.
**It is reported, not silent.** The outcome is `kept`, naming how many items are inside and saying
they are for a person to deal with. A directory quietly left behind is how a machine accumulates
things nobody can account for — which is the objection to option 2, and it is answered by saying
so rather than by deleting.
**Files are unchanged.** A declared file is the mesh's own — it wrote it, it owns it, and losing a
configuration file is not the failure this is about. The distinction is deliberate: **directories
hold what other things produced; files are what the mesh itself put there.**
## Consequences
**Moving a data directory is now safe by default.** The manifest changes, the old path stops being
declared, and the data stays where it is until somebody has looked at it. That was the operation
most likely to destroy something during the conversion, and it is now the operation that does the
least.
**Unassigning a module leaves its data.** Correct, and it means unassignment is no longer a way to
clean up — removing data is a person's act, done knowingly. Given what unassignment did before,
that is the trade being made and it is the right way round.
**A machine can accumulate directories nobody removed.** Accepted, and mitigated by saying so
every time rather than by a periodic sweep. A sweep would be the deletion this record exists to
prevent, on a timer, with nobody watching.
**It is not a backup, and must not be mistaken for one.** This stops the mesh destroying data. It
does nothing about a disk, a mistaken `rm`, or a service corrupting its own store. The conversion
still needs backups taken and **restored** before anything is moved — a backup nobody has restored
is a belief, not a copy.
## References
- [ADR 0005](0005-the-node-host.md) — the host removes what it made
- [`03-DESIGN/01-to-be/00-work-breakdown.md`](../03-DESIGN/01-to-be/00-work-breakdown.md) — the
conversion this was found by planning
@@ -0,0 +1,72 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0006-the-substrate-and-the-control-plane.md
---
# 31. The control plane authenticates nobody, so identity is a module
## Context
[ADR 0006](0006-the-substrate-and-the-control-plane.md) left one member of the substrate
conditional, and said exactly why:
| role | product | |
|---|---|---|
| identity provider | — | **conditional**: substrate only if the control plane delegates authentication, which is undecided |
[`07-the-foundation.md`](../03-DESIGN/01-to-be/07-the-foundation.md) carried it as an open question —
*whether identity is the fifth* — noting it followed from a decision nobody had taken.
**The decision is taken: the control plane does not delegate authentication.** There is no mesh
identity provider.
**Nothing in the mesh's own machinery ever needed one.** A node proves itself with a keypair it
generated, over a broker account issued at enrolment
([ADR 0004](0004-a-node-and-how-it-joins.md)). Declarations are verified by signature. None of
that touches an identity provider, and the conditional was never about machines — it was only ever
about whether a *person* signing in to a mesh surface would be authenticated by something else.
## Decision
**Identity is a module**, like the mail system and the forge. It runs *on* the mesh, not *of* it
([ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)) — a provider other modules require,
which is the ordinary shape and needs nothing new to express.
**So the substrate is three, and no longer conditional**: a relational store, a message bus, and
an image registry. Together with
[ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md), which removed the
object store, the list is settled and every member is there for the same reason — the control
plane needs it and cannot ask itself for it.
**A mesh that wants no identity provider runs none.** That is now expressible, and was not while
it sat in the substrate as a maybe.
## Consequences
**The last open question about substrate membership is closed.** Both halves of ADR 0006's test
now have an answer for every candidate, and the answer for identity is *the control plane does not
need it*.
**It does not settle how a person signs in to a mesh surface**, and that is deliberately left
open. What is settled is that whatever answers it is not part of what must exist before the mesh
does — so it can be decided late, changed, or replaced, which is precisely what being substrate
would have prevented.
**It becomes a real test of the module graph.** An identity provider is a module that *other
modules require* — the object store already consumes it — so it exercises the provider chain more
seriously than anything ported so far, where the provider was written alongside its consumer.
**Ordering follows from it rather than from preference.** Anything requiring identity has to move
after it, which is a dependency the graph can state rather than something a person has to
remember.
## References
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — the conditional this closes
- [ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md) — the other member
removed, and the test applied properly
- [ADR 0004](0004-a-node-and-how-it-joins.md) — how a node proves itself, which needs none of this
@@ -0,0 +1,79 @@
---
topic: how we work
status: superseded
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0031-the-control-plane-authenticates-nobody.md
superseded-by: 02-DECISIONS/0034-the-local-account-owns-the-mesh.md
---
# 32. The local account owns the mesh; a surface delegates to a module
## Context
[ADR 0031](0031-the-control-plane-authenticates-nobody.md) settled that the control plane
authenticates nobody, and deliberately left one thing open: **how a person signing in to a mesh
surface is authenticated.** This answers it, and answers a question 0031 did not ask — *who owns
the mesh at all.*
**There was no answer, and the absence was invisible** because every operation so far has been run
by the person sitting at the machine. Nothing had to say whether that was the design or the
circumstance.
## Decision
**The account that installed the host owns the mesh on that node.** Authority is a local login,
and there is nothing else to hold.
**No mesh user model.** No accounts, no roles, no grants, nothing to administer. A person with a
shell on a node can do anything the mesh can do there, because that is already true and pretending
otherwise would be a boundary that does not exist.
**This follows from what was already decided rather than adding to it.**
[ADR 0004](0004-a-node-and-how-it-joins.md) says there is no authorisation between nodes — every
node is the operator's own, so a message from one is a message from them, and *the mesh boundary
is therefore the security boundary*. A user model inside that boundary would guard nothing: anyone
who could be stopped by it could equally read the node's key off the disk.
**The board is different, and the difference is the network.** A surface reachable by a browser
has to know who is asking, because the people reaching it are not, by construction, people with a
shell on the machine. **So the board delegates to an OAuth provider** — which is a module.
## What this does not change
**The identity provider is still not substrate** (ADR 0031). A *surface* delegating
authentication is not *the control plane* delegating it. The control plane runs, applies
declarations and reaches nodes with no identity provider in existence; only the board needs one,
and only to decide whose browser it is talking to.
The test is unchanged and still answers no: *does the control plane need it in order to run?*
## Consequences
**The board depends on a module, and says so.** An ordinary edge in the graph, which means the
board cannot come up before the provider it authenticates against — stated as a dependency rather
than discovered as an outage.
**Moving the identity provider takes the board with it.** During that module's own conversion the
board is unavailable, and that is acceptable: it is a surface, nothing depends on it, and a brief
interruption is the trade already accepted everywhere else. Nothing that keeps a service serving
goes through it.
**Anyone with a shell on a node has full authority there.** Written down rather than left implied,
because it is the sentence that decides who gets an account on a machine. The protection is the
machine's own login, and the overlay that keeps the machine unreachable from outside
([ADR 0007](0007-connectivity.md)).
**A node cannot be operated by somebody without a login on it.** Deliberate, and the cost of
having no user model: there is no way to give a person authority over one node without giving them
a shell there. If that is ever wanted, it is a new decision and not a gap in this one.
## References
- [ADR 0031](0031-the-control-plane-authenticates-nobody.md) — the control plane authenticates
nobody; this answers what it left open
- [ADR 0004](0004-a-node-and-how-it-joins.md) — no authorisation between nodes, and why the mesh
boundary is the security boundary
- [`03-DESIGN/01-to-be/11-a-board.md`](../03-DESIGN/01-to-be/11-a-board.md) — the surface this is
about
@@ -0,0 +1,91 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md
---
# 33. The substrate is a store and a broker
## Context
Third correction to one table in one day, all found the same way: by asking whether **both** halves
of the substrate test were actually answered for a given member, or only the second.
The test ([ADR 0006](0006-the-substrate-and-the-control-plane.md)) is *what the control plane needs
in order to run, and cannot ask itself for, because it is not running yet.* ADR 0006 admits the
image registry on this line:
| role | product | |
|---|---|---|
| image registry | **an OCI registry** | it cannot grant itself a repository |
**That is the second half again.** It is true that a control plane cannot grant itself a
repository. Nothing establishes that it needs one *in order to run*.
**Counted rather than argued.** `substrate-first-node.lock` — the only bundle there is, and what a
first node actually becomes — raises twelve resources, and no registry is among them:
```
container runtime · the store · one database per context · the schemas
· the broker's certificate · the broker · the control plane
```
The registry arrives afterwards, as an ordinary module the mesh assigns. That is what the lab
asserts, in those words: *the mesh runs its own artifact store.*
**ADR 0006 half-said this already**, calling the registry *substrate by role and ordinary by
delivery, provisioned once there is a control plane to do it.* A member that is provisioned by the
thing it supposedly precedes is not a member; the phrase was carrying a contradiction rather than
resolving one.
**The registry is a closer call than the object store, and the difference is worth keeping.** The
control plane never touches an object store at all — no client, no bucket, ever
([ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md)). It genuinely
*uses* the registry: the builder pushes to it, hosts pull from it, and nothing reaches a machine
without it. **So the registry is a real dependency of the mesh operating, and not of the control
plane starting** — and it is the second that the word substrate means.
## Decision
**The substrate is two things: a relational store and a message bus.** Both are in the bundle,
both must exist before the control plane's first instruction, and neither can be asked for.
**The registry is an ordinary module.** The mesh cannot deliver anything without one, and it
installs one the way it installs everything else. The first node's chicken-and-egg is already
solved and needs nothing from this list: it fetches upstream images directly, then runs a registry
of the mesh's own.
**The test is applied to both columns, every time.** *Cannot grant itself one* is true of almost
any service and settles nothing on its own. It is what admitted the object store, and then the
registry, and both were removed by asking the other question.
## Consequences
**The substrate is now exactly what the bundle raises**, which is the strongest form this list can
take: it can be checked by counting rather than by reading an argument. A member that is not in
the bundle is not substrate, and the two statements cannot drift apart.
**A mesh that builds nothing still needs a registry** — to receive anything at all — but it needs
it as a module, on its own schedule, replaceable. That was already true and was obscured by the
list.
**The word may now be doing too little work.** "Substrate" for *a database and a broker* is a term
of art for two things everybody can name. Renaming is not taken here and is worth considering
separately; what this record fixes is the membership, not the vocabulary.
**Three removals from one table in one day is itself the finding.** Each member was admitted on the
half of the test that is easy to answer, and the design read plausibly throughout. The rule that
comes out of it is not about substrates: **a test with two conditions is a test only when both are
asked.**
## References
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — the definition, and the table this
corrects a second row of
- [ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md) — the object
store, removed for the same reason
- [ADR 0031](0031-the-control-plane-authenticates-nobody.md) — identity, which was conditional and
is now a module
@@ -0,0 +1,81 @@
---
topic: how we work
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
supersedes: 02-DECISIONS/0032-the-local-account-owns-the-mesh.md
---
# 34. The local account owns the mesh, and a web application's login is not that
*Supersedes [ADR 0032](0032-the-local-account-owns-the-mesh.md), which decided the right thing and
described it wrongly. The decision below is unchanged; what it said about the board was an
invention.*
## Context
ADR 0032 answered *who owns the mesh* — the account that installed the host — and then framed the
board as **a surface that delegates authentication**, a category it made up for the occasion. It
does not need one.
**The board is a web application.** It has a login, provided by the identity module, in the way
every web application has a login. That is a fact about an application, not a property of the
mesh, and giving it a name in the mesh's vocabulary implied a relationship that is not there.
The cost of the invented category was not cosmetic. It made the identity module look like part of
the mesh's own authority — something the mesh *depends on* to know who anybody is — when the truth
is that the mesh knows nothing about people at all, and one of the applications running on it has
a login.
## Decision
**The account that installed the host owns the mesh on that node.** Authority is a local login.
There is nothing else to hold, no user model, no roles, and nothing to administer.
**This follows from what was already decided.**
[ADR 0004](0004-a-node-and-how-it-joins.md) says there is no authorisation between nodes — every
node is the operator's own, so a message from one is a message from them, and *the mesh boundary
is the security boundary.* A user model inside that boundary would guard nothing: anyone it could
stop could read the node's key off the disk.
**A web application's login is its own business.** The board authenticates its users through the
identity module. So might anything else the mesh runs. **None of that is mesh authority**, and the
mesh does not learn who anybody is from it.
## The line this draws, which is the reason to write it down
**Signing in to an application must not, on its own, become authority over the mesh.**
Today it cannot: the board reads and does not act
([`11-a-board.md`](../03-DESIGN/01-to-be/11-a-board.md) — *not the way to change things*). Looking
at a page tells you what is true and changes nothing.
**The moment the board can assign a module, whoever it lets in has mesh authority** — and it would
arrive as a feature rather than as a decision. That is the failure this record exists to make
visible, because it is the kind that is only obvious afterwards.
So: **a surface that can change the mesh is a change to who owns the mesh**, and is taken as one.
Not forbidden — wanting to manage nodes from a browser is reasonable — but not something that
turns up in a pull request titled *add assign button*.
## Consequences
**The identity module is not special.** Not substrate ([ADR 0031](0031-the-control-plane-authenticates-nobody.md)),
not part of the mesh's authority, and nothing about the mesh stops working when it is down. Some
applications cannot be logged into, which is what it means for an application's login provider to
be unavailable.
**Anyone with a shell on a node has full authority there.** Unchanged from ADR 0032, and still the
sentence that decides who gets an account on a machine. The protection is the machine's own login
and the overlay that keeps it unreachable from outside ([ADR 0007](0007-connectivity.md)).
**A node cannot be operated by somebody without a login on it.** The cost of having no user model.
If that is ever wanted, the paragraph above says what it costs.
## References
- [ADR 0032](0032-the-local-account-owns-the-mesh.md) — superseded; same decision, invented category
- [ADR 0004](0004-a-node-and-how-it-joins.md) — the mesh boundary is the security boundary
- [ADR 0031](0031-the-control-plane-authenticates-nobody.md) — the control plane authenticates
nobody
@@ -0,0 +1,115 @@
---
topic: what runs on it
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0034-the-local-account-owns-the-mesh.md
---
# 35. One implementation, several surfaces, and what that costs
## Context
The mesh is operated from a command line today. It needs to be operable from a browser and from a
model's tools as well, and the three must not be three different systems.
**The pattern is already in the code and unnamed.** `board` serves HTTP by calling the same
functions the CLI calls; it holds nothing and decides nothing. What follows makes that the rule
rather than a property of one command.
**The board is a presentation layer over the control plane.** Not an application beside it holding
a database credential — the thing that shows what the control plane knows, and asks it to do what
a person asked for.
## Decision
**The logic lives once, in the context that owns it. A surface is an adapter with no decisions in
it.**
| surface | for |
|---|---|
| **command line** | a person on a machine, and the recovery path below |
| **HTTP** | the board, and anything else that speaks to the mesh over a network |
| **model tools** | an agent asking the mesh to do something |
**Every surface refuses identically, because the refusal is not in the surface.** An assignment
that cannot be satisfied is refused by the same resolution whichever way it arrived. The moment a
surface can accept something another would reject, the mesh has two answers to one question and
people learn which to trust.
**Reading and doing are both exposed.** The HTTP surface is not read-only: managing the mesh from
a browser is the point. This takes the decision
[ADR 0034](0034-the-local-account-owns-the-mesh.md) said had to be taken deliberately —
**a browser login now carries authority over the mesh** — and takes it knowingly rather than
letting it arrive with a feature.
**The networked surfaces authenticate through an OAuth2 identity provider.** Named by protocol
rather than by product, like every other dependency the mesh takes — AMQP for the bus, S3 for an
object store, OCI for the registry
([ADR 0006](0006-the-substrate-and-the-control-plane.md)). What fills the role today is a module
running Keycloak; what the control plane knows is that it validates a token against a provider
speaking OAuth2, and replacing that provider is a migration rather than a redesign.
**The command line does not authenticate at all**: it is already behind the machine's own login,
which is what owns the mesh (ADR 0034).
## What this is not: a kernel every module imports
**The shared library is the failure this project was started over**, and the difference has to be
stated or it will be rebuilt. The old one is 155 files and 34,636 lines *containing code from
every context* — work-domain logic sitting in the kernel every module imports, each piece landing
there to avoid a cycle between two modules that both needed it.
**Shared surfaces are not a shared library.** What is shared here is that three adapters call the
same functions. Those functions stay in the context that owns them — provisioning's logic in
provisioning, identity's in identity — and no module imports another's. A surface may call many
contexts; a context still may not reach into another's store
([ADR 0008](0008-a-context-owns-its-store.md)).
The test, when something is about to be put "somewhere shared": *does this belong to a context, or
does it only belong to the surface?* If it belongs to a context it goes there, even if two
surfaces want it.
## The loop this creates, and the way out
**The control plane's networked surfaces will depend on a module the control plane assigns.**
An identity provider is an ordinary module ([ADR 0031](0031-the-control-plane-authenticates-nobody.md)). When
it is down, or being migrated, or misconfigured, the HTTP and tool surfaces cannot authenticate
anybody — including the person trying to fix it.
**The command line is the way out, and it is why local ownership matters more rather than less.**
It authenticates through nothing, needs no network, and is available on the machine to the account
that owns the mesh. **A mesh must always be operable by somebody standing at it.**
So the rule: **no capability exists only behind an authenticated surface.** Anything the board can
do, the command line can do. That is not a courtesy to CLI users; it is the recovery path, and a
capability that exists only over HTTP is one that disappears exactly when identity does.
## Consequences
**Identity is still not substrate**, and the test still answers no: the control plane runs, applies
declarations and reaches nodes with no identity provider in existence
([ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md)). What is unavailable without it is two
surfaces, not the mesh.
**Whoever the identity provider admits has authority over the mesh.** That is now a real perimeter
with real consequences, where before it guarded a page that only read. Who may log in, and to
which realm, becomes a decision about the mesh rather than about an application.
**A surface must not grow an opinion.** The likely erosion is a validation added to the board
because it was quicker there — and then the CLI accepts something the board rejects, or worse the
reverse. Adapters hold no decisions.
**Three surfaces over one implementation is a cost paid three times if it is not one
implementation.** The reason to write this down now is that the second surface is the cheapest
moment to get it right, and the third is where the drift usually starts.
## References
- [ADR 0034](0034-the-local-account-owns-the-mesh.md) — the local account owns the mesh, and the
line this record deliberately crosses
- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — the shared library this must not
become
- [ADR 0008](0008-a-context-owns-its-store.md) — a context owns its store, which a surface does
not change
@@ -0,0 +1,99 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0035-one-implementation-several-surfaces.md
---
# 36. Bootstrap ends at a usable mesh, and the first credential comes from a person
## Context
Bootstrap currently ends when the control plane starts
([`07-the-foundation.md`](../03-DESIGN/01-to-be/07-the-foundation.md)). That is a mesh that runs and
cannot yet be used by anybody who is not standing at the machine: the networked surfaces need an
OAuth2 identity provider ([ADR 0035](0035-one-implementation-several-surfaces.md)), the provider is
a module, and no module has been assigned.
**So bootstrap should go further** — through the identity provider and the first login — and stop
at a mesh somebody can actually use.
**One thing in the way, and it is not incidental.** The mesh has never held a readable secret. The
sealing code says what it does and why:
> Make generates a secret and seals it to both ends, **keeping no readable copy.**
An initial administrator's credential is the first value a **person must read**. Everything else
the mesh generates is something no human ever sees, and everything a human provides is something
the mesh immediately stops being able to read.
## Considered Options
1. **The mesh generates it and prints it once**, to the terminal of whoever ran the bootstrap.
Convenient, and needs no prompt. **Rejected.** It would give the control plane a plaintext
secret for the first time — briefly, and only to one terminal, but the capability would then
exist. *An exception made for one case does not stay one*: the next credential that is awkward
to supply gets printed too, and the property that a copy of the mesh's database is a copy of
nothing stops being checkable by reading the code.
2. **No password: a one-time link that lets the operator set their own.** The nicest to use.
**Rejected for now** — it needs a mechanism that does not exist, and the thing it improves is
one prompt, once, on a new mesh.
3. **The operator supplies it.** **Adopted.**
## Decision
**Bootstrap runs to a usable mesh**: the substrate, the control plane, the identity provider as an
ordinary module, its realm and client provisioned, an administrator able to log in, and the
networked surfaces available.
**The administrator's credential is supplied by the person doing the bootstrap**, on standard
input and not echoed — the path that already exists for a model-access key. The mesh seals it and
cannot read it afterwards.
**What is created is an account in the identity provider, not a user of the mesh.** The mesh still
has no user model and gains none here ([ADR 0034](0034-the-local-account-owns-the-mesh.md)). What
this produces is the first login for the applications that have one.
**The provisioning is ordinary.** A realm, a client and a first account are what an identity
module's provisioner makes from what the mesh granted it — the same shape as a database and a
bucket, which are built and proven.
**The surfaces arrive when their dependency does.** The command API is not started with the
control plane and then broken until identity exists; it becomes available once it can authenticate,
the way anything else waits for a provider.
## Consequences
**The mesh still never holds a readable secret**, and that sentence needs no exception clause.
That is the whole reason for the prompt.
**An unattended bootstrap is still possible, and the value still comes from outside.** Automation
supplying the credential is the operator supplying it. What is refused is the *mesh inventing*
one — so an unattended bootstrap with no credential provided produces a mesh with no
administrator, which is correct rather than broken.
**Bootstrap gains an interactive step**, and it is the only one. Worth stating because a bootstrap
that cannot run without a person is a real constraint on how a node is stood up, and this is
deliberate rather than an oversight.
**The identity provider is still not substrate.** It is assigned by the control plane, so it comes
after it, and a thing that comes after cannot be a thing that must exist before
([ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md)). Bootstrap running through it does not
move it: bootstrap is a sequence, the substrate is a dependency.
**And the recovery path is unchanged.** When the identity provider is broken later — which is the
failure that matters, not the one at first start — the command line still works, because it
authenticates through nothing (ADR 0035).
## References
- [ADR 0035](0035-one-implementation-several-surfaces.md) — the surfaces, and why the command line
must keep working
- [ADR 0034](0034-the-local-account-owns-the-mesh.md) — the local account owns the mesh; this adds
no user model
- [ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md) — what must exist before the control
plane, which this does not change
+103
View File
@@ -0,0 +1,103 @@
---
topic: building it
status: proposed
date: 2026-09-01
deciders: jochen
reconstructed: false
rests-on: 02-DECISIONS/0009-modules-and-the-graph.md
---
# 37. Where a module lives
## The question
The mesh's own module descriptions currently sit in `examples/` inside the control plane, beside
the small programs that hand out logins. That was fine while there were three of them. It is
wrong now, and the name is doing active harm: everything in `examples/` reads as a sketch, and one
of them shipped naming a container image that nothing in the repository builds. A directory called
*the catalogue* would have made *does this actually work* the obvious question to ask of it.
So: **one repository holding the modules we ship?** And if so, where does everything that is not
ours go?
## What a module actually is, counted
The system being replaced has **126 modules** on its main branch. The shape of them is the whole
argument, so it is measured rather than assumed:
| | count | what it is |
|---|---|---|
| **the module is software** | 47 | its own source tree lives inside the module — a daemon, a service, a library |
| **helper scripts only** | 44 | no application of its own; scripts it runs at install time or offers to an agent |
| **a description and nothing else** | 35 | a package to install and some files to write |
**Two thirds of modules contain code.** The largest is a shared library of 182 source files. A
speech-capture module carries a complete daemon — audio capture, mixing, transcription, a model
runner. Treating a module as *a description of something else* is true of barely a quarter of them.
That kills the simplest answer. A catalogue cannot be "a folder of manifests" when most modules
are programs.
## The four kinds, which want different homes
**1. What the mesh is made of.** The control plane, the host, the shared library, the board.
These are not modules that happen to be ours; they are the mesh, expressed as modules so it can
install itself. They belong in the repositories that build them, which already exist.
**2. Something the world made, that we describe.** A forge, a mail system, an identity provider,
a media server. Nobody upstream ships a description; somebody has to write one, and it is the same
description for everybody who runs it. **This is what a catalogue is for.** It is also where the
small programs that create accounts belong, because such a program is part of describing that
service, not part of the mesh.
**3. Something we wrote, that runs somewhere.** An application, a site, a side project. The
description belongs **with the code, at the root of its own repository**, because the two change in
the same commit. A repository that gains an environment variable and a description that gains it
elsewhere will drift, and there is no mechanism that could stop it. This is already how it works
and it should stay that way.
**4. A package and some files.** A tool, a font, a shell. Thirty-five of these, and each is a few
lines. The catalogue.
## The proposal
**A `mesh-catalog` repository** holding kinds 2 and 4: descriptions of software we did not write,
and the programs that provision it. Not kind 1, which is the mesh itself. Not kind 3, which lives
with its own code.
**The mesh's list of modules is not this repository.** It is a table in the control plane, filled
by adding a description to a running mesh. The catalogue is a *source* to add from — one of
several, and the mesh already records which: every module carries where it came from, the branch
followed there, and the commit its description was read at. **Nothing needs inventing to support
modules from anywhere**; a repository of our own is simply the source we curate.
**A description is checked by the tool, not by a test that imports the tool.** Today a test in the
control plane parses the example manifests by reaching into the control plane's internals, and
another reads the control plane's own build file to check every image a module names can be built.
Two jobs tangled. A `module check` command on the control plane's binary would let the catalogue
hold data validated from outside, and would give the same check to somebody describing their own
application in their own repository — which is the case that matters most and currently has no
check at all.
## What this costs, and the argument against
**It is early.** Ten modules exist, four of them ours. Moving ten files is a morning; moving a
hundred is a week — but the hundred is not here yet, and splitting now adds a second repository to
release across before there is anything to release.
The counter is that the tangle is already producing faults rather than merely threatening to. A
manifest naming an unbuildable image, and a test reading a build file two directories up, are both
symptoms of one repository doing two jobs. And the moment the first module is adopted on a real
machine, the descriptions stop being examples and become the thing deployments come from. **That
is the moment this becomes urgent, and it is close.**
## What it does not settle
**Where a provisioning program's image is published**, and how a description pins it. A description
names an image by digest; the image is built from the catalogue; the catalogue must therefore both
produce an image and refer to it, which is the same knot the bootstrap has and solves by writing
the digest down after building.
**Whether kind 4 deserves a module at all.** Thirty-five descriptions that say *install this and
write these files* may be better as one module with settings than as thirty-five modules. Left
open deliberately; it is a question about the shape of the catalogue, not about whether to have one.
@@ -0,0 +1,82 @@
---
topic: what runs on it
status: proposed
date: 2026-09-01
deciders: jochen
reconstructed: false
rests-on: 02-DECISIONS/0009-modules-and-the-graph.md
---
# 38. The mesh assigns the port, and a module does not care
## The problem, as met
A database module cannot start on a machine that runs the control plane. The mesh keeps its own
store there and holds 5432; the module publishes 5432. Nothing notices until a container runtime
three layers down says `port is already allocated`
([`028`](../04-ISSUES/028-two-things-want-one-port-and-nothing-says-so/00-report.md)).
A module cannot fix this by choosing better, because **a module cannot know what else is on the
machine.** It is written once and assigned anywhere. Any number it picks is a guess about a
machine it has never seen, and two modules guessing the same number is not a mistake either of
them made.
## The number is written three times, and nothing makes them agree
Every module says its port in three places:
| where | for | example |
|---|---|---|
| `listens` | the rule set that lets traffic in | `{port: 5432, from: mesh}` |
| `serves` | what a consumer must know to connect | `{port: 5432}` |
| a container's `ports` | what the runtime publishes | `"5432:5432"` |
They agree today because one person wrote all three. Nothing checks it. A module whose `serves`
said 5432 and whose container published 5433 would resolve, compose, apply, and hand every
consumer a port that answers nothing.
## The decision
**The mesh assigns the machine-side port, and the module says only what it needs.** A module
declares that a container port must be reachable and what it is for. Which number the machine uses
is the mesh's to choose, because the mesh is the only thing that knows what else is there.
**One source, and the other two are derived.** `serves` carries the assigned port so a consumer is
told where to connect without the module having written it down; the rule set is computed from the
same assignment. Three copies become one fact.
**An assignment is made once and kept**, exactly as a credential is. A port that moved on every
push would restart both ends each time and would hand consumers a number that was true when it was
read.
## Some ports cannot move, and that is a claim
Mail is 25, submission is 587, IMAP over TLS is 993. A mail system on a strange port is not a mail
system. So a module may say a port is **fixed by the protocol** rather than assigned.
**A fixed port is exactly a claim** — the thing the mesh already has for what is singular on a
machine: one seat, one display server, one artifact store. Two modules wanting 25 on one machine is
the same shape as two wanting the seat, and gets the same answer: the second is refused, by name,
when it is assigned rather than when it is applied.
That is why this does not need a new mechanism so much as it needs the existing one pointed at
ports.
## What follows
- **A module becomes portable in a way it was not.** Two databases on one machine stop being a
collision and become two assignments.
- **The substrate has to be visible.** The mesh cannot assign around its own store while it has
never heard of it. What the bundle holds must be written down somewhere the assignment can read
— which the bundle does not say today.
- **A refusal can be useful.** *25 is held by the mail system on this machine* is a sentence a
person can act on. `port is already allocated` is not.
- **`serves` stops being written by hand**, which is a small vocabulary change with a large
consequence: what a consumer is told is now derived from what actually happened.
## What this does not settle
**Whether a module should publish to the machine at all.** Assignment makes publishing safe; it
does not make it necessary. Consumers could instead reach a provider on the module's own network by
name, with nothing published — which would make the question moot for anything inside the mesh, and
would still leave it for anything reached from outside.
@@ -0,0 +1,106 @@
---
topic: building it
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
---
# 39. What the SDK holds, and what it refuses
_Reconciliation note (2026-09-05): supersedes the earlier "repository structure" decision, which the consolidation folded; no standalone record remains to point at, so body references to it now point at the nearest surviving record, [ADR 0015](0015-applications-live-in-their-own-repository.md)._
## Context
The earlier "repository structure" decision (folded in consolidation; see the reconciliation note
above, and [ADR 0015](0015-applications-live-in-their-own-repository.md) as the nearest survivor)
named `mesh-sdk` "contracts shared across tiers:
types, not behaviour." That line is superseded here, because it draws the boundary in the wrong
place. The boundary that matters is not *types versus behaviour* — it is **how often the thing
changes**.
The current SDK is the cautionary tale, and its failure is precise. `hal/sdk` holds all the
code, including a per-module API client for every service (`clients/plex.ts`, `clients/gitea.ts`,
…) and a per-module tool implementation for each (`tools/plex.ts`, …). Every module depends on
the SDK, so **every edit to any of that per-module code rebuilds every module** — the cascade.
The SDK is under constant maintenance precisely because it became the place all the volatile
per-module logic accumulated.
The root cause is worth stating exactly, because the fix follows from it: the pressure was never
to share a client *between* modules. It was to share a client between one module's *own features*
— plex's tools, its health check and its hooks all wanted the same `PlexClient` — and the only
place to share code across a module's features was the global SDK. So **intra-module sharing
leaked out as inter-module coupling.**
## Decision
The SDK holds the **stable spine** that modules build against, and earns its place by rarely
changing. The test for membership is change-frequency, not kind.
### What it holds
- The **tool-serving harness** — the worker and registration mechanism, and the tool-definition
type. *How* a tool is declared and served is settled; it does not change when an individual
tool does.
- The **messaging and event framework** — the broker client, the event consumer, the envelope.
- The **contracts** — the manifest, declaration, provision and link shapes.
- **Core primitives** — sealing and crypto, semver, the shared resolution helpers.
These change rarely and deliberately. When one of them does change, a rebuild of everything is
the *correct* outcome, because the contract every module shares has genuinely changed.
### What it must not hold — the more important half
- **A module's API client.** A Plex client, a Gitea client, a MinIO client belong in their
module. They change when that service's API or the module's use of it changes, which is often,
and which has nothing to do with any other module.
- **A module's tool implementations.** Same reason, same place: in the module.
- **Anything volatile** — anything that changes when one service's features change.
The rule, stated so it can be applied without re-deriving it:
> If editing a thing recompiles unrelated modules **and** it changes often, it does not belong
> in the SDK.
Both conditions are load-bearing. A rare change that cascades is fine — that is a contract, and
the cascade is correct. A frequent change that stays local is fine — that is a module minding its
own business. Only **frequent *and* cascading** is the disease, and per-module clients and tools
are its carriers.
### Where per-module shared code lives instead
Code shared among a module's *own* features lives **in the module**. The default is the plainest
thing that works: an ordinary shared file the features import — `plex/client.ts`, imported by
`plex/tools/`. Within one module, features are files importing sibling files; no package
boundary, no ceremony.
A **module-local SDK** (a sub-package with its own version) is warranted only for the few modules
whose shared surface is large enough to version on its own. It is the exception, not the shape.
Either form gives the property the global SDK could not: editing a module's shared code rebuilds
**that module and nothing else**.
## Consequences
- The cascade becomes **structurally impossible for module logic**. There is no longer an edge
from one module's internals to another, so the only thing that can rebuild everything is a real
change to a shared contract in the SDK — which is rare, and when it happens, is right.
- The SDK is small and stable **by construction**, not by discipline. Its size is no longer a
thing anyone has to police.
- **Converting a module from the current system is partly a de-coupling, not just a move.** Its
client and its tools are pulled *out* of the shared SDK and *into* the module. A conversion
that copied `clients/plex.ts` into the SDK's replacement would rebuild the exact mistake.
- The host still does not import the SDK. It depends on nothing
([ADR 0005](0005-the-node-host.md)) and **mirrors** the contracts rather than
importing them, exactly as its apply-shapes table already does deliberately. The SDK is shared
by the tiers that *can* share code; the host is not one of them.
## References
- The earlier "repository structure" decision — named the repositories; its `mesh-sdk`
description ("types, not behaviour") is superseded by this record (folded in consolidation;
nearest survivor [ADR 0015](0015-applications-live-in-their-own-repository.md)).
- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — the decomposition this serves: code
belongs to the boundary that owns it.
- [ADR 0005](0005-the-node-host.md) — why the host mirrors the contracts instead of
importing the SDK.
+101
View File
@@ -0,0 +1,101 @@
---
topic: what runs on it
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
extends: 0009-modules-and-the-graph.md
---
# 40. What a module is
_Reconciliation note (2026-09-05): supersedes the earlier "grouped by domain" decision, which the consolidation folded into how-we-build.md; no standalone record remains to point at._
## Context
[ADR 0009](0009-modules-and-the-graph.md) settled that everything is a module, but never said what a
module *is* beyond "a directory the mesh processes." That gap let the catalogue's breadth read as a
smell: a module can carry a container, a built image, tools, a provisioner, migrations, health,
config, seat claims, requires and provides — so much that the unit seemed ill-defined.
The earlier "grouped by domain" decision (folded in consolidation; see the reconciliation note
above) tried to organise modules by domain, which is the wrong axis. This record states what a module is, drawn from the cases that
stress-tested it: the shell, i3-vs-sway, umami, and "database."
## Decision
**A module is one self-contained piece of software the mesh installs and manages** — everything
needed to make that one thing real and integrable: what runs, the seats it claims, what it provides
to other modules, what it requires from them, and what operates it.
The **software is the module's identity.** Capabilities, seats and provisioned resources are the
**relationships *between* modules**, not what a module is — and that is what binds a module into one
thing. umami is bound by *being umami*: its container runs umami, its provisioner creates umami sites,
its tools query umami, its `requires` gets umami a database. Every feature serves the one software.
### The three relationships
1. **Shared seat** — several modules fulfil a capability and coexist; one may be default. bash, zsh
and fish all join `shell`.
2. **Exclusive seat** — modules contend for a single slot; one holds it. i3 (needs x11) and sway
(needs wayland) contend for `display-session`.
3. **Provide / require** — a provider ships the **provisioner** that creates instances of the
resource it offers and returns sealed credentials; a consumer requires it and the mesh wires the
credential in. Symmetric: umami requires a database *and* provides analytics.
### Interfaces are mesh-owned; providers adapt to them
The mesh **defines the interface** for a capability — the provider-neutral contract of what a
consumer receives and how it integrates. Both sides conform: a provider's provisioner **adapts** its
software's real API to the mesh contract; a consumer depends on the **interface**, never on a
provider. Swap one provider for another and the consumer does not change.
### The naming rule — draw the interface at the consumer's real coupling
Name a `provides`/`requires` at the **widest boundary across which the consumer genuinely does not
care which implementation serves it**:
- Where the consumer's coupling is thin — an analytics embed snippet and dashboard, opaque to it —
the mesh defines a neutral interface (`analytics`) and providers (umami, amumi) adapt. Swappable
across vendors.
- Where the consumer **speaks a protocol** — a database's wire protocol and query dialect — the
interface *is* the protocol: `postgres-database`, `mssql-database`, `mongodb-database`. Swappable
only among protocol-compatible implementations, **never across**, because the application cannot
cross it either. "database" is not a capability; the protocol is.
- **Never false genericity.** A name must not promise a swap the contract cannot deliver
([research 005](../01-RESEARCH/005-domain-grouping/analysis.md)).
This is [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md)'s rule made general — "names the protocol, not
the product; a database names the engine because the app targets it" — with the reason stated: the
contract sits where the coupling is.
### What is not a module
- A **library** (built against, never deployed — [ADR 0039](0039-what-the-sdk-holds-and-refuses.md)).
- A **control-plane context** (the mesh itself — [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)).
A **swappable machine mechanism** (a firewall — ufw, nftables) *is* a module implementing a
capability. The host hardcodes no firewall, supervisor, package manager or runtime; it owns only the
generic apply primitives and platform detection, so it runs where none of those exist — an Android
phone has no ufw, systemd, pacman or Docker.
## Consequences
- **Supersedes the earlier "grouped by domain" decision** (folded in consolidation; see the
reconciliation note above). Modules are
organised by their relationships (seats, provisions), not grouped into domain folders.
- **Refines [ADR 0009](0009-modules-and-the-graph.md).** Everything the mesh runs and integrates is
a module — but a module is defined by the *software it delivers*, not by being a bucket of features.
- The target is a **self-fulfilling mesh**: declared wants bound to swappable modules, provisioners
wiring credentials, nothing hardcoded. The control plane's whole job is the binding.
- Converting a module from the old system includes pulling its per-module code out of the shared SDK
([ADR 0039](0039-what-the-sdk-holds-and-refuses.md)) and shipping its provisioner as an adapter to a
mesh interface — a de-coupling, not just a move.
## References
- [ADR 0009](0009-modules-and-the-graph.md) — everything is a module; this says what one is.
- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — contexts are the mesh, not modules.
- The earlier "grouped by domain" decision — superseded (folded in consolidation; see the note above).
- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — protocol-not-product, generalised here.
- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — per-module code lives in the module.
- [research 011](../01-RESEARCH/011-the-module-graph/00-overview.md) — the graph of these relationships.
@@ -0,0 +1,85 @@
---
topic: what runs on it
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
extends: 0040-what-a-module-is.md
---
# 41. Events are a relationship, the lighter sibling of provisioning
## Context
[ADR 0040](0040-what-a-module-is.md) names two relationships between modules — seats and
provide/require (provisioning). A third is latent in the mesh and worth making first-class: the
broker every node already runs ([ADR 0002](0002-nodes-communicate-over-a-broker.md)) can carry a
module's activity as **events**, which any other module reacts to. A logger that writes an audit
trail, a module that acts when another module acts, observability — all of it is one mechanism, and
today it is ambient rather than declared.
## Decision
**A module emits events and consumes events, and both are declared** — parallel to `provides` /
`requires`, so the mesh knows the event graph the same way it knows the provisioning graph.
### Events are provisioning's lighter sibling
| | provisioning | events |
|---|---|---|
| shape | **1:1**, a provider creates a resource *for* one consumer | **1:many**, a module emits, any number listen |
| credential | yes — sealed, per consumer | none — it is broadcast |
| machinery | a provisioner (the reconcile adapter) | nothing but the broker's topic routing |
| declared as | `provides` / `requires` | `emits` / `consumes` |
Because an event is broadcast and credential-free, there is no provisioner and no per-consumer
setup — only a subscription. That is why it is the *lighter* relationship, and why most
inter-module reaction should be an event, not a provision.
### An event carries what an audit needs
Every event carries its **type** (a dotted topic key, so listeners match by prefix), its **source**
module, the **node** it came from, and the **time**. A body follows. The metadata is not optional:
a reaction may only need the body, but an audit trail needs to know who did what, where and when,
and an event that cannot answer that is not auditable.
### The audit logger is just a consumer of everything
A logger that records the whole mesh's activity is **not a privileged component** — it is an
ordinary module that consumes `#` (every event) and writes them down. It holds no special access;
it only listens widely. That it falls out of the model with no new machinery is the check that the
model is right.
### `consumes` is validated like `requires`
A `consumes` for an event that **nothing** `emits` is a dangling edge, and the mesh refuses it
before deploy — the same rule that catches a `requires` for a resource nothing provides
([research 011](../01-RESEARCH/011-the-module-graph/00-overview.md)). A listener waiting for an
event that can never arrive is a silent failure, and this repository's whole discipline is against
silent failure.
### One runtime serves all three
The per-node module runtime that serves a module's tools also wires its `consumes` (subscribe,
dispatch to the handler) and lets its code `emit`. Tools are *invoked* (request/reply), resources
are *provisioned* (1:1, credentialed), events are *emitted and consumed* (1:many, broadcast) —
three relationships, one broker, one runtime, all declared on the manifest.
## Consequences
- The mesh gains a declared **event graph** alongside the provisioning graph — visible, validated,
reasoned over.
- **Reaction becomes the default coordination**: a module acts on another's event without either
knowing the other, and without a credentialed link. Coupling drops.
- An **audit trail** is a module, not a platform feature — and can be swapped, extended or run more
than once (a file logger and a queryable one) with no change to anything that emits.
- The runtime must dispatch a module's event handlers as well as its tools; that generalisation is
small (both arrive by importing the module's entrypoint) but it is real work.
## References
- [ADR 0002](0002-nodes-communicate-over-a-broker.md) — the broker events ride.
- [ADR 0040](0040-what-a-module-is.md) — the relationships this extends.
- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — `emit`/`on` are stable sdk surface; the
broker binding and the runtime are not.
- [research 011](../01-RESEARCH/011-the-module-graph/00-overview.md) — the graph these edges join.
@@ -0,0 +1,116 @@
---
topic: what runs on it
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
extends: 0041-events-are-a-relationship.md
---
# 42. The shape of an event on the wire
## Context
[ADR 0041](0041-events-are-a-relationship.md) made events a relationship — `emits`/`consumes`, the
graph, the audit logger. It did not say what an event *is* on the broker: the exchanges, the
routing keys, the headers, the queues and their configuration. That shape is a contract every
emitter and consumer conforms to, exactly as [ADR 0010](0010-delivery.md)
is for declarations — and it was being decided ad-hoc in code. This settles it, so the sdk and the
runtime implement one contract and a module never reinvents it.
## Decision
### Two exchanges, kept apart
- **`mesh.events`** — a durable topic exchange. Every event rides it: module, mesh and node.
- **`mesh.rpc`** — a durable topic exchange. Tool invocations (request/reply) ride it.
Kept separate because RPC is not an event: a `#` subscription on `mesh.events` is then a complete
audit of what happened, with none of the invocation traffic.
### The routing key is the event type, namespaced by origin
Dotted and hierarchical — `<origin>.<name>.<event…>` — with three reserved origins:
- `module.<module>.<event>` — `module.umami.site.created`
- `mesh.<context>.<event>` — `mesh.delivery.deployed`, `mesh.provisioning.granted`
- `node.<node>.<event>` — `node.anchor.joined`, `node.anchor.unreachable`
Topic matching gives a consumer `node.*.joined`, `module.umami.#`, or `#`. The origin roots are
reserved; everything after is the emitter's own namespace.
### Metadata in headers, payload in the body
An event's identity and provenance are AMQP **headers**, so a consumer — or the broker, or an
audit tool — reads who/when/what without parsing the body, and the body is only the domain payload.
**Required headers**
| header | meaning |
|---|---|
| `x-event-id` | a unique id — for dedup and audit (delivery is at-least-once, below) |
| `x-source` | the emitter: the module, context or node name |
| `x-node` | the node it was emitted from |
| `x-time` | emit time, RFC-3339 |
| `content-type` | `application/json` |
**Optional headers**
| header | meaning |
|---|---|
| `x-causation-id` | the event or command that caused this one — tracing |
| `x-schema` | a version of the body's shape, so a body evolves without silent misreads |
The routing key already carries the type; it is not duplicated as a header. An **unknown `x-`
header is ignored, not refused** — unlike a declaration, an event is observed by parties that need
not all understand every header, and refusing would couple every consumer to every emitter's
additions.
### Messages are persistent
Events are published persistent (delivery-mode 2). An audit trail that loses events on a broker
restart is not one, and the cost is disk the broker already spends on everything durable.
### Queues: one per consumer, durable, dead-lettered
- **A consumer's queue** is `<node>.<module>.events`, durable, bound to that module's consumed
patterns. Durable so a restart does not drop what arrived while it was down. **Manual ack** after
the handler succeeds — at-least-once.
- **Prefetch** bounds in-flight work (default 32) so one slow consumer does not pull the whole
backlog into memory.
- **A dead-letter exchange** `mesh.events.dead` receives a message rejected past a redelivery limit,
so a poison event is set aside for inspection rather than looping forever or vanishing silently.
- **The audit logger's queue** `<node>.audit-logger.events`, bound to `#`, is the same shape —
durable, persistent, dead-lettered — because completeness is its whole job.
- **RPC reply queues** are exclusive, auto-delete and server-named; **RPC serve queues**
`serve.<key>` are durable and shared, so several runtimes serving one tool key compete rather than
each answer.
### At-least-once, and consumers are idempotent
A handler may see an event twice — a redelivery after a crash between doing the work and acking.
Consumers must be idempotent, and `x-event-id` is what makes dedup possible. **Exactly-once is not
offered**: it is a promise no broker keeps honestly, and saying so is better than pretending.
## Consequences
- The event shape is a versioned, enforced contract, not conventions each module reinvents. The
sdk's `emit`/`on` and the runtime's AMQP binding implement it; a module never sees an exchange or
queue name.
- Metadata-in-headers means the body is exactly the domain payload, and a consumer that only wants
provenance never parses it.
- Adding a header or an origin root widens the contract and is reviewed as one — the discipline
[ADR 0010](0010-delivery.md) applies to the
declaration vocabulary.
- The sdk's first cut carried source/node/time in the *body*; this supersedes that — they move to
headers. That is code to align, in `mesh-sdk` (`emit`/`on`) and `mesh-tools` (the binding, queue
config, dead-letter).
## References
- [ADR 0041](0041-events-are-a-relationship.md) — events as a relationship; this is their wire shape.
- [ADR 0010](0010-delivery.md) — the precedent: a wire
contract, versioned, additions reviewed as security.
- [ADR 0002](0002-nodes-communicate-over-a-broker.md) — the broker.
- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — `emit`/`on` are stable sdk surface; the
binding, queue config and dead-letter are the runtime's, not the sdk's.
@@ -0,0 +1,101 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0041-events-are-a-relationship.md
---
# 43. A module's broker account is scoped by what it emits and consumes
## Context
[ADR 0041](0041-events-are-a-relationship.md) made events a relationship — `emits` and `consumes`
on the manifest. [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) gave them a wire shape — the
`mesh.events` exchange, the durable per-consumer queue, the reserved routing-key origins. Neither
said how a module *reaches* the broker: what account it holds, and what that account is allowed to
do.
As the code stands, there is no answer. The mesh can provision a **node** account (at enrolment)
and a **builder** account (scoped to the build queue), and it can *deliver* any module a sealed
own-secret at a declared path — but it has no way to provision a broker **account** for a general
module. A module that declares `own-secrets: {broker: …}` and nothing more receives thirty-two
random bytes, not a credential. So on the broker, `emits` and `consumes` are enforced by nothing: a
running module could bind any queue, consume any pattern, and publish under any origin, and the
manifest that says otherwise would be describing a boundary no code draws — the exact shape of fault
[04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) records, a scope
declared in manifests and read by nothing.
This settles it, so a module's place on the bus is a thing the broker enforces rather than a thing
the manifest merely claims.
## Decision
### A module gets a broker account when it is assigned, and its permissions are the manifest
When the mesh assigns a module to a node it provisions a broker account for that module on that node,
sealed to the node ([ADR 0004](0004-a-node-and-how-it-joins.md)) and delivered as the
module's `own-secrets` broker — `amqps://` with the mesh's fingerprint, the shape
[ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) already carries. The account's permissions are
derived from the manifest, and are exactly these:
- **What it consumes.** Read on `mesh.events`, and configure-and-read on its own queue
`<node>.<module>.events` bound to the patterns in `consumes`. It cannot bind or read another
module's queue. A module that consumes nothing gets no read on the events exchange at all.
- **What it emits.** Write to `mesh.events`, restricted to routing keys under its own origin,
`module.<name>.*`. It cannot publish as another module, and cannot publish under the reserved
`mesh.*` or `node.*` origins — those belong to the mesh and the host (ADR 0042). A module that
emits nothing gets no write.
- **Nothing else.** The events account reaches `mesh.events` and that module's own queue, and no
more. Tool serving and calling over `mesh.rpc` is a separate grant on the same principle — a
module serves the tool keys it declares and calls the ones it is bound to — and is scoped the same
way rather than folded in here.
### Consuming everything is a privilege, granted deliberately
`consumes: ["#"]` — the audit logger — is read across the whole bus: every module's events, the
mesh's, every node's. That is not a pattern like any other; it is the power to see everything, and
the account is where it becomes visible. The grant that lets one module read the entire bus is one
the mesh issues on purpose and can be audited — the answer to *who can read everything* is a row, not
a guess — rather than a breadth any manifest acquires by typing a single character. A `#` consume is
a reviewed grant, not a default one.
### The account is how the declaration is enforced
Because the account can do only what `emits` and `consumes` name, the broker itself refuses a module
that tries to consume a queue it did not declare or emit under an origin it does not own. That is what
makes an event relationship a rule and not a comment — the discipline that a stated rule says how it
is checked. A manifest that over-declares grants more than the module uses, which is visible and
reviewable; one that under-declares makes the module fail closed at the broker, which is the safe
direction to be wrong in.
## Consequences
- The control plane gains a **generic module broker-account**, derived from the manifest. The
builder stops being a special case: its access to the build queue becomes an ordinary expression of
what it consumes and serves, not a bespoke account method. One rule, and the builder is an instance
of it.
- The runtime reads its credential from a file (the broker own-secret), `amqps://` verified against
the mesh's fingerprint. The `guest` account is for raising the substrate, never for a module — a
module documented as holding its own credential and handed the broker's administrative one is worse
than one with no credential story at all.
- `emits` and `consumes` stop being advisory. They are the module's authority on the bus, so the
manifest is now a security boundary and is reviewed as one, the discipline
[ADR 0010](0010-delivery.md) applies to the declaration
vocabulary.
- *Who can read the whole bus* becomes an answerable question, because `#` is a grant and not an
accident.
## References
- [ADR 0041](0041-events-are-a-relationship.md) — events are a relationship; this scopes the account
by that relationship.
- [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) — the wire this account secures: the queue,
the origins, the `amqps` credential shape.
- [ADR 0004](0004-a-node-and-how-it-joins.md) — the link is the security boundary; a
module's account is sealed to its node the same way a node's is.
- [ADR 0010](0010-delivery.md) — a declaration is owned
and its additions reviewed; a module's broker permissions are that discipline applied to the bus.
- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — a scope declared
in manifests and enforced by no code: the fault this decision closes for events.
@@ -0,0 +1,94 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0027-a-provision-names-what-the-consumer-is-coupled-to.md
---
# 44. A public name is provisioned, not registered by hand
## Context
The mesh names and resolves its own machines internally: the overlay generates
`<service>.<node>.<suffix>` wildcards, dnsmasq answers them (`wildcard-resolution`), and the mesh
issues a certificate for each internal name. A service reachable at a *public* domain —
`plex.example.com`, not `plex.anchor.internal` — needs three things that machinery does not give it:
- a **public DNS record** at a registrar or DNS provider, so the name resolves on the internet;
- a **publicly-trusted certificate** for it, because the mesh's own authority is trusted by nobody
outside the mesh;
- and routing from that name to the module — which the reverse proxy already does: a module
`requires` the `route` capability and the proxy provides it, routing by the host it was asked for.
The routing exists. The public DNS record does not: the mesh has no way to make a name resolve on
the public internet, so today that is a step someone does by hand at a DNS provider, outside the
mesh, remembered nowhere. A public name is therefore the one part of reaching a service that the
declaration graph cannot grant or withdraw — which means it is created once and outlives whatever it
was for, the shape of drift this project exists to remove.
## Decision
### A public name is a capability, requested like any other
A module reachable at a public host declares `requires: ["public-dns"]` and contributes the hostname
it wants — beside `requires: ["route"]`, which exposes it through the proxy. The name is then
provisioned on declaration ([ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md)): created
when the module is assigned, removed when it is withdrawn, reconciled like every provision.
### The interface is neutral; the providers are the registrars
`public-dns` is drawn at the consumer's coupling: the consumer wants *a public name that resolves to
me*, and does not care whether Cloudflare, Route 53 or a registrar's own API puts the record there.
So the interface is neutral and the providers are provider-scoped — `cloudflare-dns`,
`route53-dns`, `porkbun-dns` — each implementing the one `public-dns` contract, the same way a
neutral database coupling is answered by `postgres-database` and `mssql-database`. A module names
`public-dns`; it never names a registrar.
### The record points at the mesh's public ingress, not at the node
What the name resolves to is the address the reverse proxy answers on, not the consuming machine's.
A public service is reachable only *through* the proxy — the proxy holds the `route` grant and routes
by host to the module — so the public name must resolve to the proxy. `public-dns` and `route` are
the two halves of one public exposure: the name, and what the name reaches.
### The record is a fact, not a secret
A DNS record is public by definition, so the grant returns the fully-qualified name and its TTL and
nothing sealed. The only secret is the provider's own API credential, which is the provider module's
own-secret and never leaves it — the module that wanted the name never sees it.
### Events
The provider emits `module.<provider>.record.created` and `module.<provider>.record.removed`
([ADR 0041](0041-events-are-a-relationship.md)), so *which names the mesh publishes, and where* is a
question answered from the event trail and the grants, not from a folder of records edited at a
provider.
### The public certificate is the proxy's, and is named here only to pair it
A public name without a publicly-trusted certificate is reachable and not trusted — the same pairing
the internal name and the mesh-issued certificate already have. Obtaining that certificate (ACME
against the now-resolving public name) is the reverse proxy's to do, and its mechanism is its own
decision; it is named here so the pairing is not forgotten, not resolved here.
## Consequences
- A public name is created and torn down with the module, so it cannot outlive it, and the mesh can
say which public names it publishes without anyone reading a registrar's dashboard.
- Adding a registrar is adding a provider that answers `public-dns`; the modules that want names do
not change.
- Public exposure of a service is a trio of separate, declared, enforced relationships: the firewall
opens the proxy's public port ([ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)),
`route` routes the host to the module, and `public-dns` makes the host resolve.
## References
- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — a capability is provisioned on
declaration; a public name is one.
- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the firewall, the other
half of the reachability question this was asked with.
- [ADR 0041](0041-events-are-a-relationship.md) — the provider's record events.
- [ADR 0040](0040-what-a-module-is.md) — a provider and its interface; the neutral-interface,
scoped-provider naming this follows.
@@ -0,0 +1,93 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0005-the-node-host.md
---
# 45. A machine's firewall is the sum of what its modules listen on
## Context
The reverse proxy is a *provider*: a module `requires` the `route` capability and a running proxy
provides it, routing traffic by name and reaching back to the consumer. A fair question follows —
is the firewall the same shape? Should a module *register* a port with a firewall provider the way
it requests a route?
It should not, and the difference is the point. A reverse proxy is a service another component
performs; a firewall is a property of the machine — a packet filter the host applies to itself.
Modelling it as a provider would invent a credential and a reach-back for something that has neither.
And the mesh already has the registration: a module declares `listens: [{ port, from }]` — the port
it accepts connections on, and from where. That *is* how a service says it wants a port open. What is
missing is not a model but enforcement. [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md)
records that a `scope:` key five manifests carry is read by no code: a manifest can appear to
restrict a port and restrict nothing — the exact fault
[how-we-build.md](../00-META/how-we-build.md) names, *an unenforced rule is indistinguishable from a
wrong one*, made worse because the declaration reads as a restriction.
## Decision
### The firewall is derived and host-applied, not a provider
A machine's firewall is the sum of what the modules assigned to it declare they listen on, computed
by the host and applied as one of its owned resources ([ADR 0005](0005-the-node-host.md):
the host applies, it does not decide; [ADR 0010](0010-delivery.md):
the declaration is owned resources). It is not a capability, not a per-consumer grant — opening a
port is a declarative fact about a machine, so it is computed and applied, not requested and
credentialed.
### `from` is the whole of public-versus-internal
The distinction the question is really about lives in `from`:
- `listens: [{ port: 5432, from: mesh }]` — open to the private overlay only.
- `listens: [{ port: 443, from: anywhere }]` — open to the public internet.
A module registers a port on the firewall by listening on it and saying from where. There is no
separate firewall capability, because the firewall is not a thing that reaches back or holds a
secret; it is the machine's own filter over the ports its modules named.
### The host enforces it both ways, and unknown keys are refused
A port a module listens on is opened to exactly the scope it named; a port nothing declares is
closed. And a key the firewall does not read — the `scope:` of issue 003 — is refused at the
manifest, not accepted and ignored, so a declaration that reads as a restriction is one. This is the
discipline [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) applied to the
broker account, applied here to the packet filter: the declaration is the enforcement, or it is a
comment.
### A public service is exposed through the proxy, not by opening its own port
Reaching the public internet is normally not `from: anywhere` on the service's own port. The service
listens `from: mesh` — only the proxy reaches it — and `requires: route`, so the sole machine with a
public opening is the one running the reverse proxy, and the service is exposed by name through it.
`from: anywhere` is the deliberate direct-exposure case, for a service that is its own front door.
## Consequences
- Issue 003 is closed: the firewall is computed from `listens` and enforced, so a declared scope is
real and an undeclared port is shut. Rejecting unknown manifest keys is the general fix, of which
the `scope:` key was one instance.
- The firewall and the reverse proxy stop being confused for one model: the firewall is the machine's
filter (host-derived from `listens.from`); `route` is a name-router (a provider); the public DNS
name is a third thing ([ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md)). A
public service uses all three.
- The modelling question is answered: a module registers a port by declaring `listens`, and reaches
the public internet by name through `route` + `public-dns` — never by the firewall being a
provider.
## References
- [ADR 0005](0005-the-node-host.md) — the host applies; the firewall is one of
the things it applies.
- [ADR 0010](0010-delivery.md) — the firewall is a derived
owned resource, not a grant.
- [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) — the same discipline:
a declaration is enforced, or it is a comment.
- [ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md) — the public name, the other
half of the reachability question this was asked with.
- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — the unenforced
`scope:` this closes.
@@ -0,0 +1,88 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0027-a-provision-names-what-the-consumer-is-coupled-to.md
---
# 46. A module's configuration is its assignment's, not its manifest's
## Context
A module is assigned to a node — `assign <node> <module>`, always to a machine; there is no
assignment to the mesh. "Mesh" is a *scope*, not a place: a `provides` or a `claim` scoped `mesh`
reaches the whole mesh, but the module still runs on a node. So the two kinds of thing a module can
carry are the manifest (what the module *is*) and, separately, what it should do *here* — which
differs by deployment and by node.
The mesh already has the second: **settings**. `settings set <module> [--node <node>]` — with a node
it is that machine's, without it the whole mesh's — layered over what the module declares and applied
at resolution, changeable without editing the module and without a rebuild. That is the surface a
meshboard would edit.
But settings today reach only a module's **config-file content** (a mergeable file the module owns).
Configuration that is not a file has been landing in the manifest instead, statically — a registrar's
zone and domain, the address public names point at, and, most sharply, `listens.from`. That last one
is the tell: whether a port is open to the private overlay or to the public internet is a
*per-node deployment choice* — the same database internal on one machine and public on another — and
a value fixed in the manifest is one value for every machine, so it cannot be. Static configuration in
the manifest is configuration in the wrong place: it cannot vary per node, and it cannot change
without a new module version.
## Decision
### The manifest is identity and defaults; the assignment's settings are the configuration
A module's manifest declares what it is — what it provides, requires and claims, the shape of its
resources — and, for anything configurable, a **default**. The values that make a running instance
*this* instance are settings, carried by the assignment: per-node, or mesh-wide when no node is named,
applied over the defaults at resolution. Change one and the next reconcile carries it; nothing is
edited on a machine and nothing is rebuilt.
### Settings drive the configurable fields the manifest marks, not only file content
Settings extend beyond a config file's content to the manifest fields a module declares settable —
foremost:
- **`listens.from`**: a module declares its safe default (`from: mesh`), and a per-node setting
raises or lowers it. postgres declares `listens: [{ port: 5432, from: mesh }]`; on the machine that
should expose it, a setting makes that port `from: anywhere`. Same module, different exposure, and
the firewall ([ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)) is computed
from the effective value, so the packet filter follows the setting.
- **A provider's own configuration**: a registrar's zone, domain and the ingress its names point at
([ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md)) are mesh-wide settings, not
manifest constants — one mesh's Cloudflare zone is not another's, and the module description is the
same for both.
### Unset is the default, and an unknown setting is refused
A field with no setting keeps the manifest's default, so a module runs correctly configured by nobody.
A setting that matches no settable field — like a config value that reaches no file today — is named,
not silently dropped, so a misspelled setting is found rather than believed (the discipline of
`UnusedSettings`, and of [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md):
a declaration is enforced or it is a comment).
## Consequences
- The postgres case works: one module, `from: mesh` by default, `from: anywhere` where a setting says
so — internal on ace, public on novox, changeable live.
- Provider modules stop carrying a mesh's specifics: `cloudflare-dns` describes *a Cloudflare
registrar*, and *which* zone and ingress is a setting, so the same module serves every mesh.
- Configuration becomes a thing a meshboard manages — set per node or mesh-wide, applied on the next
reconcile — rather than a manifest edit and a rebuild ([ADR 0011](0011-managed-files-are-generated-never-edited.md):
the way you change a managed thing is not by editing it).
- What a manifest may not do is grow a value that differs per machine; if it differs per machine it is
a setting, and the manifest holds only the default.
## References
- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — what is provisioned on
declaration; its per-instance values are the assignment's.
- [ADR 0011](0011-managed-files-are-generated-never-edited.md) — a managed thing is changed through
the mesh, not by editing it; settings are that, for configuration.
- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the firewall follows the
effective `listens.from`, so making `from` a setting makes exposure a setting.
- [ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md) — the provider whose zone and
ingress are settings, not manifest constants.
@@ -0,0 +1,87 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md
---
# 47. A module runs its code as its own process, with its own account
## Context
A module is one self-contained thing ([ADR 0040](0040-what-a-module-is.md)), and it gets a broker
account scoped to what it emits and consumes ([ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md)).
The catalogue now gives modules **tools** and **events** — real code, in the module ([ADR 0039](0039-what-the-sdk-holds-and-refuses.md)) —
but nothing has said what *runs* that code. The audit-logger showed one shape and was treated as an
exception: a container running the tool runtime carrying the module's compiled code, holding the
module's own scoped account. Every module with tools or events needs the same, and the tempting
alternative does not work.
**A node-wide runtime that loaded every assigned module's code cannot hold a per-module account.** It
would run under one account with the union of every module's permissions — able to emit as any of
them and read any of their queues — which is exactly the isolation ADR 0043 exists to draw. So the
runtime is per-module, not per-node, and treating the audit-logger as special left the other
modules' code with nothing to run it: the conversion produced tools and events that, as it stands,
never execute.
## Decision
### A module with tools or events runs a process of its own
A module that has tools or events runs a **runtime process** — a container, the tool runtime carrying
that module's compiled code — assigned and started like the module it is, holding the single broker
account the mesh scoped to it (ADR 0043). One module, one process, one account.
### It serves its tools, each on its own key
A tool is served on its own key (`serve.<tool>`), and a caller invokes a named tool. Only the module
that serves it answers, and the module's account is scoped to exactly its tool keys — so one module
cannot answer another's calls, the isolation ADR 0043 gives events extended to tools. This supersedes
a single `tools.invoke` endpoint that dispatched by name: that shape assumed one runtime for the
whole node, and per-module runtimes competing on one key would each be handed calls for tools they do
not have.
### It runs its events in the same process, under the same account
Emitting under the module's own origin and consuming its own queue ([ADR 0042](0042-the-shape-of-an-event-on-the-wire.md))
happen in that same process, with that same account — not a second one to scope and seal. A module's
tool code, its event code and, for a provider, its provisioner are the one module's code and run as
the one module's process.
### The runtime image is the tool runtime plus the module's code
Built from the module's source like any module image — the audit-logger's shape, made the rule, not
the exception. The module declares a `container` for it carrying `MESH_BROKER_FILE` (its sealed
credential, ADR 0043) and its compiled code. A module with **neither** tools nor events runs no such
process: a plain service module — the plex *server*, dnsmasq the resolver — is its service and files
and nothing more. A module that is both a service and code declares both containers: the service, and
the runtime beside it.
## Consequences
- The catalogue's tools and events become runnable: each tools-or-events module gains a runtime
container with its scoped credential, and the audit-logger stops being special. Until this, the
converted modules held code with nothing to execute it.
- A process, and a small image, per tools-or-events module. That is the cost of ADR 0043's isolation:
one account per module means one process per module. It is paid deliberately — a shared runtime is
cheaper and cannot be scoped, and a mesh where any module can emit as any other is not one worth the
saving.
- `serve.<tool>` per key replaces the single `tools.invoke` dispatch. The sdk's serving and a module's
account scope both come to name tools individually.
- **A provider's provisioner is a runtime process too.** It already runs as its own container; its
events (`bucket.created`, `database.provisioned`) belong to *that* process and need the same
credential. So a provisioner that emits carries `MESH_BROKER_FILE` and its scoped account like any
runtime — or it does not emit. (This is the fix for provisioners that emit today with no broker
bound: the emit is a runtime's, and the provisioner is a runtime.)
## References
- [ADR 0040](0040-what-a-module-is.md) — a module is one self-contained thing; its code runs as one
process.
- [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) — the scoped account
this process holds, and the isolation that makes it per-module.
- [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) — the events this process runs, and the
`serve.<key>` queue tools now use.
- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — the code lives in the module; this runs it.
@@ -0,0 +1,131 @@
---
topic: what runs on it
status: accepted
date: 2026-09-05
deciders: jochen
reconstructed: false
---
# 48. A provider creates the credential the mesh minted, and seals nothing
## Context
A provider module stands up a per-consumer resource — a database, a cache bucket, an object
store user — and the consumer must end up holding a credential that authenticates against it.
Building the module runtime (ADR 0047: a module runs its own code as its own process under its
own account), the provider's provisioner was run for the first time as a delivered thing, and
it did not work. It reads a seal key from the environment that nothing sets, and it seals every
credential it produces to that key with a symmetric passphrase.
Tracing the credential's path turned up something larger than a missing key. **The provisioner
harness the whole catalogue is built on describes a credential flow the mesh does not have, and
duplicates — incorrectly — one it does.**
What the sdk's `runProvisioner` does today:
- reads request files named `*.grant.json` — which nothing in the mesh writes;
- calls an adapter whose `create` **generates its own password** and returns it;
- seals that password with a symmetric key (`$MESH_SEAL_KEY`) and writes a `*.credential`
file — which nothing in the mesh reads, and no consumer ever unseals.
What the mesh already does, and has wired end to end:
- The control plane mints one password per (consumer, provider) pair (`Inventory.SecretFor` →
`secrets.Make`) and seals it to **both** node keys asymmetrically — a copy the consumer's
host can open and a copy the provider's host can open. No shared symmetric key exists
anywhere, on purpose: a key both ends hold is a key the mesh would have to distribute, which
is the same problem one level down, and the control plane's own code refuses it.
- The provider is handed, at the path its `receives` names, one contribution per consumer:
the **login to create** (`As`, derived by the mesh so the two ends agree by construction),
the consumer's address and requested values, and a **`Secret` file** holding that consumer's
password sealed to the provider and unsealed onto the machine by its host.
- The consumer is handed the *same* password, as plaintext its own host wrote by unsealing its
copy and substituting it into a config file. The consumer never unseals anything itself and
holds no key.
So the password a provider's provisioner invents is not even the password the consumer was
given: a consumer authenticating with the mesh's password against a resource the provisioner
created with its own would simply fail. The symmetric seal is not an incomplete feature to
finish delivering a key for. It is a second, contradictory credential model bolted beside the
real one, and it cannot be made to work without building the very thing the mesh was designed
not to have.
This is a decision and not a patch because the harness is the **provider contract**. Every
provider — the four that exist and the many a real mesh grows — is built on
`runProvisioner(resource, adapter)`. Whatever it says a provider is, they all inherit; and
changing it later is one migration per provider. It is cheaper and more honest to settle what
a provider is now.
## Decision
**A provider is handed the credential; it does not make one, does not seal one, and does not
hand one back.** The provisioner's only job is to make the mesh's grants true in its own
software.
Concretely, for the sdk harness and the adapter contract:
- The harness reconciles the **contributions the mesh delivers** to the provider's `receives`
path — the list of consumers, each with its login name (`As`), address, requested values,
and the path to its unsealed password (`Secret`). It does not read `*.grant.json` and it
does not write `*.credential`.
- For each consumer present, the harness reads the password from that consumer's `Secret` file
and calls the adapter to bring the resource into being under the given login. For each
consumer no longer present — the mesh drops it from the contributions file when its consumer
goes away — the harness calls the adapter to withdraw it.
- The adapter shrinks to the per-software half and nothing else. It is given the login, the
password, and the values, and it makes the resource exist or removes it. It generates no
password, derives no name, seals nothing, and returns no credential:
roughly `create({ as, password, values })` and `remove({ as })`, both returning nothing.
- `$MESH_SEAL_KEY`, the symmetric `seal()`/`writeSealedCredential` path, and the `*.grant.json`
/ `*.credential` files are removed from the provisioning path entirely. The credential
reaches the consumer through the mesh's own asymmetric channel, which already crosses node
boundaries and holds no shared secret.
Identity stays the mesh's to say. The login the provider creates is the name the mesh derived
and gave the consumer to present; the provider never invents a name, because a name the
consumer cannot learn is a name it cannot authenticate with.
## Consequences
- A provider module becomes smaller and unable to be wrong in this way: with no password to
generate and no key to seal to, the class of bug where the two ends hold different secrets
cannot be written. A provider added after this inherits the corrected contract and has no
seal to reintroduce.
- The four current providers (redis, postgres, minio, umami) each lose their `generatePassword`
+ seal code and gain a `create` that takes the password it is given. Their teardown becomes
"withdraw the login named `As`".
- The symmetric `seal()`/`unseal()` primitive loses its only caller and leaves — checked, not
assumed: nothing else in the sdk or the catalogue called it, so it is removed with the
provisioner it belonged to.
- **How this is verified:** redis is assigned as a provider in the lab, the contributions and
the unsealed password the mesh would deliver are put in its `receives` path, and a client
authenticates as that consumer with the mesh's password and gets PONG — where a provider that
invented its own password answers WRONGPASS — with `$MESH_SEAL_KEY` set nowhere and no
`.credential` file written. Proven: `provider-uses-mesh-credential` is green.
**What this does not cover — credential provisions, not data provisions.** This decision is about a
provision whose credential is a *secret the mesh mints* — a login and password (redis, postgres,
minio). A provider that instead *generates* the thing the consumer needs, and that thing is not a
secret — umami's `analytics`, where the consumer wants back a `siteId` umami assigned — does not fit,
because a contract that returns nothing has no way to hand that data back. The seal-key fault was
never umami's (it sealed no password; it returned a public id), so removing the seal does not break
it further, and it still reconciles its sites off the mesh's contributions. But delivering
provider-generated data back to a consumer is a *return path* the mesh does not have and this
decision does not build — a separate shape, left to a separate decision.
- Teardown beyond "remove the login" — data an object store leaves behind when a consumer
leaves — is named by each provider's adapter, not by the harness, and is out of scope here
except to say the contract must leave room for it.
## References
- [04-ISSUES/032](../04-ISSUES/032-provider-runtime-has-no-seal-key/00-report.md) — the
observation and the cross-repo trace this decision rests on.
- ADR 0047 (the module runtime) — what first ran a provider's provisioner as a delivered
process and exposed this; link to be filled when 0047 lands on the trunk.
- Control-plane mechanisms this relies on already existing: `mesh-control` —
`internal/inventory/secrets.go` (`SecretFor`, `SecretsFrom`), `internal/secrets/seal.go`
(`Make`, the two-blob asymmetric sealing), `cmd/mesh-control/plan.go` (`grantsFor`, the
`Grant.Sealed = ForProvider` delivery), `internal/catalogue/declaration.go` (the `receives`
contribution: `As`, `At`, `Values`, `Secret`).
- The path being removed: `mesh-sdk` — `src/provisioner/index.ts` (`runProvisioner`, `sealKey`,
`writeSealedCredential`) and the symmetric `src/primitives/index.ts` `seal()`/`unseal()`.
@@ -0,0 +1,130 @@
---
topic: what runs on it
status: accepted
date: 2026-09-05
deciders: jochen
reconstructed: false
---
# 49. A consumer's identity is bounded by the tightest backend that must accept it
## Context
The mesh says who a consumer is, once, and hands the same name to the provider (to create) and the
consumer (to present), so the two ends agree by construction rather than by two conventions (the
principle behind `ConsumerIdentity`, 04-ISSUES/023). The name is `mesh_<node>_<module>`, cleaned to
lower-case letters, digits and underscore.
Proving the provider contract per backend (ADR 0048) turned up 04-ISSUES/034: redis and postgres
create that name verbatim, but **minio refuses it** — an S3 access key is capped at 20 characters,
and `mesh_anchor_bucketuser` is 22. The provisioner then retries for ever, per consumer, and the
consumer holding that same too-long name could never present it either.
Two things about the existing derivation decide most of this:
- **The charset is already right.** `[^a-z0-9_]` is deliberately conservative, and its own comment
says it reaches "a PostgreSQL role, a MinIO access key, an LDAP uid and a Keycloak client without
quoting." That much is true.
- **The length is wrong.** `CheckIdentity` refuses names over `identityLimit = 63`, commented as
"the shortest identifier limit among the systems these names reach: PostgreSQL's". It is not the
shortest — S3's 20 is shorter — so the guard that was meant to catch exactly this lets it through,
and the failure lands at provision time as a silent retry instead of at assignment as a refusal.
So this is a small wrong constant with a real cost attached: whatever bound we set, `mesh_` (5) plus
a node name plus `_` plus a module name has to fit inside it.
## The options
**A — Bound the identity by the true minimum, and refuse early.** Lower `identityLimit` to the real
shortest (20, S3's), so `CheckIdentity` refuses an over-long name *at assignment* with a clear
message, the way it already refuses over-63 names. The derivation does not change; long names are
simply rejected before anything is provisioned.
- *For:* smallest change; keeps "the mesh says the identity once, verbatim" intact; the failure
moves from a per-consumer provision-time retry to an up-front, legible refusal — which is what
`CheckIdentity` exists to do.
- *Against:* a hard budget. `mesh_` + node + `_` + module ≤ 20 means node + module ≤ 14 characters.
`anchor` + `bucketuser` (16) is already over. It pushes the constraint onto how machines and
modules are named, which is a real limitation on legible names.
**B — Keep the readable name when it fits, compact it when it does not.** Below the bound, the name
is `mesh_<node>_<module>` as today; over it, the mesh substitutes a deterministic short form (e.g.
`mesh_` + a truncated hash of node+module) — still one derivation, so both ends still agree.
- *For:* no naming constraint; short backends always satisfied; the common case stays legible.
- *Against:* some identities become opaque, and a provisioner tracing "whose login is this" loses
the answer for exactly the consumers that overflowed. The mesh now owns a fallback format and its
collision properties (a truncated hash is not free of collisions at 15 characters).
**C — Let each interface declare its identifier bounds, and derive within the tightest a consumer
reaches.** `s3-bucket` states `identifier: { max: 20 }`; `postgres-database` states 63; the mesh
derives a name that fits the **minimum** bound across the providers a given consumer is granted.
- *For:* the most precise — each provision gets exactly the room it has, and a database consumer
keeps long legible names while an S3 consumer gets a short one; the constraint lives where the
fact does (on the interface).
- *Against:* the most work, and a consumer of two interfaces with different bounds must satisfy the
smaller — so its name shortens for both, reintroducing B's opacity in a narrower case. It also
means one consumer can hold **different** identities per provision, which the "said once" model
currently forbids.
**D — Let the provider generate a backend-valid identity and hand it back (rejected).** minio mints
its own access key and returns it to the consumer. This is the data-provision return path this era
keeps meeting — but it directly contradicts 023 and ADR 0048: the identity would no longer be the
mesh's single derivation the two ends share, it would be a value one side invents and the other must
be told. Listed for completeness; not recommended.
**E — A module (and a node) may declare a short slug; the identity is built from it.** The identity
becomes `mesh_<node-slug|node-name>_<module-slug|module-name>`: where a slug is declared it is used,
otherwise the cleaned name. A slug is a deliberately short, operator-chosen identifier — `kc` for
keycloak, `wkstn` for a workstation. It is optional: short names (`anchor`, `redis`) need none.
- *For:* this is the escape hatch B wanted to be, without the opacity. The name stays legible — a
provisioner can read `mesh_wkstn_kc` and know who is asking — because a person chose it, not a
hash function. And it makes an early refusal *palatable*: if even the slug-built identity overflows,
the refusal points at the slug, a field made for exactly this, rather than at the machine's name.
Both ends still derive it from one declared thing, so they agree by construction.
- *Against:* a new optional manifest field, and someone must pick the slug — but only for names that
would otherwise overflow, and picking a short legible identifier is a better job than being handed
a hash.
## What implementing A revealed
A was tried first. At `identityLimit = 20`, the readable budget is `mesh_` (5) + node + `_` + module
≤ 20, i.e. **node + module ≤ 14 characters** — far tighter than it looked. The catalogue's own
existing tests use `workstation`+`keycloak` (25), which compacts to `mesh_dbbc02f8dde34d3`; common
mesh names (`home-server`, `the-build-node`, `workstation`) blow the budget with any module. So B's
compact fallback would fire for the *common* case, not the rare overflow — which inverts A+B: most
identities would be opaque hashes. A alone (hard refusal at 20) would refuse most realistic names.
This is what moved the recommendation to E: the problem is not the limit, it is that the *readable
name* is the wrong source when it is long, and a slug is a better source than either a hash or a ban.
## Recommendation
**E, over a per-consumer bound (start with the global minimum, 20).** Build the identity from an
optional slug, keep it when it fits, and refuse at assignment with "declare or shorten `<module>`'s
slug" when it does not — no hash, no lost legibility, and the fix is a first-class field. Set the
bound to the true minimum (20) now; it needs no per-interface machinery to unblock S3, and a module
that consumes S3 simply declares a short slug. Graduate to **C** (per-interface bounds) later if it
turns out that non-S3 consumers are paying for S3's limit often enough to mind — E and C compose:
slugs are the mechanism, per-interface bounds refine where the ceiling sits. **B is dropped**: a
declared slug is a strictly better escape hatch than an opaque hash. **D stays rejected.**
## Consequences (of E)
- A module manifest gains an optional `slug`; a node may carry one too. `ConsumerIdentity` prefers
the slug over the cleaned name for each half. `identityLimit` becomes 20 (the true minimum), and
`CheckIdentity` refuses at `module add` / assignment — now with a message naming the slug to set.
- The common case stays legible; only names that overflow the budget need a slug, and what they get
is a name a person chose, not a hash.
- Existing modules/nodes whose names overflow declare a slug once — a migration cost paid as a clear
refusal with an obvious remedy, not a silent hash or a silent truncation.
- minio (04-ISSUES/034) is unblocked: an S3 consumer declares a short slug and its access key fits.
- **How it is checked:** the minio grant e2e — a consumer whose (slugged) identity fits reaches its
bucket with the credential the mesh delivered — plus unit tests that a slug is preferred, that an
un-sluggable over-long identity is refused (naming the slug), and that two consumers never collide.
## References
- [04-ISSUES/034](../04-ISSUES/034-mesh-login-exceeds-s3-access-key-limit/00-report.md) — the
observation.
- ADR 0048 — a provider creates the credential the mesh minted; the identity it creates it under is
the one this decision bounds.
- `mesh-control` `internal/catalogue/identity.go` — `ConsumerIdentity`, `identityUnusable`,
`identityLimit`, `CheckIdentity` — where the constant and the check live.
@@ -0,0 +1,206 @@
---
topic: what runs on it
status: accepted
date: 2026-09-05
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0024-model-access-is-a-provision.md
---
# 50. Model access is vendor-agnostic, and a vendor is an adapter
## Context
[ADR 0024](0024-model-access-is-a-provision.md) settled that model access is a provision and that
a licence is a named thing an operator uses. What shipped, and runs, is a single vendor: the mesh's
"claude" feature. A read-only trace of that feature (2026-09-05, in the code workspace) was made to
answer whether the model-access provision is Anthropic-shaped or genuinely general. The finding is
that **the vendor-agnostic layer already largely exists**, and the Anthropic specifics are a thin
band around it that a per-vendor adapter can hold.
**What is already general, with evidence.** `mesh-control internal/licences` models
`licence(name, vendor, serves)` and `licence_holder(licence, node, module, sealed)`, and each
holder's credential is sealed per-holder through `internal/secrets`. `serves` carries the
non-secret facts (a base URL, a model) and is not vendor-specific. The `accept` verb
([ADR 0024](0024-model-access-is-a-provision.md), and
[`14-model-access`](../03-DESIGN/01-to-be/14-model-access.md)) already takes an operator-supplied
value, seals it to each holder and discards the plaintext. A model the mesh runs itself answers
`model-access` at node scope with no licence at all. None of that mentions Anthropic.
**What is Anthropic-specific.** The credential is not a static key: it is a subscription OAuth grant
— an hourly access token plus a refresh token. That shape drags four things behind it that a static
key does not need: **central rotation** (one manager node refreshes under a lease and publishes the
new token), **delivery that strips the refresh token** so a consuming node holds only an access token,
an **identity guard** that reads the credential to catch a mis-binding, and a **usage** reading with
Anthropic's own `utilization%` semantics. Most vendors are a single static key, which the sealed-key
model already handles and which needs none of these four.
**The tension at the centre of this.** A refreshable credential cannot be both *sealed so the mesh
cannot read it* and *rotated centrally*. Central rotation means some node in the mesh holds the
refresh token in readable form, because that is what refreshing requires. Per-holder sealing means no
node but the holder can read the credential. For a static key the two never meet — there is nothing to
rotate. For a refreshable grant they collide directly, and this record exists to say which gives way,
and by how much.
## Considered Options
1. **Keep Anthropic special-cased in the core.** Leave the three binding columns and the `claude_*`
schema, and add other vendors beside them the same way. **Rejected.** It is exactly what
[ADR 0024](0024-model-access-is-a-provision.md) ruled against: a module that names a vendor cannot
be moved onto another model without editing it, and moving it is the point. It also grows the core
by one band per vendor, when the bands are the same shape.
2. **One provision, and refuse to hold any refresh token — re-seal only.** Make every credential
purely sealed per-holder, including refreshable ones; let each holder refresh its own grant.
**Rejected.** It throws away the hard half [ADR 0024](0024-model-access-is-a-provision.md) says
already works — the lease, the single-refresher, the switch-on-exhaustion — and replaces it with N
nodes each holding a refresh token, which is the very thing today's delivery strips on the stated
ground that *a node never holds a refresh token*. A refresh token is the long-lived secret; spraying
it across every holder is strictly worse than keeping one copy on one node.
3. **One provision, and abandon central rotation entirely** for refreshable vendors — treat the grant
as opaque and let it expire. **Rejected.** For a subscription-seat vendor an expired access token is
a dead licence; without rotation the feature that works today stops working. This is option 2's cost
without option 2's autonomy.
4. **One vendor-blind provision, plus a per-vendor adapter, with a bounded carve-out for the
refreshable case.** **Adopted**, below.
## Decision
**`model-access` is one consumer-facing, vendor-blind provision.** A consumer declares
`requires: model-access`, and is coupled to *reaching a model* — a base URL, a model name, a key —
and not to which vendor answers. That is the coupling the name is drawn at
([ADR 0040](0040-what-a-module-is.md)'s rule: name the interface at the widest boundary across which
the consumer does not care which implementation serves it). Where a consumer were genuinely coupled to
a specific wire API it could not swap across, the same rule would split the name — but the consumers
that exist reach their model through a CLI or SDK that hides the vendor, so `model-access` is the true
coupling and stays one name. This extends [ADR 0024](0024-model-access-is-a-provision.md) and
[ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) without changing them.
**A vendor is an adapter, keyed by the licence's `vendor` field.** The lifecycle a licence needs is
vendor-specific and lives in a per-vendor adapter selected by `licence.vendor`, exactly as
`public-dns` is one neutral interface answered by registrar-scoped providers —
`cloudflare-dns`, `route53-dns` ([ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md)).
A consumer names `model-access` and never a vendor, the same way a module names `public-dns` and never
a registrar.
**The field is named `vendor`, not `provider`.** The inventory already uses "provider" for the
provider-pin — *which node answers a brokered provision*. Reusing it for *which company sells this
licence* would collide two unrelated facts on one word. `vendor` is the licence's, and is separate.
### The adapter's capabilities, all but one optional
An adapter declares:
- **`shape`** — `static-key` or `refreshable-grant`. This is the switch the carve-out below turns on.
- **`accept(value) → sealed`** — take an operator-supplied credential and seal it to the holders, the
`accept` verb [ADR 0024](0024-model-access-is-a-provision.md) already defines.
- **`refresh(licence)`** — refreshable-grant only: the lease / rotate / publish machinery.
- **`identity(credential) → account-id`** — the mis-binding guard, for a vendor whose credential
carries an identity worth checking.
- **`usage(licence) → normalised rows`** — the vendor's usage reading, mapped to the common shape below.
- **`deliver`** — the credential *value* only; the destination path is the consumer's, not the
adapter's.
**A static-key vendor implements almost nothing** — `shape: static-key`, `accept` is the generic
seal, `deliver` is the value, and `refresh`, `identity` and `usage` are absent or trivial. The
abstraction earns its keep by making the common vendor small, not the rare one clever.
### The carve-out — the one place the guarantee is relaxed, said plainly
The mesh's standing principle is that it cannot read what it stores: `accept` seals to the holders and
discards the plaintext ([ADR 0024](0024-model-access-is-a-provision.md)), and a provider seals nothing
because the credential travels the mesh's own asymmetric channel
([ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)). A `refreshable-grant`
credential cannot honour that principle and be centrally rotated at the same time, and central rotation
is the working half [ADR 0024](0024-model-access-is-a-provision.md) is explicit about keeping.
**So, for `refreshable-grant` vendors only:**
- the **manager node holds the refresh token encrypted at rest** — readable by that node, because
rotation requires it. This is the bounded exception.
- **access tokens are still sealed per-holder**, as every credential is; a holder reads its own and no
other node reads it.
- the **refresh token is stripped on delivery** — it never reaches a consuming node. *A node never
holds a refresh token* stays true for every node but the one manager.
**Static-key vendors keep the full guarantee.** There is no token to rotate, so there is nothing to
hold readably, so `accept` discards the plaintext and the carve-out never fires. The majority of
vendors are static-key, and the majority therefore lose nothing.
The exception is stated rather than hidden because a relaxed guarantee that is not written down is
indistinguishable from a broken one. It is bounded on three axes at once: **refreshable-grant vendors
only, the refresh token only, the manager node only.**
### The settled details this record also fixes
- **Usage is normalised to `(licence, consumer, period, metric, value)` plus the raw response as
jsonb.** The metric is vendor-defined — Anthropic's `utilization%` is one metric, a token count is
another — and no common unit is forced across vendors. The raw response is kept so a reading can be
re-derived if the normalisation is later found wrong.
- **Binding is explicit per consumer, and an unchosen consumer is refused — no implicit fallback.**
This is the direction the resolver already takes, and it is the safe one: a mesh with several ways
to reach a model refuses a consumer that has not said which, naming the candidates and the command,
rather than silently choosing one ([ADR 0024](0024-model-access-is-a-provision.md), and
[`14-model-access`](../03-DESIGN/01-to-be/14-model-access.md)).
- **Subscription-seat authentication lives entirely inside the adapter**, never in the generic core.
So does an interactive `/login` — an adapter-specific "adopt" origin for a credential a person must
produce in a browser; the generic `licence key <name>` covers the static-key case.
- **Anthropic is the first `refreshable-grant` adapter**, carrying the OAuth refresh, the usage
reading, the identity guard and access-token-only delivery. **`anthropic-api-key` is a
`static-key` adapter for the same vendor's plain API keys**, and is the early second case that
proves the abstraction is not a single vendor wearing a coat: it exercises the whole path with the
carve-out switched off.
## Consequences
- **The carve-out is the mesh's one deliberate relaxation of "it cannot read what it stores."** It is
bounded to refreshable-grant vendors, to the refresh token, and to the manager node; the static-key
majority keep the full guarantee unchanged. This is the open risk the analysis carried here, and it
is recorded as an exception rather than pretended away.
- **Anthropic collapses from special case to adapter.** The three binding columns
(`nodes.node_license`, `nodes.hal_claude_account`, `agents.claude_account`) become three ordinary
consumers of `model-access`; the `claude_*` schema becomes the generic licence tables plus one
adapter. What was hardcoded becomes data keyed by `vendor`.
- **Adding a vendor is adding an adapter, and a static-key vendor is nearly free.** The modules that
want a model do not change when a vendor is added — they named `model-access`, not a vendor.
- **The refresh-token concentration is now a stated property to defend, not an accident.** The manager
node is a place a long-lived secret lives readably, and losing it or compromising it is a bounded,
named blast radius rather than a surprise.
### How each claim here is checked
- **Vendor-blind provision, static-key path.** A lab scenario binds an `anthropic-api-key` licence to
a consumer; the consumer resolves, receives a key sealed to its node, and reaches a model — and the
key is **nowhere in the control plane's database** nor in anything that crossed the broker. This is
the `licence_holder` sealed-per-node check that [`14-model-access`](../03-DESIGN/01-to-be/14-model-access.md)
already runs, now asserted for a second vendor.
- **The carve-out is exactly as narrow as stated.** For a `refreshable-grant` licence, a test asserts
the refresh token exists (encrypted) **only on the manager node**, is **absent from every holder's
delivery**, and that the delivered credential is access-token-only — and that for a `static-key`
licence no refresh token is stored anywhere.
- **Adapter selection is keyed by `vendor`.** A scenario with two vendors on two licences verifies each
licence's lifecycle runs its own adapter, and that a consumer naming `model-access` never names a
vendor to get one.
- **Refuse-if-unchosen.** Already checked in [`14-model-access`](../03-DESIGN/01-to-be/14-model-access.md):
a consumer with more than one candidate is refused with the candidates and the command named.
## References
- [ADR 0024](0024-model-access-is-a-provision.md) — model access is a provision, a licence is a named
thing, and `accept`; this record generalises its single vendor and keeps its working central
rotation.
- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — a provision names the
coupling; `model-access` is drawn at the consumer's.
- [ADR 0040](0040-what-a-module-is.md) — the naming rule and the neutral-interface / scoped-provider
shape a vendor adapter follows.
- [ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md) — registrar-scoped `public-dns`
providers, the precedent a `vendor`-scoped adapter mirrors.
- [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md) — the mesh seals credentials
and holds no readable copy; the carve-out here is the bounded, named exception to that for a
refreshable grant.
- [`03-DESIGN/01-to-be/14-model-access.md`](../03-DESIGN/01-to-be/14-model-access.md) — the design this
record extends, amended to describe the adapter generalisation.
- The read-only vendor-agnostic analysis, 2026-09-05 (code workspace) — the inventory and the decisions
taken on the open questions this record encodes.
@@ -0,0 +1,166 @@
---
topic: what runs on it
status: accepted
date: 2026-09-05
deciders: jochen
reconstructed: false
extends: 0030-data-outlives-the-mesh-that-declared-it.md
---
# 51. Shared data is the operator's, and a module is granted access to it
## Context
**Eight modules declared one filesystem as eight private ones.** The media stack — a library
server, the acquisition managers for films, series, music and books, a subtitle fetcher and two
download clients — shares directories on one machine: the download clients write into
`/services/media/downloads` and the managers read it; the managers write into the libraries and
the library server reads them. That sharing is the entire point of the stack. Yet each module
declared every shared directory it touched as its own `directory` resource, with an owner and a
mode. `/services/media/downloads` was written seven times, as seven private directories that
happen to be the same path.
**The resolver refuses exactly that, and is right to.** Two modules declaring one path on one node
are refused by name, with no exemption for identical content and no merge — because two owners of
one path is the class of fault this repository keeps recording ([04-ISSUES/036](../04-ISSUES/036-six-modules-own-what-they-must-share/00-report.md)).
So the stack as written refuses its own only sensible assignment: all of it on one machine,
sharing one filesystem. It passed today only because no test co-resolves any two of the eight. The
first machine assigned two of them together is where the refusal would have surfaced.
**The vocabulary had one word for two intentions, and this was already seen.**
[04-ISSUES/026](../04-ISSUES/026-the-data-directories-are-not-declared/00-report.md) found the
same gap from the other side and named it precisely: two kinds of mount are spelled identically —
*the directory my data lives in*, which the mesh creates and owns, and *a facility I was granted*,
which already exists and the mesh only reaches. That issue deferred inventing a field to tell them
apart, because doing so is a design decision and it declined to make one to get a check green. This
is that decision.
**A `directory` resource is owned, on every axis.** The host creates it, sets its owner and mode,
and removes it when it is empty and no longer declared — it *removes what it made and leaves what
it merely configured* ([ADR 0005](0005-the-node-host.md), [ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md)).
A shared media library is none of that. It existed before the mesh, several modules read and write
it at once, and losing it is the one failure that does not recover. It is not any module's
resource; it is the operator's, and a module only needs to be let at it.
## Considered Options
1. **Make the stack one module with several containers**, the way the mail system already is. The
issue raises it directly: is a set of modules that must share a filesystem really one module?
**Rejected.** The eight are independently assignable and independently useful — a person may run
the download client without the library server, or the film manager without the music one — and
folding them into a single module to express a shared directory would make *what a module is*
turn on an incidental filesystem contract. It also does not generalise: the next pipeline of
modules handing files to each other on one machine (an ingest folder, a spool, a drop directory)
would face the same wall and the same wrong remedy.
2. **One module owns the directories and the rest `require` them.** **Rejected**, and this is the
heart of the decision. Nobody owns shared operator data. The library predates the mesh and
outlives any one module, so making the library server or a manager its owner means unassigning
that module orphans everyone else's access — and the owner would set the owner and mode of a
tree it did not create. Ownership is the wrong relationship to model, because the true owner is
not a module at all.
3. **A flag on a directory resource** — `external: true`, or an owner of `operator`. **Rejected.**
It overloads one shape with a boolean that inverts every one of its semantics: created becomes
*must already exist*, owned becomes *touch nothing*, removed-when-empty becomes *never removed*.
That is the *two-kinds-spelled-identically* trap of 04-ISSUES/026 reintroduced with a single
quiet field — a reviewer reading `type: directory` would have to check one boolean elsewhere to
know whether the mesh owns the thing at all.
4. **A distinct `accesses` declaration, separate from resources.** **Adopted.**
## Decision
**Shared, pre-existing data is operator-owned and external. The mesh does not create it, does not
set its owner or mode, does not reconcile it and does not remove it.** A media library, a download
spool, an ingest directory is the operator's, and the mesh is a guest in it.
**A module declares that it needs *access* to such a path, not that it owns a resource there.** The
manifest field is `accesses`: a list of `{path, mode}`, where mode is `read` or `read-write` and
absent narrows to `read` — the safe default, because the danger with an access is being given more
than was meant, not less. A module's own configuration and state directories stay owned
`directory` resources; only the shared, pre-existing paths become accesses.
**The host mounts an accessed path and owns nothing about it.** It reaches the machine as a new
declaration shape, `access`, distinct from `directory`. The host confirms the path is present and
does nothing else — no create, no chown, no mode, no removal.
**An accessed path absent at apply time is refused, clearly, not created.** The mesh does not own
it, so conjuring it would be a lie the host then acts on — and specifically the lie 04-ISSUES/026
records, where a bind mount whose source does not exist is made by the container runtime as root
with the wrong ownership. The host says the operator must provide the path instead.
**Several modules accessing one path is normal, and never refused.** The duplicate-path refusal is
about *ownership*, not *use*: it applies to resources a module owns and to those alone. An access
is not a resource and never enters the check, so the eight-module stack co-resolves. What stays
refused is genuine rivalry — two modules owning one path — and the new contradiction it exposes: a
path one module owns while another merely accesses it, because that asserts both that the mesh owns
the directory and that the operator does.
This is a decision and not a patch because it settles *what a module may say about a path it did
not make*, which every co-located file-handoff in the catalogue now and later depends on — and
because it draws the ownership line [ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md)
started: the host owns what it made and keeps what it merely configured, and this adds the third
case it did not have a word for — what it neither made nor configured, and must not touch.
### How each claim is checked
- **The stack co-resolves.** A control-plane unit test assigns two modules that declare access to
one path on one node and asserts no refusal — the exact case the resolver refuses when the same
path is owned. The mirror test, two modules *owning* one path, still refuses, so the sharing
vocabulary does not weaken the rule it sits beside.
- **Ownership and access cannot both be claimed of one path.** A unit test asserts the resolver
refuses a path one module owns and another accesses, naming both.
- **Absent is refused, not created.** A host unit test applies an access to a path that does not
exist and asserts a clear refusal that names the operator, and that nothing was created.
- **Present is confirmed and nothing moves.** A host unit test applies an access to an existing
directory and asserts the apply reports no change and disturbs nothing.
- **Undeclaring never removes.** A host unit test drops a previously declared access and asserts
the operator's directory and its contents are left exactly as they were — the data-loss failure
[ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md) exists to prevent, on a directory the
mesh never made.
- **Every media manifest is corrected.** No `/services/media/*` path is an owned `directory`
resource in any of the eight; each is an `accesses` entry, and each module's own config and state
directories remain owned. Checked by the control-plane manifest parser, which now understands
`accesses` and refuses a malformed one.
## Consequences
**A shared filesystem between co-located modules now has a vocabulary**, and it is not the media
stack's alone: any pipeline handing files to a neighbour on one machine — an ingest directory, a
spool, a drop folder — says *I access this operator path* rather than *I own this directory*, and
several of them may say it of one path.
**Unassigning a module that reached shared data leaves the data.** Correct, and the same trade
[ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md) made for owned directories: removing
data is a person's act, done knowingly, not a side effect of unassignment.
**The operator must provision the shared paths before the stack is applied**, and a machine that
lacks one is told plainly which. That is a real new obligation, and it is the right one: the mesh
cannot own what predates it, so it cannot create it either, and saying so at apply time beats a
directory conjured as root and a service that half-works.
**The host vocabulary grew by one shape**, which is a cost — every added shape widens what a
compromised control plane can express ([ADR 0005](0005-the-node-host.md)). It is a narrow one: an
`access` is confirmed by a stat and grants the host no new action. It earns its place by letting
the host refuse to create what it must not own, which no existing shape could say.
**A path can be both owned and accessed only by refusal.** If a future manifest declares one path
as an owned directory in one module and an access in another, the resolver refuses it rather than
guessing which is meant — the two assertions about who owns the data cannot both hold.
## References
- [04-ISSUES/036](../04-ISSUES/036-six-modules-own-what-they-must-share/00-report.md) — six (in
fact eight) modules own what they must share; the problem this resolves
- [04-ISSUES/026](../04-ISSUES/026-the-data-directories-are-not-declared/00-report.md) — the two
kinds of mount spelled identically, which deferred this field to a decision
- [ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md) — data outlives the mesh; the host
keeps what it did not make. This record adds the case it had no word for
- [ADR 0005](0005-the-node-host.md) — the host removes what it made and leaves what it merely
configured; the vocabulary is finite and every shape is a security decision
- [ADR 0040](0040-what-a-module-is.md) — what a module is; an access is a new thing a module may
say about the machine it lands on
- mesh-control `feat/shared-data-access`, mesh-catalog `feat/media-access-not-ownership`,
mesh-host `feat/mount-operator-owned` — the mechanism, the corrected manifests, and the host
shape
@@ -0,0 +1,198 @@
---
topic: what runs on it
status: accepted
date: 2026-09-05
deciders: jochen
reconstructed: false
extends: 0005-the-node-host.md
---
# 52. An init step is a container run once to completion, gating what follows
## Context
**A module can declare things that exist; it cannot declare a step that runs.** The host owns a
finite vocabulary of shapes — `file`, `directory`, `service`, `package`, `container`, `action` —
and every one but `action` describes *state*: a thing that should be present, with content or a
mode or an image, which the host reconciles toward ([ADR 0005](0005-the-node-host.md)). That is
right for what it covers. But a real class of modules needs, once, to *run their own code at a
point in their own lifecycle* — and the vocabulary has no word for it
([04-ISSUES/037](../04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md)).
**mosquitto is the sharp case, and it fails silently without this.** Its Dynamic Security plugin
loads at broker start and refuses to come up unless `dynamic-security.json` already holds an admin
client. Seeding that file is a step that must happen *after* the data directory exists and *before*
the broker container starts. The manifest can declare the directory, the config file and the broker
container; it cannot declare "seed this, once, before that container starts." Written as it is
today, the broker starts against an unseeded store and the plugin aborts — and the next reconcile
does not fix it, because nothing in the declaration ever seeds the file.
**It is not one module's defect.** The database providers need the same to run a first-boot
migration, an extension enable, or a health gate before they are announced ready; today that works
only where the *image* happens to seed itself from an environment variable, and anything the mesh
must run once against the server has no home. This is the timing face of the same gap
[04-ISSUES/035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md) records
from the content side: the manifest needed *the file to exist before first start*, and had only
*the file has this content, forever*.
**The obvious answer is the one that already went wrong.** An earlier mesh had exactly this as a
feature — event-driven hooks that ran custom code at phases of build, publish and deploy. It was
powerful and it was *complex to set up and flaky*, and that fragility, not the need, is the content
of the issue. Whatever this becomes must not rebuild that engine.
**The ground has shifted since that engine, in a way that makes a much smaller answer possible.** A
module with tools or events now runs a **process of its own** — a container carrying the module's
compiled code, holding the single broker account the mesh scoped to it, isolated from every other
module ([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). The
code that must run at first boot is *already in that image, under that account*. So the mesh does
not need a way to run a module's code — it has one. It needs a way to say **run this container to
completion, and start the one that depends on it only after it has.**
## Considered Options
1. **A host `run` shape — a command the host executes on the machine.** The direct reading of
"run code at a lifecycle phase." **Rejected.** The host's vocabulary is finite and every added
shape is a security decision, because it widens what a *compromised control plane* can express
([ADR 0005](0005-the-node-host.md)). A general "run this command as the host" is the largest
such widening there is: the blast radius is the whole machine, as root. The mesh already drew
this exact line for `action` — it runs a command, and so it is *permitted from the bundle and
refused from the link*, because the bundle arrives with the binary and the link is a separate
party with an unbounded reach. A new host-command shape usable by an ordinary module would be an
`action` from the link by another name, which is precisely what is refused.
2. **Per-phase lifecycle hooks on a module** — `pre-start`, `post-start`, `pre-remove`, and their
build/publish cousins, each naming code the mesh runs at that phase. The general answer, and the
old feature. **Rejected for now.** It is the flaky engine the issue warns against, and most of
its phases have no present need. Deciding the full set of phases, where each one's code runs, and
how each is made idempotent is a large design taken to buy capability nothing yet asks for. The
three blocked modules all need one phase — *before a container starts* — and a mechanism narrow
enough to be obviously correct beats a general one that is not.
3. **A distinct one-shot resource type** — a new shape, sibling to `container`, that names an image
and runs it once. **Rejected.** It grows the host vocabulary by a whole shape (a `Type`, a
struct, an applier, a place in every host's shape list) to express something a `container` almost
already is. A one-shot *is* a container — a pinned image, an account, volumes, an environment —
that happens to exit. Spending a new shape on the difference is the cost of option 1 in smaller
type, for a capability the existing shape can carry with one modifier.
4. **A modifier on the existing `container` shape: this container runs once, to completion, and the
host gates the apply on it.** **Adopted.** It reuses the shape the host already has, adds no new
host action, and leans on two guarantees the host already gives — *apply in declared order,
never sorted*, and *a failed step fails the apply* — to turn "before that container starts" into
an emergent property of ordering rather than a dependency graph the host must resolve.
## Decision
**A run-once step is an ordinary `container`, marked to run to completion.** The manifest sets
`run-once: true` on a container resource. Everything else about it is a container as before — a
digest-pinned image ([ADR 0006](0006-the-substrate-and-the-control-plane.md)), volumes, an
environment, and for a module's own code the same scoped account its runtime already holds
([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). The mesh adds
no way to run code; it marks a container the host must **run to completion and require to exit 0**,
rather than start and leave running.
**The gate is declaration order, not a named dependency.** The host applies a declaration in the
order it is given, does not sort, and does not resolve dependencies — ordering is a decision, and it
is the control plane's ([ADR 0005](0005-the-node-host.md),
[04-ISSUES/013](../04-ISSUES/013-a-file-arrives-after-the-service-that-needs-it/00-report.md)). A
run-once step is placed *before* the container that depends on it, and **a failed run-once step
halts the apply**, exactly as a failed `action` does — so everything the declaration places after
it, the broker included, is never reached until the step has completed. "Before the broker starts"
is therefore expressed by list position plus completion, and the host cross-references nothing.
**Completion is recorded, and a re-apply does not re-run it.** The host records what it applied only
after the fact, as the digest of the declaration that produced it
([ADR 0018](0018-a-picture-is-read-from-what-runs.md)) — a run-once step no differently. Because the
step leaves nothing running to inspect, that persisted digest, not a live container, is the marker
that it happened. On a later apply the host finds the digest already recorded for this exact
declaration and does nothing; it re-runs only when the declaration's digest has changed, and a step
that exited non-zero recorded nothing and so is retried next apply. This is the reconcilable,
idempotent discipline the state shapes get for free, made explicit for a step.
This is a decision and not a patch because it settles **what a module may say about running its own
code**, which the whole catalogue of providers — a seed before start, a first-boot migration, a
health gate — now and later depends on, and because it draws the line the issue asked for: the
narrowest sound mechanism that unblocks the three modules without rebuilding the hook engine whose
fragility is the warning.
### The security bound, stated plainly
**The host gains no new action and no new shape.** `run-once` is a boolean modifier on the
`container` shape that already exists. A run-once container is strictly *less* powerful than an
`action`: it cannot run an arbitrary host command, only a digest-pinned image under an account the
mesh scoped — which is exactly the capability `container` already grants from the link. A control
plane that is compromised can express nothing through `run-once` it could not already express by
declaring an ordinary `container`. The dangerous expansion of option 1 — a command the host runs on
the machine — is not made.
### How each claim is checked
- **A run-once step runs to completion and its exit 0 is required.** A host unit test applies a
run-once container whose image exits 0, asserts the host ran it to completion (not detached, not
left running) and reported it done; a sibling test applies one that exits non-zero and asserts
the apply fails, naming the step.
- **A failed run-once step gates what follows.** A host unit test places a run-once container that
exits non-zero before another container and asserts the second is never started and the apply is
reported gated — the mirror of the existing test that a failed action stops what follows.
- **It is not re-run once it has completed.** A host unit test applies a run-once step, then applies
the identical declaration again with the first run's record present, and asserts the second apply
runs nothing and reports the step unchanged.
- **A changed declaration re-runs it.** A host unit test applies a run-once step, then applies one
whose image or environment differs, and asserts it runs again — the digest moved, so the marker no
longer matches.
- **The vocabulary carries the field end to end.** A control-plane unit test resolves a module whose
manifest marks a container `run-once` and asserts the rendered host declaration carries the field,
in author order before the container it gates; the manifest parser refuses a `run-once` that is
not boolean.
- **mosquitto seeds before the broker.** mosquitto's manifest declares a run-once init container,
before the `server` (broker) container, that writes the admin client into `dynamic-security.json`
and exits — checked by the control-plane resolver, which now understands the field, and by the
ordering of the rendered declaration. The end-to-end proof that it runs exactly once, at the right
phase, and converges on re-apply is owed to a lab scenario ([04-ISSUES/037] open question), which
this record does not close.
## Consequences
- **The three blocked modules gain a home for their step.** mosquitto seeds its dynsec admin before
the broker; a provider that must migrate or health-gate at first boot declares a run-once step in
its own runtime image, under its own account, before the container that depends on it.
- **The general lifecycle hook is deferred, deliberately.** Only *before a container starts* is
bought here. `post-start`, `pre-remove` and the build/publish phases remain unbuilt, and the day
one is genuinely needed it is decided then, against a need, not speculatively — the same restraint
that kept this from being the old engine.
- **The seed-then-mutate file is safe if the step is written to be.** A run-once seed writes
`dynamic-security.json` only when it is absent and never reconciles it, so what the running plugin
grows in that file afterward is never wiped
([04-ISSUES/035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md)).
The host's marker guarantees the step is not re-run; the step's own code guarantees it does not
clobber on the pass it does run.
- **The host vocabulary did not grow, and that is the point.** The cost of a run-once step is one
boolean and a completion path in the container applier, not a new shape and not a new action. The
mesh expresses ordering and completion; the module runs its own code, where it already runs it.
- **A run-once step that never converges is a stuck apply, loudly.** A step that exits non-zero
every time halts the apply every time, and the container it gates never starts — which is the
correct failure, reported, rather than a broker that half-starts against an unseeded store and a
reconcile that reports success. It is failed forward, not failed silent.
## References
- [04-ISSUES/037](../04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md) — a
module cannot run its own code at a lifecycle phase; the gap this resolves, and the warning about
the old hook engine
- [04-ISSUES/035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md) — the
content face of the same gap: a file needed before first start, that the running program then
mutates
- [04-ISSUES/013](../04-ISSUES/013-a-file-arrives-after-the-service-that-needs-it/00-report.md) — the
order of a declaration is the control plane's, and the host applies it as given; the gate rests on
this
- [ADR 0005](0005-the-node-host.md) — the host's finite vocabulary, the ordered declaration it does
not sort, and `action` as the shape a command already is and why it is refused from the link
- [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) — a module runs
its code as its own process under its own account; a run-once step is that process, run to
completion
- [ADR 0018](0018-a-picture-is-read-from-what-runs.md) — what was applied is recorded after it works;
the completion marker is that record
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — a container is pinned by digest; a
run-once container no differently
- mesh-control `feat/lifecycle-run-once`, mesh-host `feat/apply-run-once`, mesh-catalog
`feat/mosquitto-bootstrap` — the vocabulary, the apply support, and mosquitto's seeded broker
@@ -0,0 +1,181 @@
---
topic: what runs on it
status: accepted
date: 2026-09-06
deciders: jochen
reconstructed: false
extends: 0052-a-step-that-runs-once-before-a-container.md
---
# 53. A scheduled step is a container run on a recurring schedule
## Context
**[ADR 0052](0052-a-step-that-runs-once-before-a-container.md) gave the mesh a step that runs *once*;
a real class of modules needs one that runs *again and again*.** kometa reconciles a media library
against its lists on a timer; a ticketing integration polls its source for new work every few minutes;
a backup, a cache warm, a metrics roll-up all recur. The host's vocabulary describes *state* — a file,
a directory, a container that should be running — and 0052 added *a step that happens once and is
done*. Neither says *this should happen every night at 3, forever*
([04-ISSUES/037](../04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md) named
the lifecycle gap; 0052 closed the once-before-start face of it and left the recurring face open).
**The modules that need it already run their own code.** As with run-once, the ground has shifted
since the old mesh's flaky hooks: a module with tools or events runs a **process of its own** — a
container carrying the module's compiled code under the single scoped account the mesh gave it
([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). The work that
must recur is already in that image, under that account. The mesh does not need a way to run a
module's code on a timer — it has the code and the account. It needs a way to say **run this
container again on this cadence**.
**The shape of the answer is already decided, one modifier over.** 0052 rejected a host `run` command,
per-phase hooks, and a distinct one-shot resource type, and adopted *a modifier on the `container`
shape the host already has*, because a one-shot is a container that happens to exit. A scheduled step
is the same container that happens to exit — run repeatedly. Reusing that shape keeps the host
vocabulary flat and inherits 0052's security bound whole. The only thing 0052's `run-once` does not
carry is *when to run it again*.
**One difference from run-once changes a rule, and it is the reason this is its own record.** A
run-once step **gates the apply**: it is placed before the container that depends on it, and a failure
halts everything after it, because "seed the store before the broker starts" is a correctness
precondition ([ADR 0052](0052-a-step-that-runs-once-before-a-container.md)). A scheduled step is the
opposite: it runs *after* the machine is up and converged, on its own clock, and a single failed run
is an ordinary operational event — the next run comes anyway. A scheduled step that halted the apply,
or that a failed run marked the node not-current over, would make a routine poll into a reason the
whole machine reads as broken. So the gating rule 0052 established is exactly the rule this record
must **not** inherit.
## Considered Options
1. **A host `cron`/`timer` shape — the host installs a system timer that runs a command.** The direct
reading. **Rejected**, for the reason 0052 rejected a host `run` shape: it widens what a
*compromised control plane* can express toward "run this command on the machine, forever," which is
the largest widening there is, and it is `action`-from-the-link by another name
([ADR 0005](0005-the-node-host.md)). A recurring command is worse than a one-off, because it
persists.
2. **Per-phase lifecycle hooks** — `on-schedule` joining `pre-start`/`post-start` as named code the
mesh runs. **Rejected for now**, as in 0052: it is the flaky hook engine the issue warns against,
and the three modules that need this need one thing — *run this container on a cadence* — which a
narrow modifier expresses without deciding a whole hook vocabulary.
3. **A distinct `scheduled` resource type**, sibling to `container`. **Rejected**, as 0052 rejected a
distinct one-shot type: it spends a whole new host shape (a `Type`, a struct, an applier, a place
in every host's shape list) on something a `container` already almost is — a scheduled task *is* a
container (pinned image, account, volumes, environment) that runs on a clock.
4. **A modifier on the existing `container` shape: `schedule`, a cron expression the host runs the
container on.** **Adopted.** It reuses the shape the host has, adds no new host action, and sits
beside `run-once` as its recurring twin — the same container, exited, run again.
## Decision
**A scheduled step is an ordinary `container`, marked with a `schedule`.** The manifest sets
`schedule: "<cron>"` on a container resource — a standard five-field cron expression. Everything else
about it is a container as before: a digest-pinned image
([ADR 0006](0006-the-substrate-and-the-control-plane.md)), volumes, an environment, and for a module's
own code the same scoped account its runtime already holds
([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). The mesh adds no
way to run code; it marks a container the host must **run on that cadence, each time to completion**,
rather than start once and leave running (a service) or run once and gate (a run-once step).
**A run is fired by the clock, not by the apply, and does not gate it.** Applying the declaration
installs the schedule; it does not run the step. The machine converges — reports applied and current —
as soon as the schedule is installed, exactly as it does for a service that is running. Thereafter the
host fires the container when the cron expression is due. This is the deliberate inversion of 0052:
a scheduled step is downstream of convergence, not a precondition of it.
**A failed run is recorded and the next run still comes; it never marks the node not-current.** A run
that exits non-zero is logged against the module — visible, auditable — but it does not fail the apply,
does not halt other resources, and does not flip the node's reported state. A poll that fails at 03:00
and succeeds at 03:05 is the system working, not a machine that needs attention. A step that fails
*every* time is a loud, repeating log entry, which is the correct signal for "this recurring job is
broken" — distinct from "this machine did not converge."
**Runs do not stack.** If a run is still going when the next is due, the host skips the due run rather
than starting a second copy, and logs the skip. A slow nightly job that occasionally overruns must not
spawn a growing pile of concurrent containers competing for the same account and volumes — the failure
mode that made the old timers dangerous.
**Each run is independent and idempotent by the module's own code.** The mesh guarantees only *the
container is run on the cadence*; that a run does the right thing when the previous one half-finished
is the module's contract, the same discipline a run-once seed owes
([ADR 0052](0052-a-step-that-runs-once-before-a-container.md),
[04-ISSUES/035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md)).
This is a decision and not a patch because it settles **what a module may say about running its own
code on a cadence**, which every recurring provider job — a sync, a poll, a roll-up — now and later
depends on, and because it draws the line 0052 left: the recurring twin of run-once, with the gating
rule deliberately reversed so a routine job's failure is never a machine's failure.
### The security bound, stated plainly
**The host gains no new action and no new shape.** `schedule` is a string modifier on the `container`
shape that already exists. A scheduled container is strictly *less* powerful than an `action`: it runs
a digest-pinned image under an account the mesh scoped, which is exactly what `container` already
grants from the link, and it cannot run an arbitrary host command. A compromised control plane can
express nothing through `schedule` it could not already express by declaring a `container` — the cron
string only says *how often*, not *what*. The dangerous expansion of option 1 — a command the host
runs on the machine on a timer — is not made.
### How each claim is checked
- **A scheduled step runs when the schedule is due.** A host unit test installs a container with a
schedule that is due immediately (or advances a injected clock to when it is due) and asserts the
host ran it to completion; a sibling test with a schedule not yet due asserts it has not run.
- **Installing it does not run it, and the node is current without a run.** A host unit test applies a
scheduled container and asserts the apply reports current *before* any run has fired — the schedule
is state that is present, not a step that gated.
- **A failed run does not fail the apply or the node.** A host unit test fires a scheduled container
that exits non-zero and asserts the failure is recorded against the module, the apply is not failed,
and the node stays current — the mirror of the run-once test where a non-zero exit *does* halt.
- **Runs do not stack.** A host unit test fires a scheduled container whose run outlasts its next due
time and asserts the host skipped the due run and logged the skip, rather than starting a second
container.
- **The vocabulary carries the field end to end.** A control-plane unit test resolves a module whose
manifest sets `schedule` on a container and asserts the rendered host declaration carries the field;
the manifest parser refuses a `schedule` that is not a valid cron expression, and refuses a container
that is both `run-once` and `schedule` (a step is one or the other, never both).
- **A real module recurs in the lab.** A converted module declaring a scheduled step (kometa's library
sync, or a poller) is assigned in a lab scenario, and the scenario asserts the scheduled container
fires on its cadence and its account and volumes are the module's — the end-to-end proof this record
owes, as 0052 owed its run-once lab proof.
## Consequences
- **The recurring providers gain a home for their cadence.** kometa reconciles on its schedule; a
poller polls; a roll-up rolls up — each a scheduled step in its own runtime image, under its own
account, on the cron it declares.
- **The general lifecycle hook is still deferred.** Only *run once before* (0052) and *run on a
cadence* (this) are bought. `post-start`, `pre-remove` and the build/publish phases remain unbuilt,
decided when a real need arrives, not speculatively — the restraint that kept both from being the old
engine.
- **A recurring job's failure is loud but not fatal.** The node stays current while a scheduled step
fails and retries; a step that fails forever is a repeating log entry, not a machine marked broken.
This is the correct separation — a machine's convergence and a job's success are different questions —
and it is why this could not simply be `run-once` without the schedule.
- **The host vocabulary did not grow, again, and that is the point.** The cost of a scheduled step is
one string field and a cron loop in the container applier, not a new shape and not a new action. The
mesh expresses cadence; the module runs its own code, where it already runs it.
- **`run-once` and `schedule` are exclusive and complete for now.** A container runs once and gates, or
runs on a cadence and does not, or runs and stays up (a service). A manifest that asks for two of
these at once is refused, because the three are distinct answers to "how does this container run."
## References
- [ADR 0052](0052-a-step-that-runs-once-before-a-container.md) — the run-once step; this is its
recurring twin, reusing the `container`-modifier shape and inheriting its security bound, and
reversing its gating rule
- [04-ISSUES/037](../04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md) — a
module cannot run its own code at a lifecycle phase; 0052 closed the once-before-start face, this
closes the recurring face
- [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) — a module runs
its code as its own process under its own account; a scheduled step is that process, run on a cadence
- [ADR 0005](0005-the-node-host.md) — the host's finite vocabulary, and why a host command (even on a
timer) is refused from the link
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — a container is pinned by digest; a
scheduled container no differently
- [ADR 0018](0018-a-picture-is-read-from-what-runs.md) — what was applied is recorded after it works;
the installed schedule is state, a fired run is an event
- mesh-control `feat/schedule-container`, mesh-host `feat/apply-schedule`, and the converted module
that first declares a scheduled step — the vocabulary, the apply support, and the recurring proof
@@ -0,0 +1,178 @@
---
topic: what runs on it
status: accepted
date: 2026-09-06
deciders: jochen
reconstructed: false
extends: 0050-model-access-is-vendor-agnostic.md
---
# 54. Model usage is a vendor-neutral record, produced by the adapter, at two grains
## Context
**[ADR 0050](0050-model-access-is-vendor-agnostic.md) fixed the *shape* of a usage reading and left
its *home* open.** It decided that an adapter may expose `usage(licence) → normalised rows`, that a
row is `(licence, consumer, period, metric, value)` plus the raw vendor response as jsonb, and that the
metric is vendor-defined (Anthropic's `utilization%` is one metric, a token count another). It did not
say where those rows are stored, how they get produced on a cadence, or whether the only grain is the
whole licence — and the mesh's first real consumer needs all three answered.
**The predecessor mesh recorded usage at two grains, and both are wanted.** It polled the vendor for a
licence-level reading (an account's `utilization%`), and it also attributed **per-session** token and
cost — which model, how many input and output tokens, what it cost — parsed from the agent's own
transcript, including the case where a long session switched the account it billed against mid-way. The
licence-level reading answers "how close is this subscription to its cap"; the session-level reading
answers "what did this piece of work cost, and against which account." A mesh that kept only the first
could not bill a project or notice a runaway session; keeping only the second could not see a cap
approaching. Both are load-bearing and neither subsumes the other.
**A session is already a consumer, so the second grain needs no second vocabulary.**
[ADR 0026](0026-the-mesh-has-a-session-of-its-own.md) and design 15 establish that an agent session is
a thing in its own right, and that **model access is bound to the session, not to the machine** — a
session is a *consumer* of `model-access`, identified by `(node, module)` and its own id. That is
exactly the `consumer` column ADR 0050 already put in the usage row. So the two grains are not two
schemas; they are the same row at two consumer resolutions: the holding module for the licence grain,
the session for the finer one. The design's own still-open worker-naming gap (design 14) is the same
gap here and is left where it is — a session id distinguishes what `(node, module)` cannot.
**The pieces this needs already exist.** A periodic reading is a **scheduled step**
([ADR 0053](0053-a-step-that-runs-on-a-schedule.md)) — the adapter's `usage()` poll is a container the
mesh runs on a cadence, which is precisely what 0053 was built for. A durable audit of what happened is
**an event the audit trail records** ([ADR 0041](0041-events-are-a-relationship.md),
[ADR 0042](0042-the-shape-of-an-event-on-the-wire.md)); the `audit-logger` module already consumes every
event. And a queryable current picture is **a context store**, the same shape the licences themselves
live in ([ADR 0008](0008-a-context-owns-its-store.md)). Nothing new in kind is required; what is missing is
the decision to point them at usage.
## Considered Options
1. **Licence grain only — a single `utilization%` poll, nothing per session.** Rejected: it cannot
attribute cost to a piece of work or catch a session that is burning an account down, which is half
of why usage is recorded at all.
2. **A bespoke `sessions` schema mirroring the old mesh's `*_sessions` / `*_session_account_usage`
tables.** Rejected: it reintroduces a second, vendor-shaped vocabulary for something the mesh
already names — a session is a consumer, and its usage is a usage row. A parallel schema would drift
from the `model-access` vocabulary and force every reader to learn two.
3. **Store usage only as raw vendor blobs, normalise later.** Rejected as the *whole* answer (kept as a
fallback within the chosen one): a reader that must parse Anthropic's response shape to answer "what
did this cost" has the vendor coupling the whole feature exists to remove. The raw blob is kept
beside the normalised row (0050 already requires this), not instead of it.
4. **One vendor-neutral usage record at two consumer grains, produced by the adapter, recorded as
both an event and a queryable row.** Adopted.
## Decision
**Model usage is one vendor-neutral record — `(licence, consumer, period, metric, value)` plus the raw
response — recorded at two grains that differ only in the `consumer`.** At the **licence grain** the
consumer is the holding module and the metric is the vendor's own account reading (Anthropic:
`utilization%`). At the **session grain** the consumer is the agent session — `(node, module)` and its
session id ([ADR 0026](0026-the-mesh-has-a-session-of-its-own.md)) — and the metrics are the ones a
session bills: input tokens, output tokens, model, and cost. The row shape is 0050's, unchanged; the
grain is which consumer the row is *for*.
**The adapter is the only thing that knows the vendor, and it produces both grains.** Reading an
account's cap is the adapter's `usage(licence)` verb ([ADR 0050](0050-model-access-is-vendor-agnostic.md));
attributing a session's cost is the adapter reading that vendor's transcript or usage API and emitting
rows keyed to the session. The mesh defines the row and the plumbing; the adapter fills it from whatever
the vendor exposes, and a static-key vendor that exposes nothing simply produces no rows — usage is an
optional reading, not a requirement of holding a licence.
**A reading is taken on a schedule, not on a request.** The licence-grain poll is a **scheduled
container** ([ADR 0053](0053-a-step-that-runs-on-a-schedule.md)) the adapter runs on a cadence; the
session-grain rows are produced as sessions progress, from the transcript the session already writes.
Neither blocks anything: a poll that fails is a logged, retried scheduled run (0053's rule), and a
session whose cost cannot yet be attributed is a row not yet written, never a session refused.
**Usage is recorded two ways, for two audiences.** Each reading is **emitted as an event**
([ADR 0041](0041-events-are-a-relationship.md)) — an immutable "this was observed at this time" that the
`audit-logger` already records, so the history of what an account did is in the audit trail by default,
under nobody's special arrangement. And the **current** picture — the latest reading per
`(licence, consumer, period, metric)` — is upserted into a **usage context store**, so "how close is
this cap" and "what has this project spent this month" are a query, not a fold over the event log. The
event is the record of what happened; the store is the answer to what is true now.
**Usage is not a credential, and is recorded in the clear.** The one thing the mesh must not read is the
sealed key ([ADR 0050](0050-model-access-is-vendor-agnostic.md), the refresh-token carve-out aside). A
token count and a cost are not secrets; they are the operator's own operational facts, recorded openly
so they can be queried, audited, and charged against. This is the deliberate opposite of the credential
rule, and stating it prevents a later reader assuming usage inherits the key's secrecy and hiding it
from the person who is paying.
This is a decision and not a patch because it settles **where a model's usage lives and at what grain**,
which every reader — a bill, a cap alarm, a per-project report — depends on, and because it closes the
half [ADR 0050](0050-model-access-is-vendor-agnostic.md) explicitly left open, reusing the session
([ADR 0026](0026-the-mesh-has-a-session-of-its-own.md)), the schedule
([ADR 0053](0053-a-step-that-runs-on-a-schedule.md)), the event ([ADR 0041](0041-events-are-a-relationship.md))
and the store ([ADR 0008](0008-a-context-owns-its-store.md)) the mesh already has rather than inventing a
vocabulary beside them.
### How each claim is checked
- **A usage row is vendor-neutral and carries the raw beside it.** A unit test constructs an Anthropic
`utilization%` reading and a token/cost reading and asserts both render to
`(licence, consumer, period, metric, value)` with the vendor response preserved in the raw column;
a reader that answers "what did this cost" touches only the normalised columns.
- **The two grains differ only in the consumer.** A unit test records a licence-grain row (consumer =
the module) and a session-grain row (consumer = a session id) for one licence and asserts both are
the same shape and both are returned when the licence's usage is asked for, distinguishable by
consumer.
- **A poll is a scheduled run and its failure is not fatal.** The adapter's usage container declares a
`schedule` ([ADR 0053](0053-a-step-that-runs-on-a-schedule.md)); a host test (0053's) already proves a
scheduled step runs on cadence, does not gate, and logs rather than fails on a non-zero run — the poll
inherits this and adds nothing to check.
- **Each reading is an event the audit trail records.** An integration check asserts a usage reading
emits an event that the `audit-logger` receives (it consumes `#`), so the history is present without
the usage module and the audit module knowing about each other beyond the event.
- **The current picture is a query.** A store test upserts two readings for one
`(licence, consumer, period, metric)` and asserts the later replaces the earlier, so "what is true
now" is one row, while the event log keeps both.
- **A static-key vendor with no usage reading records nothing, and that is fine.** A test resolves a
`static-key` model-access consumer whose adapter has no `usage` verb and asserts the licence works and
no usage rows or poll are required — usage is optional, holding a licence is not conditioned on it.
- **Usage is readable in the clear; the key is not.** A test asserts a usage row is stored unsealed and
is returned to an ordinary query, while the licence key remains sealed and absent from the same
surfaces — the deliberate inversion of the credential rule.
## Consequences
- **A bill and a cap alarm are both queries.** "What did project X spend this month" reads the
session-grain rows; "how close is account Y to its cap" reads the latest licence-grain metric — both
from the usage store, neither a fold over events or a call to the vendor.
- **The session becomes the unit of cost, which is what it already is.** Because a session is the
consumer, attributing cost needs no new identity — and the worker-granularity gap
([design 14](../03-DESIGN/01-to-be/14-model-access.md)) surfaces here exactly as it does for access,
to be closed once, for both, when a session id is threaded through.
- **The adapter carries the vendor's usage quirks alone.** Anthropic's `utilization%`, its transcript
shape, a mid-session account switch — all live in the Anthropic adapter; the mesh, the store, and
every reader see only rows. A second vendor adds a second adapter and no new table.
- **Usage history is durable and tamper-evident by reuse, not by a new mechanism.** It rides the event
trail the mesh already keeps, so an operator who wants the whole history has it, and one who wants the
current number has the store — without usage owning either mechanism.
- **The refresh-token carve-out is untouched by this.** Usage is read *from* an authenticated adapter;
it neither holds nor exposes the credential, so the one place the mesh reads what it stores
([ADR 0050](0050-model-access-is-vendor-agnostic.md)) is not widened by recording what that credential
was spent on.
## References
- [ADR 0050](0050-model-access-is-vendor-agnostic.md) — model access is vendor-agnostic; fixes the
usage row shape and the `usage(licence)` adapter verb, and leaves its home open — which this closes
- [ADR 0026](0026-the-mesh-has-a-session-of-its-own.md) — the mesh has a session of its own; a session
is the consumer the finer grain attributes to
- [ADR 0053](0053-a-step-that-runs-on-a-schedule.md) — a scheduled step; the licence-grain poll is one
- [ADR 0041](0041-events-are-a-relationship.md) — events are a relationship; a usage reading is one, and
the audit-logger records it
- [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) — the shape of an event on the wire; the form a
usage reading takes to reach the audit trail
- [ADR 0008](0008-a-context-owns-its-store.md) — a context owns its store; the current usage picture
lives in one
- [ADR 0024](0024-model-access-is-a-provision.md) — model access is a provision; usage is a reading of
what that provision was used for
- [03-DESIGN/01-to-be/14-model-access.md](../03-DESIGN/01-to-be/14-model-access.md) — the model-access
design; this fills its usage section and shares its open worker-naming gap
- [03-DESIGN/01-to-be/15-the-agent-session.md](../03-DESIGN/01-to-be/15-the-agent-session.md) — the agent
session; the consumer the session grain is keyed to
@@ -0,0 +1,146 @@
---
topic: what runs on it
status: accepted
date: 2026-09-07
deciders: jochen
reconstructed: false
extends: 0050-model-access-is-vendor-agnostic.md
---
# 55. Model access is answered by a licence, or by a node that hosts the model
## Context
**[ADR 0024](0024-model-access-is-a-provision.md) made model access a provision, and
[ADR 0050](0050-model-access-is-vendor-agnostic.md) fixed what answers it: a record — a licence — that
an adapter turns into a sealed vendor credential.** A consumer requires `model-access`, is put on a
licence, and is delivered a key (a static API key, or an access token a manager refreshes). Every
answer so far has been a credential to reach a vendor's API across the internet.
**But a model need not come from a vendor. A node in the mesh can host one.** An operator with a GPU
runs Ollama or vLLM, which serves an OpenAI-compatible API on that node. A consumer that wants that
model does not need a vendor credential — it needs the model server's **endpoint**: the base URL and
the model name, and a key only if the server is configured to want one. This is model access answered
by a **node**, not by a record.
**The mesh already knows how a node answers a provision — it is the ordinary provider/consumer path.**
A provider `provides` a provision at a scope, `serves` its connection facts, and the mesh fills the
consumer's bound facts with the provider's `at`/`port` and each served fact — exactly how a Postgres
consumer learns where its database is. Nothing reserved `model-access` to records: a node offering
`provides: ["model-access"]` resolves through this path, and the resolver already **prefers a local
answer over a licence** — its own comment names the case, "a model the mesh runs itself." So the
capability exists; what is missing is the decision to use it, and the statement of what a node-answer
delivers and where it stops.
**A node-answer and a record-answer are the same provision with two shapes of answer.** This is the
same move [ADR 0050](0050-model-access-is-vendor-agnostic.md) already made for shapes within the
vendor path (static-key vs refreshable-grant): one provision, more than one way it is answered. A
consumer written against `model-access` should not care whether the model behind it is a vendor's or
the mesh's own — it asks for model access and is given what reaches a model.
## Considered Options
1. **A separate provision for the local case (`local-model`, `model-endpoint`).** A node answers that;
`model-access` stays record-only. Rejected: it splits "where my model comes from" into two
provisions a consumer must choose between in its manifest, when the mesh already models a
record-answer and a node-answer to **one** provision. A consumer would have to know, at authoring
time, whether its model will be a vendor's or the mesh's — the exact coupling the provision was
meant to remove. It is the safer implementation (see the limitation below) but the worse interface.
2. **Model access answered by either a licence or a node, under the one provision.** A consumer
requires `model-access`; the operator answers it with a licence (a vendor) or by assigning a
node that hosts a model. Adopted: one interface, and the answer is an operator's deployment choice,
not a consumer's authoring choice.
## Decision
**Model access is one provision answered two ways: by a licence (a record, turned into a sealed vendor
credential by an adapter — [ADR 0050](0050-model-access-is-vendor-agnostic.md)) or by a node that hosts
the model (a provider that serves an endpoint).** A consumer requires `model-access` and is delivered
whichever the operator assigned; it does not name the kind.
**A node-answer delivers an endpoint, not a credential.** The provider `provides: ["model-access"]`
and `serves` its connection facts — the port it listens on and the model it runs — and the mesh fills
the consumer's bound facts with the provider node's `at`, the served `port`, and the served `model`,
the same way every provider consumer learns where its provider is. The consumer assembles a base URL
(`http://<at>:<port>/v1`) and points an OpenAI-compatible client at it. If the local server wants a
key, the provider mints one the ordinary way (a per-consumer secret, sealed and host-unsealed); if it
does not — the common Ollama case — the consumer lists `model-access` under `binds` and **not** under
`secrets`, and no key is delivered. Secret delivery and fact delivery are already independent, so a
keyless endpoint is expressed by asking for the facts and not a secret.
**A node-answer uses no adapter.** The adapter registry ([ADR 0050](0050-model-access-is-vendor-agnostic.md))
is the vendor-credential machinery — accept-and-seal, refresh, usage. A node-hosted model has no vendor
secret to seal; its endpoint is served, and its key (if any) is minted like any provider's. The adapter
is consulted only for the record/vendor answer. So the vendor-agnostic decision is untouched, and the
node-answer adds no vendor logic anywhere.
**The resolver prefers a local answer.** When a node's own set answers `model-access` — a model the
mesh runs itself — a licence for it is not consulted. This is already the resolver's behaviour and is
made a decision here: a mesh that runs a model uses it, and a licence is the answer for a consumer that
has no local model, not a competitor to one that does.
**One limitation, stated so it is not found as a bug.** The local-preference above is exact for a
**node-scope** provider co-located with its consumer, and for any node-answer in a mesh that holds no
`model-access` licence. It is *not* yet exact for a **mesh-scope** model server — one node serving the
model to others — **while a licence for `model-access` also exists in the same mesh**: the record pass
that turns a licence into an answer keys on same-node satisfaction and would still demand the licence be
used, double-answering. Until that pass is taught to stand down when a brokered node need already
answers, a mesh-scope local model and a vendor licence must not both answer `model-access` in one mesh.
A node-scope local model has no such constraint. This is named because an unstated limitation is
indistinguishable from a bug, and costs more.
This is a decision and not a patch because it settles **what may answer model access** — a question
every model-access consumer's meaning depends on — and because it lets the mesh's own hosted models sit
behind the same provision as the vendors', which is what makes "the mesh can run its own model" a
deployment choice rather than a second interface to build against.
### How each claim is checked
- **A node answers model access without a licence, and is preferred over one.** A resolver test
assigns a consumer and a module that `provides: ["model-access"]` at node scope on the one node, with
a licence also present, and asserts no resolved need is answered by the record — the local model
answers and the licence is ignored. (This test exists; the decision adopts what it proves.)
- **The consumer is delivered an endpoint, not a credential.** A mesh bed assigns a model-server
provider and a consumer that binds `model-access` and does not list it under `secrets`, and asserts
the consumer's config carries `OPENAI_BASE_URL` built from the provider's served `at`/`port`, and
that no key file was delivered to its secret path.
- **The local endpoint is reachable through what the consumer was given.** The bed makes a request to
the base URL the consumer wrote and asserts the model server answers — the wiring, not a model's
output, is what is proven (the server may be a stub; a real model is not needed to prove the mesh
routed the consumer to it).
- **A node-answer consults no adapter.** A node-answered `model-access` need is resolved with the
vendor registry never read — asserted by the absence of any vendor on a node-answered need and the
ordinary served-facts delivery.
- **The scope limitation holds where stated.** The node-scope case is what the bed and the resolver
test exercise; the mesh-scope-plus-licence collision is recorded here and left for the resolver
change that reconciles the two answer passes, not worked around in a module.
## Consequences
- **The mesh can run its own model, and a consumer reaches it through the same `model-access` it uses
for a vendor.** One interface, two answers; a consumer moves between a vendor and a local model by an
operator reassigning its provision, not by a code change.
- **A local model is keyless by default and keyed by the ordinary path when it must be.** Nothing new
is invented for the local server's credential: it either has none, or mints one the way every
provider does.
- **The vendor path is untouched.** Adapters, the refresh carve-out, and usage
([ADR 0054](0054-model-usage-is-recorded-at-two-grains.md)) are the record answer's business; a
node-answer neither uses nor changes them. A local model that exposes usage would serve it as facts,
not as an adapter's usage verb.
- **The two answer passes meet in one place, and must be reconciled there.** The mesh-scope limitation
is the single point where a node-answer and a record-answer to the one provision can collide; it is
named, and its fix is a resolver change, not a per-module workaround.
## References
- [ADR 0024](0024-model-access-is-a-provision.md) — model access is a provision; this decides a node
may answer it, not only a record
- [ADR 0050](0050-model-access-is-vendor-agnostic.md) — model access is vendor-agnostic; the record
answer and its adapters, which the node answer sits beside and does not use
- [ADR 0054](0054-model-usage-is-recorded-at-two-grains.md) — model usage; a vendor's business on the
record path, served as facts (if at all) on the node path
- [ADR 0008](0008-a-context-owns-its-store.md) — a context owns its store; a licence is a record in one,
a node-answer needs none
- [03-DESIGN/01-to-be/14-model-access.md](../03-DESIGN/01-to-be/14-model-access.md) — the model-access
design, which this extends with the node answer
@@ -0,0 +1,115 @@
---
topic: the tiers
status: accepted
date: 2026-09-09
deciders: jochen
reconstructed: false
extends: 0007-connectivity.md
---
# 66. Public routing is name-agnostic, its names are resolved inside the mesh, and an internal authority can certify them
## Context
**[ADR 0007](0007-connectivity.md) and [connectivity §3](../03-DESIGN/01-to-be/08-connectivity.md)
made a public route a grant: a workload that must be reachable requires a route, the proxy provides
it, the consumer contributes the name it wants and the port it listens on.** What was never pinned
is **what that name is** — and building a whole mesh in the lab showed the gap costs more than it
looks.
**The catalogue shipped each route as a full domain.** A module that needed a public name carried
that name, in full, as a literal in its manifest. Running the same catalogue against a different
domain — a lab standing in for production, or a second operator's mesh — meant overriding that
literal on every routed module, per node. The mesh was, in effect, carrying a **map of names to
services**: the one thing it should never hold, because a name is the operator's choice (one runs
the forge at `git`, another at `code`) and the domain is the node's, and neither is the mesh's to
know.
**And a second gap surfaced the moment an internal issuer tried to certify those names.**
[Connectivity §5](../03-DESIGN/01-to-be/08-connectivity.md) already states the issuer must be
configurable and that the lab runs its own ACME authority. With that authority wired to the proxy,
issuance still could not complete: the authority accepted the order and offered a challenge, then
**could not connect to the validation target.** Nothing inside the mesh resolved the public route
name. The mesh publishes each `<node>.internal` name into every container, but not the public names
the proxy serves — so a validator living in the mesh had no address to reach, and a name the mesh
cannot resolve is a name it cannot have certified.
**The two are one problem.** A name the mesh can *compose* from parts it is given, and *propagate*
to whoever needs to resolve it, is exactly a name it can also have *certified* — and the reverse:
without the composition and the propagation, neither the routing nor the certificate is the
operator's to move between meshes.
## Considered Options
**1. Keep the full domain in the manifest, override per node.** The status quo. It works, and it is
wrong in the specific way this repository cares about: the catalogue holds a domain map, lab and
production differ by an override on every routed module rather than one fact, and a module manifest
names something — the public domain — that belongs to the node, not the module. An unowned name in
the wrong place is the shape of a leak.
**2. A module declares a label; the node declares its public domain; the mesh composes.** The route
contribution carries a subdomain the operator chose, the node carries its public domain as
node-level configuration, and the mesh joins `<label>.<public-domain>` and grants exactly that. The
mesh interprets nothing. Lab-versus-production becomes one node setting. Chosen.
**3. For certification, issue only from a publicly reachable node against a public authority.**
This is already true for public meshes and stays true. It is not an option for a lab or an
internal-only mesh: there is no public authority to answer, and no public reachability to validate
against. An internal issuer is required there — and an internal issuer must be able to *validate*,
which it cannot do unless the routed name resolves and is reachable **inside** the mesh. So the
resolution gap is not optional to close; it is what makes an internal authority possible at all.
## Decision
**The mesh core holds no map of hostnames, subdomains or domains.** A route contribution carries a
**label** (the subdomain) chosen by the module's operator. A node contributes its **public domain**
as node-level configuration. The mesh composes `<label>.<public-domain>`, grants exactly that name,
and never interprets what it means. One operator's forge at `git.example.tld` and another's at
`code.other.example` are the same module with two facts supplied around it.
**When the proxy is granted a name, the mesh publishes that name → the node that serves it into
internal resolution, mesh-wide** — the same mechanism, and the same "given by the mesh, not chosen
by a module," that already writes `<node>.internal` into every declared container
([connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md)). The mesh propagates the names it was
told to serve. It still knows nothing about what any of them mean.
**An internal authority certifies those names by the same path a public one would.** The proxy is
pointed at whichever issuer the mesh names — a public ACME authority, or an internal one — and
trusts that issuer's root; nothing else about issuance changes. The internal authority validates by
reaching the routed name, which the clause above has just made resolvable inside the mesh. So the
three are one decision: **compose the name, propagate it, certify it** — each is meaningless without
the one before it.
## Consequences
- **Lab-versus-production is one node-level `public-domain` setting**, not an override on every
routed module. The same catalogue runs against any domain.
- **The manifest layer needs composition it does not yet have.** Today a route name is stored as a
literal, with no interpolation of a node's domain into a module's label. Until that exists, the
composed name is produced by a per-node settings override — a stopgap that reproduces option 1's
per-module cost and is explicitly *not* the design.
- **An internal issuer depends on route-name resolution.** Its challenge validates against the
routed name; without that name in internal resolution, issuance for it cannot complete inside the
mesh. The lab found this as a live failure, not a theory.
- **Nothing about the public path changes.** A publicly reachable node issuing a public name from a
public authority is untouched; this widens the same shape to names and meshes that are not public.
**How each is checked** — an unenforced rule is indistinguishable from a wrong one:
- **Name-agnostic:** the same catalogue resolves against two different public domains by changing
one node setting and nothing else; and no module manifest contains a full public domain. A
manifest that pins an FQDN is the smell the check looks for.
- **Resolution:** a request to a routed name, made from inside the mesh, reaches the workload that
serves it — and, the sharper check, issuance for that name against the internal authority
completes, which it cannot unless the validator resolved and reached the target.
- **Internal authority:** a TLS handshake to a routed name verifies against the internal root and
nothing else, the same shape §5 already uses for internal node-to-node names.
## References
- [ADR 0007 — connectivity](0007-connectivity.md), which made a route a grant and named exposure,
resolution and certificates as one context.
- [ADR 0009 — modules and the graph](0009-modules-and-the-graph.md), the provide/require/contribute
vocabulary a route and an authority both use.
- [Connectivity design §2, §3, §5](../03-DESIGN/01-to-be/08-connectivity.md), amended alongside this
record.
+129
View File
@@ -0,0 +1,129 @@
---
topic: the tiers
status: accepted
date: 2026-09-10
deciders: jochen
reconstructed: false
extends: 0006-the-substrate-and-the-control-plane.md
---
# 67. Genesis is a pivot: a temporary control plane installs the registry that makes it permanent
## Context
**The mesh builds its own modules into its own registry, and there is no public registry for them
— that is the point, not an omission.** A module is cloned from the forge, built, and published
where the mesh can move, replace and back it up. Nothing about that arrangement wants a copy of the
mesh's code hosted by somebody else.
**Every image must be pinned by digest** ([ADR 0006](0006-the-substrate-and-the-control-plane.md)),
and the reasoning is exactly right: a bundle is applied where no mesh exists to check anything
against anything, so what it names must be exact.
Those two sentences are individually correct and together they close a door. The digest a pin means
is a *manifest* digest, and a manifest digest is **assigned by a registry when something is pushed
to it**. The control plane's image is built from source and pushed nowhere, so it has no such
digest, so it cannot be named — and a registry cannot be installed without a control plane to
install it. That is not a pin. It is a dependency the pinning rule created by accident.
**It went unnoticed because the lab hid it.** The lab raised a disposable registry, stocked it from
a workstation, and rewrote every image reference to point at it — so the lab bootstrapped along a
path no real machine has. A first node in the lab always worked, and a first node anywhere else had
no path at all. Every bootstrap fault found this year was found late for the same reason: **the
install procedure existed only as a test fixture**, and a fixture is free to invent what it needs.
## Considered Options
**1. Publish the mesh's own images to a public registry.** Rejected on the premise: there is no
public registry for the mesh's modules and there is not meant to be. It would also make raising a
mesh depend on somebody continuing to host its code, which is the dependency the whole arrangement
exists to remove.
**2. Build the control plane from source on the first machine.** Rejected. A bare machine would
need a toolchain and a working tree — and worse, the source lives in a forge **that runs on the
mesh**. A total rebuild would then need the mesh it is rebuilding. Acceptable for adding a node to
a healthy mesh; useless for the case that matters.
**3. Keep a disposable registry as an install step.** Rejected. It exists in no production, and
concealing this problem is precisely what it has been doing.
**4. Pivot through a temporary control plane.** Chosen.
## Decision
**An image may be named by the digest of its own configuration.** A bare `sha256:…` names an image
the machine already holds — content-addressed, immutable, unforgeable, and requiring nothing to
have served it. It satisfies what the pinning rule asks for; the rule simply never contemplated an
image that no registry had ever seen. It is legal exactly where nothing could have served one.
**Genesis is a pivot**, in this order:
1. the installer **carries the control-plane image** and loads it onto the machine
2. a **temporary** control plane is raised from it, named by that image's own digest
3. the **registry module is installed** — its image is upstream and it is never built, which
[`04-ISSUES/029`](../04-ISSUES/029-the-artifact-store-cannot-be-delivered-by-the-artifact-store/00-report.md)
already settled: a module that provides the artifact store cannot be delivered through it
4. the control-plane image is **pushed into the mesh's own registry**, which assigns it a manifest
digest — the first one it has ever had
5. the control plane is **reinstalled as an ordinary module** pinned to that digest
**The host performs the replacement, not the control plane.** Tier 0 outlives tier 2: the control
plane composes a declaration naming the registry-pinned image, and the host applies it and recreates
the container. Nothing is asked to replace itself while running, and the control plane is stateless
— what it knows is in the store.
**The installer is tier 0, and a separate program from the host.** Bootstrapping is by hand and
changes the machine, which is tier 0's definition. But the host states that it *connects to nothing
and listens on nothing*, and that claim is what makes the one thing running forever on every
machine auditable. An installer connects to plenty. Same tier, same delivery, different program.
## Consequences
- **The control plane stops being a special case.** It becomes an ordinary module with an ordinary
image in the mesh's own registry — so the mesh can build and roll out **its own upgrades**, which
is what a mesh that runs itself was always reaching for.
- **The bundle's job shrinks** to raising a temporary control plane exactly once.
- **The source builds the installer; it does not run it.** Cloning moves to a release machine, where
a forge being available is an ordinary working assumption, and leaves the disaster-recovery path
where it very much is not.
- **The lab's disposable registry is deleted.** The lab bootstraps by running the same program a
bare machine runs — the only arrangement in which the installer cannot quietly drift out of truth
again.
- **Two public images remain at genesis** — the store and the broker. An air-gapped install would
embed those too, at a much larger artifact; that is a build variant, not a different design.
- **A control-plane module manifest must exist**, and did not.
- **The registry may require nothing.**
[`04-ISSUES/029`](../04-ISSUES/029-the-artifact-store-cannot-be-delivered-by-the-artifact-store/00-report.md)
states that a module providing the artifact store may not *build* artifacts, because there is
nowhere to put them until it runs. The pivot shows that is the narrow case of a wider rule: **it
may not require anything the store is needed to deliver.** Found the hard way — an unrelated
change gave the registry a public name and, with it, a route requirement. At genesis nothing
provides a route, and nothing can, because the routing stack needs images and images need the
store. The same cycle, re-entered through a door the existing wording did not cover.
- **The handover is the sharp edge.** For one moment the bundle and the module both describe the
same container, and the host tracks what it owns. If a safe handover is not expressible with what
exists, the install **stops before it** and says what is missing. A machine left without a control
plane cannot be fixed remotely, so a partial install that halts cleanly is the better outcome.
**How each is checked** — an unenforced rule is indistinguishable from a wrong one:
- **Naming by its own digest:** a machine that can reach no registry at all raises a control plane.
- **The pivot completed:** after installing, the running control plane's image is pinned by a digest
**the mesh's own registry assigned** — not by an image id. If it is still the image id, the pivot
did not happen and the mesh cannot upgrade itself.
- **No fiction left in the lab:** the scenario declares no registry machine, and the bed bootstraps
through the installer rather than around it.
- **The handover:** a machine whose control plane has been replaced still has one, and it answers.
- **The registry requires nothing:** its manifest is resolvable on a mesh that has no other module
in it. A requirement added to it later is caught where it is written, rather than by a genesis
that cannot complete — which is how this one was found.
## References
- [ADR 0006 — the substrate and the control plane](0006-the-substrate-and-the-control-plane.md),
which pinned images by digest and named what a first node fetches.
- [ADR 0005 — the node host](0005-the-node-host.md), which makes tier 0 the one thing installed by
hand and the only thing that changes a machine — the property this keeps true by shipping the
installer beside the host rather than inside it.
- [`04-ISSUES/029`](../04-ISSUES/029-the-artifact-store-cannot-be-delivered-by-the-artifact-store/00-report.md),
the same cycle one layer down, and the rule that a registry module is named and never built.
+112
View File
@@ -0,0 +1,112 @@
---
topic: building it
status: proposed
date: 2026-09-12
deciders: jochen
reconstructed: false
extends: 0016-the-lab.md
---
# 68. The lab takes requests, one at a time, and runs each from its own copy
## Context
**The lab is exclusive hardware, and today a person holds it.** Raising a scenario takes over
addresses and names on the workstation for as long as it stands, and only one scenario can stand
at a time. So a run is not merely slow — it occupies the machine and the person who started it,
who then waits rather than works.
**Running it in the background against the working copy is worse than waiting.** The obvious fix
is to start a run and carry on editing. But a run reads the working copy as it goes: binaries are
rebuilt from it, manifests are read out of it, and the bed's own code is loaded from it. Edit
while it runs and the result describes a state that never existed — a mixture of what was there
when each file happened to be read. A green result obtained that way is not evidence, and a red
one costs a day to disbelieve.
**Nothing today records what was asked for.** A run is a command line in somebody's terminal. What
commit it exercised, what it was trying to find out, and what it answered all live in scrollback,
which is why the same question gets re-run rather than looked up.
**Most of the parts already exist.** The lab writes a receipt of its last run. The mesh already
carries messages between nodes and can notify a person. The machine already runs work on a
schedule. What is missing is the thing in the middle.
## Considered Options
**1. Leave it as it is — a person drives the lab and waits.** Rejected. It is the loop
[ADR 0010](0010-delivery.md) removed everywhere else, kept here by habit rather than by argument,
and the cost compounds: because a run is expensive to start and blocks the person, fewer are run,
so faults are found later and in larger batches.
**2. Run in the background against the working copy.** Rejected on the reasoning above. The
failure is silent, which is the kind this repository exists to refuse.
**3. Put the lab behind the ordinary build pipeline.** Rejected for now. The pipeline builds
artifacts and does not own a machine that can raise virtual machines; giving it one makes the
pipeline's slowest job the lab's, and couples every push to hardware only one machine has. This
may become right later; it is not the smallest thing that works.
**4. A queue in front of the lab, and an isolated copy behind it.** Chosen.
## Decision
**The lab accepts requests rather than commands.** A request is recorded, queued, and answered.
The person who made it is told when it is answered and does not wait.
**A request names a bed and a commit, and nothing else.** This is the load-bearing restriction. A
request may say *run this bed, at this version of these repositories*. It may not say what to
install, on which machine, or with which settings — because a request that could say those things
would be a second way of installing a mesh, and the whole reason the installer exists is that the
lab already was one ([ADR 0067](0067-genesis-is-a-pivot.md)). The bed decides what is installed;
the request only decides which bed and which version.
**Requests are released one at a time.** The hardware admits one standing scenario, so the queue
enforces what the hardware already requires, rather than leaving it to whoever remembers.
**Every run happens in a copy the lab owns.** The lab checks the requested commit out into its own
path and builds and runs from there. A working copy is never read by a run. This is what makes the
queue safe to use while work continues, and without it the rest of this record is not worth
having.
**The lab is reached through tools, not only a command line.** A command line is available only
to whoever is sitting at the machine, which is the constraint this record exists to remove. The
lab answers three questions to anything that can reach the mesh — *what is standing now*, *what is
queued or running*, and *what did this request answer* — and accepts a request and a cancellation.
An agent can therefore start a run, stop attending to it, and come back; and somebody who did not
start a run can still see it, which is the difference between a shared lab and a private one.
**The restriction holds at every door.** A tool submits a bed and a commit, exactly as a command
line does. A tool that could name a module, a node or a setting would reintroduce the second
installer through a different entrance, and the entrance is not what made it dangerous.
**Every run leaves a record that outlives the terminal**: what was asked, which commit, when it
ran, what it answered, and where its output went. A question already answered is looked up rather
than re-run.
## Consequences
Work continues while the lab runs, which is the point. A second session may edit freely, because
nothing it edits is what the lab is reading.
A request is reproducible by construction: it names a commit, so the same request can be asked
again and compared. Today two runs of "the same thing" are only as alike as the tree happened to be.
The lab gains a second copy of every repository it exercises, costing disk and needing to be kept
from drifting into a place people edit by hand.
Anything that can reach the mesh can now see what the lab is doing, including an agent working on
something else. That is the intended gain and also the obvious hazard: a thing that is easy to ask
is easy to ask too often, and the hardware still admits one scenario at a time.
The queue becomes a thing that can fail — stuck, backed up, or lost — and a queue nobody watches
is worse than no queue, because it absorbs requests silently.
## How this is checked
| Rule | Checked by |
|---|---|
| A run never reads a working copy | The runner is given a path it owns and no other; a run started while a working copy is deliberately dirtied produces a result matching the commit, not the edits. |
| One scenario stands at a time | A second request submitted while one runs is observed to wait, not to raise. |
| A request cannot say what to install | The request format admits a bed and a commit only. A request naming a module, a node or a setting is refused, and the refusal is exercised. |
| A request is answered | Every queued request reaches a terminal state with a record. A request that vanishes is a failure of the queue, not a quiet nothing. |
| The lab can be asked from elsewhere | What is standing is asked from a session that did not raise it, and the answer matches the machine. A lab that only answers its own caller has not left the terminal. |
@@ -0,0 +1,89 @@
---
topic: building it
status: accepted
date: 2026-09-12
deciders: jochen
reconstructed: false
extends: 0009-modules-and-the-graph.md
---
# 69. A module is a repository and a path within it
## Context
**The builder clones one repository and reads `module.json` at its root.** The to-be design says
so in as many words — *"one file at the root"* — and the code implements it: clone, read the root
manifest, build what it declares.
**Nothing that exists is shaped that way.** The catalogue holds sixty-seven modules, each in its
own directory, and has no manifest at its root. None of the five code repositories has one either.
So today the builder cannot be asked to build any module that exists: pointed at the catalogue it
finds no manifest, and pointed at a module's source it finds no manifest.
**The system being replaced already works the other way**, and has for years: a monorepo with one
directory per piece of software, and the coordinator builds a module from a repository and a path
inside it. The root-only assumption is not a simplification of that — it is a different model that
was never reconciled with it.
**And it splits what a build needs into two places.** The control plane's manifest sits in the
catalogue; the source it describes sits in the control plane's own repository. A build must read
one tree, so under the root-only model neither location can be built from.
## Considered Options
**1. One repository per module.** Rejected. Sixty-seven repositories for sixty-seven modules, most
of which are a single manifest naming a public image, and every one needing its own creation,
permissions and lifecycle. It also contradicts [ADR 0015](0015-applications-live-in-their-own-repository.md),
which put *applications* in their own repositories precisely because modules do not need one.
**2. Keep manifests in the catalogue and source elsewhere, and have a build fetch both.** Rejected.
A build would clone two trees whose versions can disagree, so "what commit is this module?" stops
having one answer — and that question is the whole basis of knowing when to rebuild.
**3. A module is a repository and a path within it.** Chosen. It is what the current system does,
what the catalogue already looks like, and it keeps a module's description beside the thing it
describes.
## Decision
**A module is named by a repository and a path within it.** The path holds `module.json`, and
everything that manifest declares is produced from that path. A module whose path is the root is
the ordinary case of this, not a separate one.
**A module's manifest lives beside its source.** Where a module has code, its directory holds both,
so one commit answers "what is this module, and what is it made of". Where a module has no source —
a manifest naming a public image — the directory holds only the manifest, and there is nothing to
build.
**This moves the core modules.** The control plane and the builder are built from the control
plane's repository, so their manifests belong in that repository at their own paths, not in the
catalogue. The catalogue keeps the modules whose source it holds, and the modules that are only a
manifest.
**One commit, one module version.** Because a module is one path in one repository, the commit that
built it identifies it exactly, and "the source has moved ahead of what the mesh holds" stays a
question with a yes or no answer.
## Consequences
The builder gains a path alongside the repository and the ref. A build is `repository, path, ref`,
and the manifest it returns is the module the mesh records.
The catalogue stops being the place every manifest lives, and becomes the place manifests live
*when their module has no other home*. That is a smaller claim than it sounds: most of the
sixty-seven stay exactly where they are.
Two repositories change shape — the control plane's gains manifests for the modules built from it.
Nothing else moves.
A repository can hold modules that are built and modules that are not, and no rule distinguishes
them beyond whether their manifest declares anything to build.
## How this is checked
| Rule | Checked by |
|---|---|
| A module is buildable from its repository and path | The builder is asked for a module by repository and path, and returns a manifest whose artifacts are pinned to digests the mesh's registry assigned. |
| A manifest sits beside what it describes | A module declaring something to build, whose path holds no source to build it from, is refused at build time rather than producing an empty result. |
| One commit identifies one module | Two builds of the same repository, path and commit produce the same digests. |
| The core modules are built like any other | The control plane is rebuilt from its own repository and path, and the running mesh is upgraded to it — the same path an ordinary module takes. |
@@ -0,0 +1,107 @@
---
topic: the tiers
status: accepted
date: 2026-09-12
deciders: jochen
reconstructed: false
extends: 0067-genesis-is-a-pivot.md
---
# 70. The catalogue owns the module graph, and genesis builds rather than carries
## Context
**The module graph has no owner.** What modules exist, what each requires and provides, what each
claims, what each is made of — all of it lives inside the control plane because that is where it
was first written, not because anything decided it belonged there.
**The control plane's own test says it does not belong there.**
[`06-the-control-plane`](../03-DESIGN/01-to-be/06-the-controller.md) defines the tier as
*everything that needs to know about more than one node*, and states the corollary plainly:
anything a single machine could answer alone is not the control plane's. What a module is, and
what it needs, requires no knowledge of any node whatsoever.
**And nothing can query it.** The graph is the thing that answers *what must be rebuilt when this
changes*, *what would break if this were removed*, and *what can be installed here* — and today it
is reachable only as control-plane internals. [ADR 0009](0009-modules-and-the-graph.md) already
decided that a build edge is derived, that an artifact is stale when anything it was built against
moved, and that the rebuild set is therefore computable. None of that has anywhere to live.
**Genesis currently carries an image, and cannot produce a builder at all.**
[ADR 0067](0067-genesis-is-a-pivot.md) has the installer carry the control plane's image. That
works, but it leaves the builder with no route onto a fresh mesh — it cannot be fetched from the
public internet, because the mesh builds it, and the installer carries one image only. So a raised
mesh cannot build anything, including the modules it is made of.
## Decision
**The catalogue is a core module, beside the control plane and the builder, and it owns the module
graph.** Modules, what they require and provide, what they claim, their dependencies and their
build edges, assignments and configuration — the catalogue holds them and serves tools over them:
install a module, query what is available and what it needs, query the graph, query and update
settings.
**It is one per mesh**, expressed the way the control plane already expresses it — a claim scoped
to the mesh, not a new mechanism.
**The control plane consumes it.** Resolving what a module requires into an actual binding, and
composing what a machine should be, both need to know what modules are. So the dependency runs
from the control plane to the catalogue, which is the opposite of what the tiers suggest and is
therefore written down here rather than left to be inferred.
**Genesis builds the core modules rather than carrying them.** The installer ships an *init
builder* — the one thing carried — which is started, clones the source, and builds the control
plane, the catalogue and the builder. The installer then raises a temporary control plane, which
installs the catalogue, registers the permanent control plane and the builder, and assigns all
three to the first machine. The temporary control plane stops; the installer verifies that the
permanent one answers and can query the graph. The builder then sees a catalogue in its initial
state and builds the core modules into it.
**So exactly one thing is carried, and it is a builder rather than a result.** That is the
difference from [ADR 0067](0067-genesis-is-a-pivot.md), which carried the control plane's image:
carrying a builder produces every core module on the machine, including the builder itself, so
there is no component left without a route.
## Consequences
The catalogue joins the small set of things that cannot arrive through the ordinary path, because
it cannot be installed by something that needs it in order to install anything. It arrives the same
way everything else does under this record — built by the init builder before the mesh can install
anything — so the set is answered by one mechanism rather than three special cases.
The rebuild fan-out gains a home. *What was this built against* and *what must rebuild now* are
questions about the graph, and the graph now has an owner to hold the edges and answer them.
The catalogue holds state, so it owns a store in the substrate's database, the same way the control
plane's contexts do. That is the mesh's own store and not the `postgres` module, which is a
provider other modules consume.
A mesh without a catalogue cannot resolve anything, where previously it merely lacked an interface.
That is the cost of ownership over surfacing, and it is deliberate: one owner beats two copies.
## Open, and to be settled before this is built
**Where the init builder clones from.** [ADR 0067](0067-genesis-is-a-pivot.md) rejected building
from source at genesis partly because the source lives in a forge that runs on the mesh, so a
total rebuild would need the mesh it is rebuilding. Carrying a builder answers the toolchain half
of that objection and not this half. Genesis must therefore name a source that exists before the
mesh does.
**What the init builder publishes into.** Building produces artifacts that must be pinned by a
digest a registry assigned, and today the registry is installed after the machine has joined. If
the core modules are built first, the registry has to exist first, so the substrate's order needs
restating rather than assumed.
**Where the line falls between the catalogue and the control plane.** Claims and assignments need
to know about every node, which by the control plane's own test is its work. Whether the catalogue
holds them and asks, or the control plane holds them and the catalogue surfaces them, is not
settled here.
## How this is checked
| Rule | Checked by |
|---|---|
| The catalogue owns the graph | The control plane answers *what does this module require* by asking the catalogue, and a mesh whose catalogue is stopped cannot resolve — observed, not assumed. |
| One per mesh | A second catalogue assigned anywhere in the mesh is refused by the claim, and the refusal is exercised. |
| Genesis builds rather than carries | The installer carries exactly one artifact, and after installing, every core module is pinned to a digest the mesh's own registry assigned. |
| The builder has a route | A mesh raised by the installer, with no hand-placed image, can build a module. |
@@ -0,0 +1,80 @@
---
topic: the tiers
status: accepted
date: 2026-09-12
deciders: jochen
reconstructed: false
extends: 0070-the-catalogue-owns-the-module-graph.md
---
# 71. Genesis clones from a mesh, and checks what it got
## Context
**[ADR 0070](0070-the-catalogue-owns-the-module-graph.md) has the init builder clone the source,
and does not say from where.** [ADR 0067](0067-genesis-is-a-pivot.md) had already rejected building
at genesis partly for that reason: the forge holding the source runs *on* the mesh, so a total
rebuild would need the mesh it is rebuilding.
That objection is real but narrower than it reads. It only binds when the mesh being raised and the
mesh holding the source are the same one, which is true exactly once.
## Decision
**Genesis clones from a mesh's forge, reached by name.** Any mesh that holds the source can serve
it. The first mesh is not structurally special — it is simply the only one that existed when there
was nothing else to clone from.
**If the mesh serving the source is lost, the name moves to another mesh that holds a copy.**
Recovery is a name pointing somewhere else, not a backup being restored. This is what makes the
source's survival a property of there being more than one mesh, rather than a property of somebody
having remembered to take a copy. A mesh that has installed from that name holds the source
afterwards, so every installation adds a place the name could point.
**Genesis names a commit and checks what it got.** It does not clone whatever a branch happens to
point at. The forge a mesh installs from is the trust anchor for everything that mesh will ever
run, and a branch is a moving target that somebody else controls.
*This is not hypothetical.* On 2026-09-11 the forge that would serve this role was running a
cryptominer, and its git operations were being tampered with in flight — output injected into the
protocol stream by a hook that fired on every fetch. Nothing was altered: the repositories were
verified against local copies and found byte-identical. But a mesh installing from that name during
those hours had no way to establish that for itself, and would have had none.
## Consequences
The init builder needs a name it can resolve and a commit it can verify, and nothing else. It does
not need to know which mesh answers.
Whoever operates the mesh that name points at carries a responsibility to everyone installing from
it, and should know that. It is not merely a convenience host.
A mesh that cannot reach any forge cannot be raised. That is a real limit and it is accepted: the
alternative is carrying the whole source in the installer, which makes the installer a release
artifact that goes stale rather than a program that fetches what it was told to.
## Open — what relationship a mesh keeps afterwards
**Not decided, and named here so it is not decided by accident** by whoever writes the init
builder. Two shapes, and they are meaningfully different:
**A snapshot, and then independence.** A mesh installs once, mirrors the source into its own forge,
and has no upstream afterwards. It is fully self-hosted, in the sense that nothing it needs lives
anywhere else. Updates are then something an operator does deliberately, by pulling changes in —
tooling for which is possible and is not a priority.
**A continuing upstream for core modules**, the way a distribution serves packages and a separate
collection serves everything else. A mesh keeps looking at the origin for the modules that make a
mesh a mesh, and holds its own for the rest.
The first is more obviously aligned with the rest of this design, which is arranged so nothing a
mesh needs depends on somebody else continuing to host it. The second is more convenient and makes
a security problem in one forge everybody's problem. Neither is chosen here.
## How this is checked
| Rule | Checked by |
|---|---|
| Genesis needs only a name and a commit | A mesh is raised with the name pointed at a different mesh than the last time, and the result is identical. |
| What was cloned is what was asked for | Genesis is pointed at a commit and refuses a forge serving different content under it, rather than building what it received. |
| Losing the serving mesh is survivable | The name is repointed at a mesh that installed from it earlier, and a raise succeeds. |
@@ -0,0 +1,86 @@
---
topic: the tiers
status: accepted
date: 2026-09-12
deciders: jochen
reconstructed: false
extends: 0070-the-catalogue-owns-the-module-graph.md
---
# 72. Two graphs, and a build chain that orders itself
## Context
[ADR 0070](0070-the-catalogue-owns-the-module-graph.md) gave the catalogue the module graph and
said the control plane consumes it — that resolving a requirement and composing what a machine
should be *"both need to know what modules are, so the dependency runs from the control plane to
the catalogue."*
**That paragraph is wrong, and this record corrects it.** It was written before the two graphs had
been told apart, and it creates a dependency that does not need to exist: a catalogue that is down
would leave the control plane unable to compose the declaration that would repair it.
## Decision
**There are two graphs, with different owners, and they meet only when something is installed.**
| Graph | Owner | What it links | Answers |
|---|---|---|---|
| the module graph | the **catalogue** | module-versions to each other | what was this built against · what must rebuild now · what does this need |
| the runtime graph | the **control plane** | module-versions to nodes | what runs where · who consumes this provision · what breaks if this machine goes |
The catalogue does not know nodes exist. The control plane holds module-versions and nodes, along
with capabilities and claims, because deciding whether a machine qualifies needs every node.
**So the control plane never asks the catalogue anything.** It holds what it needs to compose a
declaration. A catalogue that is down stops new installs and stops the rebuild fan-out, and does not
touch anything already running or the ability to repair it.
**Everything between them travels as events over the broker**, like all module-to-module
communication. The builder finishes and announces that a module was built. The catalogue registers
it, places it among what it depends on, and announces that a module was upgraded. The control plane
reacts to *that*, not to build output — a semantic fact rather than an artifact.
**Nothing is lost if a receiver is down.** A consuming module's queue is durable with a dead-letter
exchange, declared by the mesh rather than by the module, so an event waits for a consumer that is
not there.
**Build order is not computed. It emerges from the chain.** The builder never consults the graph: it
builds what it is asked for, one at a time. The catalogue asks for the next build after the previous
registration, so *"do not start this until that is registered"* holds by construction rather than by
a schedule somebody maintains.
**The catalogue's rule is a condition, not an order:** ask for a module to be rebuilt once everything
it was built against is current. That handles a chain and a diamond with one rule, where an
order-based approach needs to know the shape in advance.
**Two refusals belong to the catalogue.** A cycle, because the chain would never settle. And a
rebuild whose artifacts are identical to what it replaced, which is not an upgrade and must not be
announced as one — or a single change ripples outward forever through modules that did not change.
**Genesis does none of this.** The init builder has a fixed, short list — control plane, catalogue,
builder — in a written order, because there is no catalogue yet to ask.
## Consequences
The builder stays simple, and independent of the catalogue. That is what makes genesis possible at
all: the thing that builds the catalogue cannot require the catalogue.
The two sides can be briefly out of step — the catalogue may hold a module-version a moment before
the control plane knows of it. Assigning in that instant fails, and should say why rather than
report that no such module exists.
A module's declared events stop being documentation and become its permissions: an account is
scoped from what a module emits and consumes, so a builder that announces what it built is granted
what it needs by the ordinary mechanism rather than by a special case.
## Open
**Whether an upgrade is applied or merely noticed.** Today the system this replaces deploys
automatically, and that is a defensible default for core modules on a mesh its operator runs. But
the design as it stands does the opposite: it records that the source moved ahead, makes it visible,
and waits to be told. This must become a setting with a chosen default rather than inherited
behaviour — and the choice matters most on the day a bad commit reaches something that carries mail.
**Whether a module assigned to several machines upgrades on all of them at once.** Doing so makes
one bad commit simultaneous everywhere. Doing one machine and pausing turns it into one casualty.
@@ -0,0 +1,94 @@
---
topic: the tiers
status: accepted
date: 2026-09-13
deciders: jochen
reconstructed: false
extends: 0070-the-catalogue-owns-the-module-graph.md
---
# 73. The installer carries a builder, and the registry stays where it is
## Context
[ADR 0070](0070-the-catalogue-owns-the-module-graph.md) decided that genesis builds rather than
carries, and [ADR 0071](0071-where-genesis-gets-its-source.md) settled where it clones from. Two
questions were left open, and the design record names them as the one gap that stops a fresh mesh
from being able to produce anything at all: **how the builder arrives**, and **what it publishes
into**.
Today the installer carries the control plane's image inside itself. That works, and it is why
genesis needs no registry: nothing is ever fetched, because the one image that matters is already
present. The cost is that the mesh which results holds an artifact it did not make, cannot rebuild,
and knows nothing about — no version, no source, no edges. That is the same shape as the fault
[issue 044](../04-ISSUES/044-the-runtime-every-module-builds-on-cannot-be-built-by-the-mesh/00-report.md)
recorded for the shared runtime, and fixing it there while shipping it here on every new mesh would
be a strange place to stop.
## Decision
**The installer carries a builder, and nothing else.** One artifact, not a growing set. It clones
the source at a named commit, checks what it got ([ADR 0071](0071-where-genesis-gets-its-source.md)),
and produces the control plane from the same repository and path that any later rebuild of it would
use. What raises the mesh is therefore the same thing that will maintain it, and there is no second
mechanism kept in step with the first.
**The registry does not move, and the argument for moving it does not survive being made.**
It was put this way: a produced image has to be put somewhere before anything can fetch it, so the
registry must now precede the control plane, and
[ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md)'s answer to that question has to flip.
It does not, because the premise is false. **The thing that builds the image and the machine that
runs it are the same machine.** A built image is already in that machine's container runtime, and
the temporary control plane names it exactly as it names a carried one — by the digest of its own
configuration, a local identity that requires nothing to have served it. Building changes where the
bytes came from. It does not change where they are.
| | is it substrate? | must it precede the control plane? |
|---|---|---|
| the store | yes | yes — there is nowhere else to put the control plane's state |
| the broker | yes | yes — the control plane reaches a machine only over it |
| the image registry | yes — it cannot grant itself a repository | **still no** — the first machine neither fetches the control plane nor needs to, whether the image was carried in or made here |
So the registry stays where [ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md) put it:
substrate by role, ordinary by delivery, installed by the temporary control plane as its first act.
The bundle carries two services and a control plane, as it did. **What publishes into the registry
is unchanged too** — the existing step that pushes the control plane's image into it, which is the
moment that image first receives a digest assigned by something other than itself. It now pushes
something this mesh built rather than something it was handed.
## Consequences
**Genesis gains one step and changes no others.** A build happens before the image is loaded. The
pivot described in [ADR 0067](0067-genesis-is-a-pivot.md) survives exactly as written, because the
step it pivots on never cared where the image came from.
**A fresh mesh can produce from the moment it exists.** The builder is present before the control
plane is, so the core modules, the catalogue and the builder's own module can be built in the
ordinary way rather than waiting for somebody to carry them in. The paragraphs in
[`17-raising-a-mesh`](../03-DESIGN/01-to-be/17-raising-a-mesh.md) that describe this were describing
something that could not start; they can start now.
**Genesis needs more of the outside world.** Carrying an image needed nothing but the installer.
Building one needs the source, and whatever the build itself reaches for. This is a real cost and
is not waved away: it makes genesis fail in more ways, all of them at a step that says what it was
doing. It is accepted because the alternative is a mesh that cannot rebuild its own control plane,
which fails in exactly one way, silently, later, and for ever.
**A pre-built bundle remains possible and is not this.** Nothing here forbids delivering artifacts
rather than building them; it fixes where they may come from. A bundle of pre-built core modules is
an **export of a mesh that built them**, carrying what the catalogue knows about each alongside the
artifact itself — so that loading one leaves the graph in the state building would have left it. A
bundle that carries images without that is the thing this decision rejects, whoever ships it.
## What this does not decide
**Whether the builder's own module is carried or built.** It builds everything else; what installs
*it* as an ordinary module afterwards, so that it too can be upgraded, is the same closed-list
question [`12-a-module-repository`](../03-DESIGN/01-to-be/12-a-module-repository.md) already holds,
and is unchanged by this.
**How a machine authenticates to a registry that asks it to.** Genesis raises its own and reaches it
over the loopback, so this remains a joining problem
([issue 042](../04-ISSUES/042-nothing-gives-a-node-an-account-for-a-registry/00-report.md)).
@@ -0,0 +1,172 @@
---
topic: the tiers
status: accepted
date: 2026-09-15
deciders: jochen
reconstructed: false
extends: 0039-what-the-sdk-holds-and-refuses.md
---
# 74. The mesh defines a module protocol; an SDK is an implementation of it
## Context
[ADR 0039](0039-what-the-sdk-holds-and-refuses.md) says what the SDK holds: the tool-serving
harness, the messaging and event framework, the contracts, and core primitives. It settles what
belongs in *an* SDK. It does not say what happens when there is more than one.
There is already more than one. **The contracts are expressed twice** — as Go types in the control
plane and the host, and as TypeScript types in the SDK — and nobody has felt it because both live
in one repository and one head.
**A correction, made after inspecting the wire rather than the types** (2026-09-16). This record
first claimed the two implementations already disagreed — `resource` vs `Provision`, `consumer`
meaning the module in one and the node in the other, headers declared on one side and emitted by
neither. **On inspection the live wire agrees**, and the claim was wrong:
- The grant types that disagreed (`Grant`, `Interface`, `Credential` in the SDK's `contracts`)
were **dead** — exported and imported by nothing. The live provisioning wire is the contributions
file, whose shape (`as`, `secret`, `node`, `at`, `values`) is the same on both sides. Those dead
types have been removed.
- The envelope agrees too: Go emits all five required headers, and `x-causation-id`/`x-schema` are
**optional** — the SDK sets them when a handler has a causation or a schema, and a bare event
carrying neither is correct, not a drift.
So the danger was never live disagreement. It was **dead types that contradicted the live wire**,
which read as the contract and were not — and are exactly what led this record to assert a drift
that inspection did not find. That is a sharper reason for the decision below, not a weaker one: a
type is only as good as its being the wire, and the way to guarantee that is to specify the wire and
check implementations against it, rather than to trust a hand-kept type to still describe it.
A failure of this kind does not announce itself. Two implementations that disagree about an
envelope do not fail to compile — they ignore each other's messages, and a mesh where a module
stops reacting looks exactly like a mesh where nothing happened.
## The question this settles
A module may be written in any language the mesh can build
([`18-building-a-module`](../03-DESIGN/01-to-be/18-building-a-module.md)). Every language needs an
SDK. What is an SDK *of*?
Two answers were available, and the obvious one is wrong.
**Shared types, generated.** Write the shapes once — a schema, an IDL — and generate Go, TypeScript,
Rust. It is the familiar answer and it solves the smaller half of the problem. The shapes are not
where the difficulty is.
**A specified wire, with a conformance suite.** The shapes are a consequence; what an SDK must get
right is *behaviour*.
## Decision
**The mesh defines a module protocol. An SDK is an implementation of that protocol in one
language, and nothing more.**
That is the whole of what an SDK is. Not a library a language happens to have, not a convenience
layer, not a place for helpers to accumulate — an implementation of a specified protocol, finished
when it implements it and correct when it agrees with every other implementation.
### The protocol is split per capability
**A module does not use all of it, so an SDK need not implement all of it.** A module that only
consumes events uses the event capability. One that serves tools uses the tool capability. A
provider uses provisioning. Nothing about consuming an event requires knowing how a grant is
answered.
So the protocol is a floor plus capabilities:
| part | what it covers | who needs it |
|---|---|---|
| **connection** — the floor | reading the sealed credential, pinning the certificate fingerprint, taking identity from the credential rather than the environment | everything |
| **events** | the envelope and its headers, the durable per-consumer queue, binding, at-least-once with dedup on `x-event-id` | a module that emits or consumes |
| **tools** | registration, the shared durable `serve.<key>` queue, request and reply | a module with a surface |
| **provisioning** | a grant in, a credential out, and what each carries | a module that provides something |
**This is the same shape the host already has.** A host declares which resource kinds it can apply,
and a partial host — one that can write files and run things but not manage users or containers —
is a real thing rather than a broken one ([ADR 0005](0005-the-node-host.md)). An SDK that implements
the floor and events is exactly as legitimate, and a module written against it is a module that
does events.
**So a language arrives in pieces rather than all at once.** A Rust SDK implementing connection and
events is useful the day it exists; tools and provisioning follow when something needs them. The
alternative — a language is unsupported until it is entirely supported — is what makes adding one a
project rather than a contribution.
**And what a language can be used for is then a fact the mesh can state**, rather than something an
author discovers by writing a module that cannot be built: the toolchain list says which languages
exist, and the conformance results say what each can do.
### What the specification covers
Per capability, what two implementations can disagree about:
- **the exchanges and queues** — which exchanges exist, that a consumer's queue is durable and
named `<node>.<module>.events`, that a tool is served from a shared durable `serve.<key>`
- **the envelope** — every header, which are required, what an unknown `x-` header means, and that
ignoring one is correct rather than lax
- **identity** — that a module's node and module name come from its sealed credential and not from
its environment, so what it emits matches what the mesh authorised
- **delivery** — at-least-once, and that dedup is on `x-event-id`, which only the emitter can make
- **the credential** — the sealed document's fields, and that a connection pins a certificate
fingerprint rather than trusting an authority
- **provisioning** — a grant in, a credential out, and what each carries
- **the vocabulary** — that `consumer` is one thing, named once
### Conformance is per capability
**A suite per part, and an SDK claims the parts it passes.** A monolithic pass/fail would make a
partial implementation indistinguishable from a broken one, which is the distinction this is built
on.
**And the suite is executable, not prose.** A specification nobody can run is a document two
implementations drift from while both believe they conform. Conformance is a set of fixtures — an
emitted event, a served tool call, a grant and its answer — that every SDK must produce and consume
byte-for-byte.
**The existing two implementations are the first two to be made to pass it.** Not a future language:
the drift above is present, and a suite that only new SDKs must satisfy would leave the disagreement
that already exists in place while certifying everything added afterwards against it.
## Why not generated types
Generation makes the shapes agree and leaves everything that matters unspecified. Two SDKs
generated from one schema can still name their queues differently, take identity from the
environment, dedup on the wrong field, or omit a header the other requires — and every one of those
is a mesh that runs and quietly does not work.
It also makes the contract into whatever the generator supports, which is a decision nobody made
about a boundary everything else depends on.
**The shapes are worth generating once the wire is specified.** That is a convenience, and it comes
second.
## Consequences
**A language is a commitment, and now a divisible one.** Adding one means implementing the protocol
and passing the suites for the parts it claims. That is more work than transliterating types, and
it is the work that was always there — the difference is that it can be finished, and finished in
pieces, rather than believed.
**Versioning becomes possible.** `x-schema` exists for it and is never written. A specified envelope
with a version on the body is what lets a mesh hold a module built against an older SDK, which is
the ordinary state of any mesh that has been running for a while.
**The two current implementations agree on the live wire** — inspection showed it. What was wrong
was a set of dead types beside the wire, now removed. The suite's job here is therefore prevention:
to keep that agreement true as the wire changes, and to hold a new language's SDK to it, rather than
to repair a break that exists today.
**This does not make the mesh polyglot by itself**, and should not be reported as though it does. It
makes polyglot possible to do correctly. A Rust SDK is still a Rust SDK.
## How this is checked
| Rule | Checked by |
|---|---|
| One vocabulary | A word means one thing across implementations, checked by the fixtures using it. |
| The wire is what is specified | Both existing SDKs run the conformance suite in their own test suites, and a change to one that breaks a fixture fails there rather than in a mesh. |
| A new SDK is a passing SDK | A language is not listed as buildable for a capability until its SDK passes that capability's suite; the toolchain list and the conformance results name the same set. |
| A partial SDK is a real thing | An SDK implementing the floor and one capability passes, is listed for that capability, and a module using another is refused with the reason — rather than failing at runtime in a language nobody said was finished. |
| An unknown header is ignored | A fixture carries one, and every implementation accepts it. |
| Identity comes from the credential | A fixture sets an environment that disagrees with the credential, and the emitted event carries the credential's. |
@@ -0,0 +1,129 @@
---
topic: the tiers
status: accepted
date: 2026-09-15
deciders: jochen
reconstructed: false
extends: 0014-no-npm-workspace.md
---
# 75. An artifact store is a provision; a package registry is a different one
## Context
Two questions have been circling, and they turn out to be one question asked twice.
**"Should gitea be the mesh's registry?"** It serves OCI images and a dozen package ecosystems, it
is already needed — genesis clones from one — and the mesh's own registry has neither
authentication ([issue 042](../04-ISSUES/042-nothing-gives-a-node-an-account-for-a-registry/00-report.md))
nor a transport a runtime will accept over a network
([issue 048](../04-ISSUES/048-nothing-makes-a-machine-trust-the-mesh-registry/00-report.md)). Gitea
has both.
**"Where does the SDK come from?"** [ADR 0014](0014-no-npm-workspace.md) already answers it — each
module consumes its dependencies, the mesh's own shared library included, *from the private
registry* — and nothing installs one, so today it comes from a git URL, which is
[issue 053](../04-ISSUES/053-the-sdk-is-pinned-twice-and-the-two-disagree/00-report.md).
**The framing that dissolves both:** `artifact-store` is already a provision, and `registry`
already provides it. So "should gitea be the registry" is not a question about replacing a
component. It is a question about **a second provider of an existing provision** — which this mesh
has a mechanism for, and uses for certificate authorities and VPNs already.
## Decision
**Two provisions, because they are two jobs.**
| provision | is | for |
|---|---|---|
| `artifact-store` | content-addressed blobs, pinned by digest, no versions, no ranges | what the **mesh** delivers to **machines** |
| `package-registry` | an ecosystem's own registry — npm, cargo, PyPI, Go | what **code** resolves when it is compiled |
They are not the same store with different clients. One is addressed by digest and immutable by
construction; the other is addressed by name and version, and resolves ranges. Conflating them is
how a mesh that pins everything ends up rebuilding one commit into two different things.
**`registry` remains the provider genesis installs.** Not because it is better, but because of what
it is: a directory and one container, no database, no control plane, installable at step 8 of an
install where neither exists yet. Gitea needs a store and provisioning, which means a control plane,
which means the pivot has already happened — and the pivot needs somewhere to publish to.
**Gitea also provides `artifact-store`, and a mesh may choose it.** Two providers of one provision
is a thing the mesh understands: it refuses, names both, and choosing is assigning the one you want.
**And it does not claim `the-artifact-store`.** That claim is node-scoped, so a module holding it
cannot share a machine with another that does — and a machine running gitea for git and packages
*alongside* a registry serving artifacts is an ordinary arrangement, not a conflict. They are
different ports doing different jobs.
The exclusivity that matters is mesh-wide and is already expressed: `provides` at mesh scope means
two providers are two answers, and the resolver refuses until one is assigned. Forbidding
co-residence adds nothing to that and forbids something reasonable. **Whether `registry` should
still hold that claim is left open here** — it may be protecting something about the port or the
data directory that is not written down, and removing a claim is not a thing to do from the outside
of a manifest.
A mesh that assigns gitea gets authentication and TLS for its artifacts — which is to say, **issues
042 and 048 are answered by choosing a provider that already solved them**, rather than by
reimplementing accounts and certificates in a registry that has none.
**Gitea provides `package-registry`.** That is ADR 0014's private registry, and it is one service
rather than one per ecosystem. `verdaccio` may provide it too, for npm alone, and is then a choice
somebody makes rather than the answer.
**The registry is not removed at the end of installing.** A mesh that never runs gitea still has an
artifact store. Retiring it is a migration a mesh performs, not a step an installation ends with.
## Why not simply gitea, from the start
Because genesis would need a control plane before the thing that stores the control plane's image,
and that is circular rather than merely awkward. It would also make one of the three things the
build loop cannot produce for itself into a stateful application with a database — the pivot is the
hardest part of this design already.
And it puts every artifact in the service that is also the trust anchor for everything the mesh will
ever run ([ADR 0071](0071-where-genesis-gets-its-source.md)), which records that forge serving a
cryptominer with tampered git operations. Two blast radii are better than one.
## Moving from one provider to the other is a designed act
**Not a removal.** Every image a machine runs is pinned to a digest at a named store, the control
plane's own included. Changing the provider means:
1. gitea installed, reachable, and holding an account the builder may publish with
2. every artifact mirrored
3. every declaration re-pinned, the control plane's **last**, because it is what performs the others
4. **every machine verified to have converged and to be able to pull from the new store**
5. only then the old provider unassigned, and its volume kept ([ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md))
Step 4 is the one that is easy to skip and the only thing between this and a mesh that cannot
restart its own control plane. A machine that reboots mid-migration pulls from a store that no
longer exists, and a local image cache hides that until exactly the moment it matters.
## Consequences
**The bootstrap is unchanged**, which is the point of keeping the small provider.
**042 and 048 gain a second answer.** They can be fixed in the registry, or dissolved by choosing a
provider that already has accounts and TLS. The second is less work and more service.
**ADR 0014 becomes satisfiable.** There is a provision for the private registry, something that
provides it, and a module may depend on it — so the SDK can be published and consumed rather than
cloned, and issue 053 has somewhere to go.
**A mesh can be minimal or complete, and both are legitimate.** One with the small registry and no
gitea builds and runs modules and cannot serve packages. That is a real configuration, not a broken
one — the same way a partial host is real.
**And the bootstrap still has no package registry.** The first build of the shared base happens
before anything has installed one. That is the same pivot as everything else and it is **not solved
here**: it is named, so the next person does not discover it.
## How this is checked
| Rule | Checked by |
|---|---|
| Genesis needs no database | The installer raises a mesh of one on a machine with nothing, and the artifact store it installs has no store of its own. |
| Two providers are a choice, not a conflict | A mesh holding both is asked to resolve `artifact-store` and refuses, naming both, until one is assigned. |
| Providers may share a machine | A node is assigned both gitea and a registry, and both run — only one of them answers `artifact-store`. |
| The two stores are not interchangeable | A module depending on `package-registry` is not satisfied by `artifact-store`, and the refusal says why. |
| A migration is verified before it is finished | The old provider cannot be unassigned while any machine's declaration still names it. |
@@ -0,0 +1,87 @@
---
topic: building it
status: accepted
date: 2026-09-16
deciders: jochen
reconstructed: false
extends: 0075-two-stores-and-which-provides-what.md
---
# 76. The SDK is a published package, and the toolchain resolves it by version
## Context
[ADR 0014](0014-no-npm-workspace.md) decided a module consumes its dependencies — the mesh's own
shared library included — from the private registry. [ADR 0075](0075-two-stores-and-which-provides-what.md)
decided the private registry is a `package-registry` provision, and that gitea provides it. What
neither settled, and what [issue 053](../04-ISSUES/053-the-sdk-is-pinned-twice-and-the-two-disagree/00-report.md)
left open, is the one build where the rule cannot simply be obeyed: **the first one.**
The TypeScript toolchain image is built *from* the SDK — it carries the SDK so that every module
compiled inside it resolves the shared library without each build fetching it. So the thing that
compiles TypeScript and the thing that contains the SDK were the same object, and that object
cannot be what builds the SDK. Stated as a question — "how does the SDK reach the registry before
the toolchain exists, when the toolchain is what builds it?" — it reads as a paradox.
It is not one. The paradox exists only because the toolchain *bakes a git-cloned copy* of the SDK.
The SDK itself is plain TypeScript: it needs `node` and `tsc` and nothing the mesh makes. A public
base image can build it. The circularity is a property of the workaround, not of the SDK.
## Decision
**The SDK is an ordinary published package in the mesh's `package-registry`, consumed by version.**
The git URL in the toolchain's manifest and the sibling-path lock beside it — the two halves of
issue 053 — are both removed. A build resolves the SDK the way it resolves any dependency, with a
lock that agrees with its manifest, so `npm ci` is the command and reproducibility is by
construction rather than by the machine the build ran on.
**The SDK is built with a public base image, not with the mesh's toolchain.** It is *not* one of
the components the loop cannot build — the control plane, the registry, the builder, the catalogue
([`12-a-module-repository`](../03-DESIGN/01-to-be/12-a-module-repository.md)), which arrive by
carrying an init builder because they are the loop's own machinery. The SDK is machinery for
nothing; it is an ordinary dependency the loop builds and publishes like any other. The only
constraint is narrow: it cannot be compiled *in the mesh toolchain*, because that toolchain is built
from it. So it is compiled on a public base image instead — which needs nothing the mesh makes — and
published before the toolchain that consumes it. It is not carried, because building it does not
wait on a mesh existing first.
**The toolchain base stays, thinned.** mesh-tools remains the image bundles are compiled in and the
one place the SDK is resolved — but it `npm ci`s the SDK by version from the registry instead of
baking a copy cloned from a git URL. Bundles keep borrowing its resolved dependencies; what changes
is that the version they borrow is named and honest. This was the shape chosen over dropping the
shared base entirely and having every bundle resolve the SDK itself: one resolution point, one
place to be right about the version.
**Genesis orders the publish before the first compile.** The package-registry provider is a public
image (gitea), so it comes up needing no toolchain; the SDK is published into it; only then is the
toolchain built, so the first `npm ci` has a registry to read from. Nothing in that chain is
circular, because the only thing that needed the toolchain — baking the SDK — is gone.
## Consequences
Each language's toolchain repeats the shape: its own SDK, built from that language's public base
image, published to the same registry, resolved by version with that ecosystem's lockfile-honest
install (`npm ci`, `cargo` against a vendored or registry source, `pip` against a pinned set). The
warning in issue 053 — that whatever the TypeScript repository does the others will copy — is
answered by making the copied thing the correct one.
A change to the SDK is publish-then-consume, exactly as [ADR 0014](0014-no-npm-workspace.md) already
priced it: publish the new SDK version, then bump the toolchain (and any module pinning it directly)
to consume it. There is no shortcut that resolves an unpublished SDK, which is the property that was
missing.
mesh-tools is no longer an SDK carrier in the sense that mattered — it does not contain a copy
whose provenance is a branch head somebody force-pushes. It contains a version.
A mesh with no package-registry cannot build TypeScript. This is accepted and is not new: it is the
same shape as a mesh that cannot reach a forge being unable to be raised
([ADR 0071](0071-where-genesis-gets-its-source.md)). Installing brings the registry up first.
## How this is checked
| Rule | Checked by |
|---|---|
| The SDK a build compiles against is named, not cloned from a branch | The toolchain manifest pins `@novox/mesh-sdk` to a version, and the build runs `npm ci`, which refuses a lock that disagrees with the manifest. Issue 053's two checks become this one. |
| The SDK builds without the mesh's own toolchain | The SDK's build recipe names a public base image. A recipe that named the mesh toolchain would reintroduce the cycle and is refused in review. |
| The registry is up before the first compile | The genesis bed asserts the package-registry answers, and the SDK is published, before the base build runs. A base build that ran first would fail its `npm ci` with no registry, which is the positive control. |
| A second language repeats the shape, not a new one | When a second SDK is added, its recipe is compared to this one: public base, publish by version, lockfile-honest install. |
@@ -0,0 +1,61 @@
---
topic: the mesh
status: accepted
date: 2026-09-16
deciders: jochen
reconstructed: false
extends: 0006-the-substrate-and-the-control-plane.md
---
# 77. The parts are named controller, foundation, node — not control plane, substrate, master
## Context
The words drifted. In conversation and in code the same thing was called *control plane*,
*controller*, *master*, and *hub*; the store-and-broker pair was called *substrate* and
*foundation*; a machine was a *node*, a *worker-node*, a *peer*, a *slave*. A mesh named
differently by two people is a mesh they describe differently, and the drift was worst on the
parts talked about most.
Three of the terms carried wrong ideas. *Control plane* is borrowed from networking's
control-plane/data-plane split and means nothing here. *Master/slave* and *hub/peer* imply a
subordinate — but no node is: a node applies its own declaration and keeps running when the
control-node dies, so it is as much its own machine as any other. *Substrate* is a biology
metaphor that landed for no one.
## Considered Options
1. **Keep the inherited words.** Rejected: they are the source of the drift, and two of them
(control plane, substrate) are metaphors that teach the wrong shape to anyone reading them cold.
2. **master / slave, or hub / peer, for the nodes.** Rejected: both name a hierarchy the mesh does
not have. The control-node owns no other node; lose it and the rest keep running what they were
last told.
3. **controller / foundation / node + control-node.** Chosen.
## Decision
The component that decides what each node should be, holds the mesh's records, and tells nodes is
the **controller** — the module `mesh-controller`, which claims the mesh-scoped `the-controller`
seat. The store and broker raised at genesis are the **foundation**. Machines are **nodes**;
there are 0..n of them, and exactly one — the one running the controller — is the **control-node**.
Retired: *control plane*, *substrate*, *master/slave*, *hub/peer*, *worker-node*.
[`00-META/glossary.md`](../00-META/glossary.md) is the authority, and a new name for an existing
thing lands there in the change that introduces it in code.
## Consequences
`mesh-control` became `mesh-controller` across the module, container, image, binary, `cmd/` dir and
the git repository; `substrate` became `foundation` in the embedded base bundles, the default
template and the example lock; the seat `the-control-plane` became `the-controller`. The 03-DESIGN
prose and 00-META follow the new words.
What got harder: the records under `02-DECISIONS/` are immutable, so they keep the words they were
written with — this record included, whose own title names what it retires. A term retired here
still appears there, and the glossary is how to read it. The git repository on the forge was
renamed `mesh-control` → `mesh-controller`.
## References
- [`00-META/glossary.md`](../00-META/glossary.md) — one name per thing, and the words retired.
- The rename shipped across all six code repositories and hq (main).
@@ -0,0 +1,63 @@
---
topic: the tiers
status: accepted
date: 2026-09-16
deciders: jochen
reconstructed: false
extends: 0033-the-substrate-is-a-store-and-a-broker.md
---
# 78. The store and the broker are ordinary modules
## Context
[ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md) settled that the foundation is a store
and a broker, raised at genesis; [ADR 0006](0006-the-substrate-and-the-control-plane.md) settled
that the controller cannot grant itself either, because it consumes them and is not running yet to
ask. Both were raised as bundle resources — plumbing, with no record in the mesh's module graph.
That left two costs, named in [issue 051](../04-ISSUES/051-the-mesh-cannot-update-what-it-depends-on/00-report.md).
The foundation's own store and broker could not be upgraded — nothing owned them as modules. And a
mesh that wanted a database or a queue for its modules installed the `postgres`/`lavinmq` modules,
each of which raised a **second** server: a mesh ran two postgres and two brokers.
## Considered Options
1. **Leave them as bundle-only plumbing.** Rejected: they cannot be upgraded, and the second
server stays. The floor keeps a permanent specialty in it.
2. **The control-plane pivot verbatim — raise a temporary one, install the module, retire the
temporary** ([ADR 0067](0067-genesis-is-a-pivot.md)). Rejected for a *stateful* server: it means
a handover with real downtime, tearing down the store the controller is mid-read of.
3. **Adopt in place.** Chosen.
## Decision
The foundation's store and broker are **adopted in place** as the ordinary `postgres` and `lavinmq`
modules. Genesis still raises them first (nothing else can — ADR 0006), then each module declares a
container with the **same name, image and spec** the foundation raised, so the applier — which keys
on the container name and compares a spec digest — reconciles it rather than raising a second. The
credentials are the foundation's, made at genesis and carried in through `secret accept`, because
the mesh cannot invent a credential that already made the databases. The servers bind mesh-wide so
a consumer on any node can reach the one shared server.
A mesh runs **one postgres and one lavinmq**, and each is upgradeable through a stated window: the
store's is a connection-pool reconnect; the broker's is the harder case of recreating the bus the
push travels over, so the mesh reconnects to the one that returns.
## Consequences
The twelve-module floor has no specialty left in it — the store and broker are moments in a
module's life, not a separate kind of thing. Two follow-ups are tracked:
[issue 054](../04-ISSUES/054-the-adopted-store-and-broker-are-open-before-the-filter/00-report.md)
(the adopted servers bind `0.0.0.0` before the packet filter is installed) and
[issue 055](../04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md)
(whether a consumer on another node reaches them over the overlay).
What got harder: a foundation upgrade recreates the very server the controller reads from, or the
bus the instruction to upgrade travels over — a window that a stateless module upgrade does not have.
## References
- [issue 051](../04-ISSUES/051-the-mesh-cannot-update-what-it-depends-on/00-report.md) — the gap this closes.
- [`03-DESIGN/01-to-be/07-the-foundation.md`](../03-DESIGN/01-to-be/07-the-foundation.md) — the amended design.
- Shipped across mesh-host, mesh-catalog and mesh-lab (main); proven 22/22 in the one-node lab.
@@ -0,0 +1,63 @@
---
topic: the tiers
status: accepted
date: 2026-09-17
deciders: jochen
reconstructed: false
extends: 0078-the-store-and-broker-are-modules.md
---
# 79. The foundation seats are named after their servers
## Context
[ADR 0078](0078-the-store-and-broker-are-modules.md) settled that the store and broker are the
ordinary `postgres` and `lavinmq` modules, adopted in place on the control-node, and that a mesh
runs **one postgres and one lavinmq**. But that singularity held only by convention: genesis
assigns them to the control-node alone. [Issue 056](../04-ISSUES/056-an-adopted-module-assigned-to-a-second-node-raises-a-second-server/00-report.md)
recorded the gap — a second `assign` to another node finds no container of that name there and
raises a *second* server, holding none of the first's data, and nothing refuses it.
The mesh already has the mechanism for "there is one of me": a mesh-scoped exclusive **seat**.
[ADR 0077](0077-the-controller-and-the-foundation.md) gave the controller one, which it named
`the-controller`. Nothing gave the store and broker theirs.
## Decision
Each foundation module claims a mesh-scoped seat named **after the server it guards**:
`mesh-controller`, `mesh-store`, `mesh-broker`. So the `postgres` module claims `mesh-store` and
the `lavinmq` module claims `mesh-broker`; the resolver refuses a second holder mesh-wide, the same
way it keeps one hub and one controller. A second `assign` is now a refusal at resolution — *one
per mesh* — not a silent second server.
The controller's seat, which [ADR 0077](0077-the-controller-and-the-foundation.md) named
`the-controller`, is renamed `mesh-controller` under this same rule, so all three foundation seats
follow one convention: the seat is the server. (The module `mesh-controller` and its seat now share
a name, which is the point — there is one of that server, and the seat says so.)
This does not tie a foundation module to the control-node — a mesh-scoped seat forbids a *second*
holder, not a wrong single one. Adoption still requires the container to already be running where
the module lands, which genesis arranges; the seat closes the "two servers" gap, and the
control-node convention remains what puts the one holder in the right place.
## How this is checked
- The resolver's `checkClaims` refuses two holders of a mesh-scoped seat
(`mesh-controller/internal/catalogue/resolve.go`).
- `TestAFoundationModuleCannotBeRaisedOnASecondNode` asserts each foundation module's second
assignment is refused with *one per mesh*.
- The `postgres`, `lavinmq` and `mesh-controller` manifests declare the seat.
## Consequences
"One store, one broker, one controller" is now a property the mesh enforces rather than a
convention it hopes for. The silent operational edge that remains — a cross-node consumer
provisioned only when the provider is pushed again — is separate, and tracked as
[issue 057](../04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md).
## References
- [issue 056](../04-ISSUES/056-an-adopted-module-assigned-to-a-second-node-raises-a-second-server/00-report.md) — the gap this closes.
- [ADR 0077](0077-the-controller-and-the-foundation.md) — named the controller's seat, here renamed.
- [ADR 0078](0078-the-store-and-broker-are-modules.md) — asserted one postgres and one lavinmq, here enforced.
- [`00-META/glossary.md`](../00-META/glossary.md) — seat and claim.
@@ -0,0 +1,55 @@
---
topic: how we work
status: accepted
date: 2026-09-17
deciders: jochen
reconstructed: false
extends: 0019-how-this-repository-works.md
---
# 80. The development cycle is checked, not trusted
## Context
[ADR 0019](0019-how-this-repository-works.md) made this repository the source of truth, and the
process overview drew the flow work must follow: an idea or a symptom, a decision, a to-be design,
a build in a code repository, an as-is update on shipping. The playbooks describe every step, and
frontmatter carries every status.
But the flow itself was enforced by nothing. A design could appear citing no decision; a design
could sit `in-progress` naming no code; an issue could be `fixed` by nobody knows what. Each is
indistinguishable from correct work until somebody reads carefully — and the whole point of the
playbooks is that nobody should have to hold this repository in their head. A session that starts
cold (or an agent after a context clear) must be able to *find* the chain by following frontmatter
pointers, which only works if the pointers are reliably there.
## Decision
The development cycle is enforced mechanically, to the extent frontmatter can carry it:
- **No design without a decision** — every to-be design names at least one record in `decisions:`.
- **No development without a design that says where** — an `in-progress` or `implemented` design
names its owning code in `code:`.
- **No owner-less diagnosis, no fix-less fix** — an issue marked `located` or `fixed` names
`located-in:`; one marked `fixed` or `resolved` says `fixed-by:` (prose counts — "nothing, the
capability existed" is an answer).
- **No silent graduation** — a `graduated` research overview says what it `became:`, and the
targets exist.
[`00-META/checks/cycle.py`](../00-META/checks/cycle.py) refuses violations, beside `records.py`
and `index.py`; all three run before any merge here. What frontmatter cannot see — that code work
actually started from a handoff — remains held by playbooks 04 and 07: a feature branch exists
because a design or an issue sent it, and a merge is a human checkpoint.
## Consequences
A `/clear` costs little: [`AGENTS.md`](../AGENTS.md) now carries the cycle and a where-to-look
table, and the chain a fresh session needs is guaranteed present in frontmatter rather than
reconstructed from memory. The checks are the floor, not the ceiling — they verify pointers exist,
not that their content is true; reading remains the job.
## References
- [`00-META/process/00-overview.md`](../00-META/process/00-overview.md) — the flow, and its new
"The cycle is checked" section.
- [ADR 0019](0019-how-this-repository-works.md) — the repository this disciplines.
@@ -0,0 +1,51 @@
---
topic: how we work
status: accepted
date: 2026-09-17
deciders: jochen
reconstructed: false
extends: 0080-the-development-cycle-is-checked.md
---
# 81. A decision nothing cites is not yet in the chain
## Context
[ADR 0080](0080-the-development-cycle-is-checked.md) made the development cycle checked — but its
checks covered designs, issues and research, not the decisions themselves. Measuring showed why
that matters: 19 of 70 records were cited by nothing — no design doc's `decisions:`, no research
`became:`, no issue, no other record. Among them sat load-bearing decisions (the credential flow,
the module-runtime cluster), and the cost had already been paid once in practice: a stale premise
about an orphaned decision survived in working memory precisely because no pointer led to the
record that had settled it.
## Decision
Every **accepted** decision must be reachable from the cycle: cited by a design doc's frontmatter
(`decisions:` — a governing citation, not a prose mention), a research overview, an issue report,
a `00-META` document, or another record's `extends`/`supersedes` chain.
[`cycle.py`](../00-META/checks/cycle.py) refuses orphans. Proposed records are exempt — a record
under consideration has no home yet — and superseded records are reachable through their
supersession chain by construction.
The 19 orphans were given true homes in the same change: the module-runtime cluster
(0044–0049, 0053–0055) into the connectivity, controller, protocol, writing and model-access
designs; the build decisions (0072, 0076) into the building design; 0021 into the playbook that
implements it; the process records were already reachable once `00-META` counted as a source.
Taken together with 0080, the practice has a name the industry will recognise:
**spec-driven development, with provenance** — a decision is the *why*, the design doc is the
spec, `code:` names the implementation, and the lab beds are the conformance tests. What the
common form leaves implicit, the cycle makes checked: the spec itself must trace to a decision,
and the decision must be findable from the work it governs.
## Consequences
Following pointers now reaches every accepted decision, so a cleared session (or a person) can
trust the frontmatter graph as the whole map. The check is reachability, not truth: a citation
placed wrongly still lies, and reading remains the job.
## References
- [ADR 0080](0080-the-development-cycle-is-checked.md) — the cycle this completes.
- [`00-META/process/00-overview.md`](../00-META/process/00-overview.md) — the flow.
@@ -0,0 +1,84 @@
---
topic: building it
status: accepted
date: 2026-09-17
deciders: jochen
reconstructed: false
extends: 0072-two-graphs-and-the-build-chain.md
---
# 82. The registry is reached by name, and the overlay is its security
## Context
Delivery ends at a node pulling an image, and for every node but the one that built it, that step
has never worked. Three facts conspired, each recorded separately:
- Artifact references are written under `127.0.0.1:5000` — a deliberate parking (the catalogue
says so in the commit that reverted the mesh-reachable name: *"the runtime refused it: 'http:
server gave HTTP response to HTTPS client' … the binding expression returns then"*), so a joined
node is told to pull from its own loopback
([issue 048](../04-ISSUES/048-nothing-makes-a-machine-trust-the-mesh-registry/00-report.md)).
- A container runtime treats any non-loopback registry as HTTPS, and the mesh's registry serves
plain HTTP; nothing the mesh writes tells any runtime otherwise — the one place that file
existed was the lab's, which is why this never failed in a bed (issue 048).
- Nothing gives a node an account for the registry
([issue 042](../04-ISSUES/042-nothing-gives-a-node-an-account-for-a-registry/00-report.md)) —
and nothing has ever said whether one is required.
The tempting fix is TLS from the mesh's own authority. The machinery even exists — the mesh CA
issues leaves for `<node>.internal` names, and a `certificate:` manifest field delivers them. But
the registry is reached **over the overlay**, and the overlay is WireGuard: every byte is already
encrypted and already authenticated to a peer the mesh admitted. The mesh's own code has carried
this position for weeks: *"the registry is reached over the mesh's own private network … a second
layer inside it would be certificates to issue and rotate for no property the first does not
have."* TLS inside the tunnel would also re-order genesis (a certificate needs the network, the
registry precedes it) and put CA handling into three clients (the runtime, the archive fetcher,
the builder) — cost with no new property.
## Decision
**The overlay is the registry's transport security, and the mesh writes the trust it means.**
1. **References name the registry by its mesh name.** The builder publishes to
`<provider>.internal:<port>` (the binding's `at` and `port`) — the parked one-line change
lands. Genesis still publishes to loopback on the first machine, before any overlay exists;
references minted at genesis stay loopback and are valid where they matter — on that machine.
A rebuild re-pins to the mesh name.
2. **Every node on the private network is told the mesh's registry speaks plain HTTP.** The
controller injects, into every such node's declaration, a merged `/etc/docker/daemon.json`
naming `<provider>.internal:<port>` under `insecure-registries`, and a service resource that
restarts the runtime when that file changes — the `/etc/hosts` pattern for the content, the
nftables pattern for the reload. No module author is involved; being on the network is what
grants the trust, because being on the network is what the trust *is*.
3. **No accounts (issue 042), recorded as the position it always was.** Reading and pushing
require presence on the overlay and nothing else. The boundary is enforced, not assumed: the
registry's `listens` is `from: mesh`, the firewall derives from it, and the overlay admits only
peers holding mesh-issued keys. An operator's *external* registry credential is an ordinary
operator-supplied secret (`secret accept`), owned by whichever module names that registry.
Per-node accounts return as a decision, not a patch, if the boundary assumption ever changes —
the shape would be the existing provision flow.
## How this is checked
- The no-fake lab bed: a joined node pulls a mesh-built image by the registry's `.internal` name
over the overlay — an address outside every range the lab's own runtime configuration trusts,
so the mesh-written trust is what makes it work or nothing does.
- The firewall half is generated from `listens` and visible in the node's ruleset (`from: mesh`).
- The plaintext-inside-tunnel position holds exactly as long as the registry is unreachable off
the overlay; the `listens` stanza and the derived ruleset are the check.
## Consequences
A runtime restart when the trust first lands on a node — at joining, before workloads, where it
is free. The registry's HTTP is exposed to whatever stands on the overlay, which is the stated
boundary; a mesh that wants defence in depth inside its own tunnel reopens this record rather
than bolting certificates on quietly. Issues 042 and 048 close on this record.
## References
- [issue 042](../04-ISSUES/042-nothing-gives-a-node-an-account-for-a-registry/00-report.md),
[issue 048](../04-ISSUES/048-nothing-makes-a-machine-trust-the-mesh-registry/00-report.md).
- [ADR 0072](0072-two-graphs-and-the-build-chain.md) — the build chain this completes.
- [ADR 0029](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md) — the overlay as a
boundary.
@@ -0,0 +1,59 @@
---
topic: the mesh
status: accepted
date: 2026-09-18
deciders: jochen
reconstructed: false
extends: 0010-delivery.md
---
# 83. One push leaves the mesh consistent
## Context
A provision is minted while composing the *consumer's* node; the provider's grant list is a pure
read of secrets already issued from it. So assigning a cross-node consumer and pushing its node
produced a consumer that retried forever against a provider that had never heard of it, until the
provider's node was pushed a second time — an action with no signal to take, documented nowhere,
and invisible whenever consumer and provider share a machine
([issue 057](../04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md)).
Two remedies were on the table: **cascade** — a push also delivers to the machines its compose
changed — or **report** — a push says "now push the provider" and leaves the act to the operator.
## Decision
A push finishes what it starts: after composing and sending the named node, the controller flushes
every *other* machine that is now behind — whose declaration differs from what it was last sent —
by name, in the push's own output, converging over a bounded number of rounds (a flushed send may
itself mint).
Behind is measured against what a machine was last *sent*, not against a before/after snapshot of
this push. The mint that makes a provider behind happens when the consumer is assigned or its
account issued — before `push` runs at all — so by push time the provider already differs from
what it holds, with no in-command delta to detect. The only durable signal is "what it should be"
versus "what it last received", which is the same comparison `push --behind` already makes.
A machine behind for an unrelated reason is flushed by this too, and that is correct rather than a
cost: a named push that knew a machine was behind and left it so would be the very silence this
decision removes. The narrower reading — flush only what this push provably changed — was
rejected because it cannot see a mint that a prior command performed, which is precisely the 057
case.
Reporting alone was rejected because it converts a derived fact the controller already holds into
an operator obligation, and an obligation enforced by nothing is issue 057 restated. The
declaration is computed from the whole mesh; delivering a mesh that is knowingly inconsistent and
merely saying so would make "push succeeded" mean less than it says.
## Consequences
- One push is sufficient for a cross-node consumer: the provider's grants arrive from the same
act that minted the provision. The undocumented rule "push the provider node too" ceases to
exist rather than becoming documentation.
- A named push delivers to every machine that is behind, not only the one named — each named in
the output, never silent. `push --behind` remains the way to reconcile the mesh without naming
a node; a named push now carries the same guarantee for the machines its work touched and any
others already waiting.
- How this is checked: the built-store-cross-node bed registers a cross-node consumer, pushes
only the consumer's node, and asserts the provider minted its vhost — the workaround push is
removed, so a regression fails the bed.
@@ -0,0 +1,108 @@
---
topic: what runs on it
status: accepted
date: 2026-09-20
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md
---
# 84. Which provider serves a consumer, when the mesh runs more than one
## Context
[ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) settled what a provision
is *named* for — the thing the consumer's code is coupled to, so `postgres-database` and
`mssql-database` are different provisions and a wrong match is refused at resolution. It said one
thing more, in passing, and left it: *"Two providers of `postgres-database` — a container on this
node and a managed instance elsewhere — are interchangeable and should both match."* Naming was
decided; **which of several providers serves a given consumer was not.**
The mesh assumes there is only one to choose. Provisions are mesh-scoped: the adopted store is a
single `mesh-store` a consumer on any node reaches, and `scope: "mesh"` is written into the
provision definitions. **That assumption is false on day one, and was always meant to be.**
Node-specific services delivered to the mesh is the plan, not an edge case:
- Both control-capable nodes already run their **own** general-purpose relational store (the same
engine, one instance each), serving that node's own applications.
- Each runs its **own** SQL server, its **own** cache, its **own** object store. One node alone
runs six separate relational-store instances, each raised by the module that needed it.
- The one provision that is currently single — the identity provider, one instance on one node —
already authenticates applications whose home is a **different** node.
So several providers of one provision name genuinely coexist, and they are **not** interchangeable
the way 0027's aside supposed. They differ by node, by the data they hold, and by locality. A
consumer bound to the wrong one reads the wrong database, or takes a cross-node hop it did not
need, or cannot be moved without silently rebinding. The model has no field in which to say which
one. This is 0027's own fault — *a match that resolves and is wrong* — one level up: 0027 refused
the wrong **dialect**; nothing refuses, or even asks about, the wrong **instance**.
## Considered Options
1. **Keep `scope: "mesh"` — one provider per provision, mesh-wide.** Rejected: it is false on day
one, and making it true would force every node's applications onto one node's server — the
exact opposite of node-specific services delivered to the mesh, and a single point of failure
the topology was built to avoid.
2. **Resolve to any provider of the name (0027's "both match").** Rejected: when providers hold
different data and live on different nodes they are not interchangeable, and picking one
arbitrarily is a wrong-instance match — the confidently-wrong answer 0027 exists to prevent,
restated at the level of the instance rather than the dialect.
3. **Always require the consumer to name the provider explicitly.** Rejected: needless ceremony in
the common case, where the consumer wants the provider on its own node; and a field every
manifest must carry is a field an author forgets, which then matches everything again — the
failure 0027 warned about for qualifiers.
4. **A provision is node-scoped; the consumer selects the provider, defaulting to co-location.**
Adopted.
## Decision
**A provision is served by a provider identified by its node, and the consumer selects which one.**
A provider is a (node, module) pair, not a mesh-wide singleton. A consumer's binding resolves to a
specific provider, and the selection is part of the assignment
([ADR 0046](0046-a-module-configuration-is-its-assignments-not-its-manifest.md)), not the
manifest.
**The default is co-location.** A consumer that names no provider is served by the provider of
that provision on **its own node**. This is the common case and needs nothing said. A mesh with
one provider of a kind is just the case where co-location and "the only one" coincide — expressed
as *there happens to be one*, not as a scope.
**A consumer coupled to a provider's data names it.** Where two consumers must share one database,
or a consumer must reach a provider on another node, the assignment names that provider — because
that coupling is exactly what may not be guessed, and naming it is what makes a later move safe.
**A module need not consume a provider at all.** It may carry its **own** instance inside its own
composition — on its own module network, publishing no host port, **not** declared as a provision —
when a genuine engine fork or a pinned server version makes the shared provider unusable. Such an
instance is invisible to resolution and can be bound by nothing else. The rule is *share by
default; embed only when a fork or a version forces it* — most of the per-module stores that exist
today are vanilla engines on stale pins that a consolidation onto the node's provider would absorb.
## Consequences
The mesh can carry its real topology **deliberately** rather than by the accident of which
provider happened to be the single one. A consumer's data-coupling becomes a stated fact, which is
what lets a provider be moved without a consumer silently following the wrong one — provided a
provider keeps its identity across a relocation, which is a follow-up this record opens rather than
closes. Rotation ([13](../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md)) addresses a
specific provider's holders, not "the provision's".
What got harder: an assignment now may carry a provider selection, and a wrong one is a new way to
misconfigure. It is mitigated the way 0027 mitigated its own: the co-location default removes the
choice in the common case, and genuine ambiguity — several providers, none named, none co-located —
is refused with the candidates named, never resolved by picking.
## References
- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — the naming this extends;
its "both match" aside is the gap closed here.
- [ADR 0031](0031-the-control-plane-authenticates-nobody.md) — identity is exactly such a
provider, and already serves consumers on another node.
- [ADR 0046](0046-a-module-configuration-is-its-assignments-not-its-manifest.md) — where the
selection lives.
- [ADR 0078](0078-the-store-and-broker-are-modules.md) — the adopted store, whose mesh-wide
binding is the single-provider assumption this record replaces.
- [issue 067](../04-ISSUES/067-a-provision-cannot-name-which-provider-serves-it/00-report.md) — the
gap, and the day-one evidence.
- [`03-DESIGN/01-to-be/23-choosing-a-provider.md`](../03-DESIGN/01-to-be/23-choosing-a-provider.md)
— the design.
@@ -0,0 +1,153 @@
---
topic: what runs on it
status: accepted
date: 2026-09-20
amended: 2026-09-20
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0031-the-control-plane-authenticates-nobody.md
---
# 85. A secret is a provision, and the vault is the module that provides it
## Context
[ADR 0031](0031-the-control-plane-authenticates-nobody.md) decided that identity *runs on the
mesh, not of it* — a module other modules require, rather than a privileged part of the
controller. [ADR 0078](0078-the-store-and-broker-are-modules.md) did the same for the store and
the broker: the twelve-module floor has no specialty left in it. **Secrets are the exception that
survived.** No module owns a secret.
Secret handling is smeared across three built-in parts of the runtime:
- the **controller mints** one credential per consumer↔provider pair and seals it to both node
keys ([ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md));
- the **mesh generates local secrets** — *"generated secrets are the mesh's, never authored"*
(as-is, `06-configuration-and-secrets.md`) — for a value a single module needs for its own use;
- the **synchroniser injects** both into a node's generated files.
Nothing is the owner of "a secret" the way the store module is the owner of "a database", and the
cost is recorded rather than hypothetical. The as-is design names the weakness in its own words:
*"Rotation is not a mesh operation… there is no mechanism that rotates one and informs everything
holding it."* A `rotate` command has since been built and proven, but it reaches **only** the
provisioned pairs; a secret a module generates for its own fully-local use — the password of a
version-pinned embedded store ([ADR 0084](0084-which-provider-serves-a-consumer.md)), an internal
token — is minted by the mesh and then has no operation that can remake it. Three species of
secret, and only the first has an owner:
| species | minted by | rotates? |
|---|---|---|
| a provisioned credential (a database login) | controller, sealed to nodes (0048) | yes — `rotate`, per pair |
| a module's own local secret | the mesh, as a generated value | **no owner, no rotation** |
| an operator-delivered secret (an external key) | a person, sealed in (`secret accept`) | no rotation, no audit |
## Considered Options
1. **Leave it a property of the controller.** Rejected: it is the smear above — no owner, local
secrets that cannot be rotated, operator secrets that cannot be audited — and it is exactly the
specialty 0078 removed for the store and broker, kept here for no reason anyone recorded.
2. **A dedicated vault built into the foundation, not a module.** Rejected: it reintroduces a
privileged built-in, the thing 0031 and 0078 went out of their way to remove, and a mesh that
wants none would still carry it.
3. **Fold all minting, the controller's provisioning credentials included, into the vault.**
Rejected: the controller must mint in order to **deliver** any provision — the vault's own
credential among them — so making the vault mint the credential of its own delivery is the
store/broker chicken-and-egg for no gain. The provisioned-pair credential already has an owner
and a rotation ([13](../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md)); this record
does not disturb it.
4. **A secret is a provision; the vault is an ordinary module that provides it.** Adopted.
## Decision
**Secret-holding is a module, parallel to identity.** A module that needs a secret **for its own
use** — a local service's password, an internal token, an external key it was handed — requires a
`secret` provision from a vault provider, exactly as it requires a database from the store. The
vault generates the value (or holds one it was given), and because the credential belongs to the
consumer↔vault pair it **rotates, backs up and is audited through the same per-pair machinery**
[13](../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md) already defines. That is what
gives the second and third species the rotation and audit they lack. A module's own secret stops
being a generated value that nothing owns and becomes an ordinary provision with a provider.
**The controller's minting of provisioning-pair credentials is unchanged** ([ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)):
it is how every provision, the vault's own included, is delivered. The vault does not mint the
mesh's delivery credentials; it provides secrets to modules, and is itself provisioned the ordinary
way.
**The vault is node-scoped like every provider** ([ADR 0084](0084-which-provider-serves-a-consumer.md)):
each node its own, selected the same way, so a module's own secret is held by the vault on the
module's node. **A mesh that wants no vault runs none** — a module requiring no secret needs
nothing, which is the same test 0031 applied to identity.
## Consequences
The gap that opened this — a generated local secret with no rotation — **closes without new
machinery**: rotating such a secret is the vault's provisioner remaking a pair credential, the
operation 13 already specifies. Backup, audit and break-glass gain an owner — the vault module —
and become things a design specifies rather than absences. The as-is sentence *"rotation is not a
mesh operation"* is already false for provisioned pairs and, once the vault ships, for local
secrets too; the as-is document is updated when it does, not before.
What got harder: a break-glass path — recovering a secret when the sealed delivery path is
unavailable — must not reintroduce a key that one place holds, which is the property
[ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md) was built to preserve; how
the vault offers recovery without it is left to the design as an open question. And a module that
today bakes a password into its own composition must instead require it from the vault — a
migration taken module by module, not a flag day.
## Amendment — 2026-09-20, before anything shipped
Recorded on the record itself rather than as a supersession, by its decider, on the day it was
accepted and before any code was merged against the sentences that change. The original text above
is left as written; this section says what it got wrong and what stands instead.
**What it got wrong.** The decision treated the vault as a provider like the store — optional, and
one per node — and left the mesh's own root secrets outside it: the store's superuser, the broker's
administrator, the controller's contexts, sealed to a node key and nothing else. Those are the
secrets with no rotation and no recovery, and they are the ones a vault exists for. At genesis they
are not even secret: the foundation raises its store and broker with fixed, well-known credentials
and carries those into the mesh. Leaving that floor in place gave module secrets an owner and the
root secrets none.
**What stands instead.**
- **The vault is a foundation module.** It is installed at genesis as part of the foundation
([ADR 0078](0078-the-store-and-broker-are-modules.md) is the precedent: a foundation piece is
still an ordinary module), not assigned later by a mesh that happens to want one. *"A mesh that
wants no vault runs none"* is withdrawn. A mesh has root secrets, so a mesh has a vault.
- **One per mesh, on the control-node.** *"The vault is node-scoped like every provider"* is
withdrawn. Node scoping ([ADR 0084](0084-which-provider-serves-a-consumer.md)) exists because a
store holds data a consumer is coupled to; the vault holds nothing a consumer is coupled to, and a
second one would be a second place to lose. Consumers on other nodes reach it as they reach
identity.
- **The vault holds the mesh's root secrets under an operator-held key.** The controller mints and
delivers exactly as before; in addition, every secret a module holds for itself is sealed a second
time, to an **operator sealing key** whose private half never enters the mesh. The vault keeps
those operator-sealed copies on its own disk, outside the store, and can hand them out — they are
ciphertext to everything but the operator. This is the break-glass path the original text left
open, and it does **not** reintroduce a key one place holds: the mesh holds blobs it cannot open,
and the operator holds a key with nothing to open until given a blob. Recovery needs both.
- **Genesis mints real root secrets and seals them to the operator key first**, so the fixed
credentials the foundation is raised with are replaced before the mesh is handed over.
**Unchanged.** A module's own secret is a `secret` provision the controller mints and the vault
records; the provisioned-pair path of [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)
is untouched; the vault stores no plaintext, ever.
**What it costs.** An operator key is a thing a person must keep, and a mesh whose operator key is
lost has root secrets that can be rotated but not recovered — the same standing as today, stated.
Sealing every own secret twice is a column and a call. Genesis grows a step. A module's
vault-provided secret (a pair credential) is not yet sealed to the operator key; that is the next
increment, not this one.
## References
- [ADR 0031](0031-the-control-plane-authenticates-nobody.md) — identity is a module; this is the
same move for secrets.
- [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md) — the provisioning-credential
path, left unchanged.
- [ADR 0078](0078-the-store-and-broker-are-modules.md) — the de-specialisation this completes.
- [ADR 0084](0084-which-provider-serves-a-consumer.md) — the node-scoping the vault obeys.
- [issue 068](../04-ISSUES/068-secrets-have-no-owning-module/00-report.md) — the gap.
- [`03-DESIGN/01-to-be/24-the-secrets-vault.md`](../03-DESIGN/01-to-be/24-the-secrets-vault.md) —
the design. The source mesh's `secret_locate` / `secret_backup` / `secret_verify` /
`secret_breakglass` subsystem is the prior art it draws on.
@@ -0,0 +1,75 @@
---
topic: building it
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0085-a-secret-is-a-provision.md
---
# 86. A secret reaches a process as a file, and an exception is declared
## Context
The mesh seals a secret to the machine that uses it and discards the plaintext; the host unseals it
into a file at 0600 ([ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)). Then
the module hands it to its container through an env-file, and the runtime puts it where anything
on the machine that can talk to the runtime can read it: `docker inspect` prints it, and
`/proc/<pid>/environ` holds it for the life of the process
([issue 041](../04-ISSUES/041-a-sealed-credential-ends-up-in-the-process-environment/00-report.md)).
Two facts made this a decision rather than a fix. **It is a supported shape**: placing
`${secret:…}` in a file a container reads as its environment is what playbook 06 shows, and 36
containers in 25 catalogue modules do it, the controller's own among them. And **nothing says which
secrets are exposed**: a reader cannot tell from a manifest whether a credential is protected from
`inspect` or not, because the manifest looks the same either way. The vault
([ADR 0085](0085-a-secret-is-a-provision.md)) rests on the seal this undoes.
## Considered Options
1. **Leave it: the environment is where configuration goes.** Rejected — it gives the sealed value
back to a routine operation, and the care taken to seal it is then theatre.
2. **Refuse every secret in an environment.** Rejected — some software reads its configuration
from the environment and nothing else, and a rule the catalogue cannot obey is a rule that gets
switched off.
3. **A secret reaches a process as a file; an environment exception is declared, with a reason,
and refused otherwise.** Adopted.
## Decision
**A secret reaches a process as a file.** A module mounts the file the host wrote and points the
program at it; the mesh's own programs accept a `_FILE` twin for every variable that carries a
credential, the way the controller's store connections already did.
**A container that reads a secret from its environment says so.** The manifest key
`secrets-in-environment` on the container carries the reason. It is catalogue-level — the host
never sees it — and it exists so a reader can tell from the manifest which secrets are exposed
that way and why.
**Everything else is refused.** A `${secret:…}` placeholder inside a container's `env` is never
filled and is refused outright. A file carrying a secret that a container names in `env-file` is
refused unless the container declares the exception.
## Consequences
The property *a sealed credential is readable only where it is used* becomes a manifest-level
fact: true where no exception is declared, stated where one is. The controller reads all six of
its credentials from files; the 35 other containers are marked with their reason and convert one
by one where their software accepts a path, which is per-module work under
[issue 041](../04-ISSUES/041-a-sealed-credential-ends-up-in-the-process-environment/00-report.md).
What got harder: a module author meets one more refusal, and the reason they write is only as
honest as they are. What is not decided: the reach of the exposure on a node — which identities can
talk to the runtime — which decides whether a declared exception is a hardening item or something
sharper.
## How it is checked
The catalogue engine refuses at composition, before anything reaches a machine, and its tests
refuse both shapes and prove the reason is stripped from what is sent.
## References
- [issue 041](../04-ISSUES/041-a-sealed-credential-ends-up-in-the-process-environment/00-report.md)
- [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md), [ADR 0085](0085-a-secret-is-a-provision.md)
- [`03-DESIGN/01-to-be/13-credentials-and-their-rotation.md`](../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md)
@@ -0,0 +1,63 @@
---
topic: what runs on it
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0010-delivery.md
---
# 87. A seeded file is created once, and what grows in it is not the mesh's
## Context
A declaration is complete for what the host owns, and the host reconciles what is declared
([ADR 0010](0010-delivery.md)): a file with this content, held to it. That is the only thing a
manifest could say about a file, and it is the wrong thing for a file a module needs to **exist
before first start** and something else then legitimately writes into — an access list a
provisioner appends consumers to and the program persists back, a bootstrap configuration a
program rewrites. Every reconcile restored the seed behind the running program, erased what had
grown in it, and reported success
([issue 035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md)).
A run-once step ([ADR 0052](0052-a-step-that-runs-once-before-a-container.md)) can write a seed
only if absent, and that closed the instance for a broker whose seed is a program's job. It left
the general case: a plain file the mesh writes and never overwrites.
## Considered Options
1. **Two owners never share a file: the provisioner owns it, and first-start ordering is solved
another way.** Rejected as the only answer — some software refuses to start without the file,
and a module that must ship a program merely to write an empty file has been made to write a
program to say one word.
2. **A create-once semantic on a file.** Adopted.
## Decision
A file resource may say `create-once`. The host writes it when it is absent and, when it is
present, leaves it entirely alone — content, mode and owner — and reports it as **kept**, not
corrected. What is in the file then is somebody else's work the mesh asked for. The mesh removes
nothing it did not create ([ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md)); it now
also does not overwrite what it created once and handed over.
On the security question ADR 0010 asks of every new resource behaviour: this **narrows** what a
declaration can do to a machine. A create-once file gives a compromised control plane one fewer
way to change a machine repeatedly — it can seed, once, and never again.
## Consequences
A module says which of its files are seeds, and the difference is visible in the manifest rather
than in whether the file happened to be revisited. A later change to a seed's declared content
does not reach a machine that already has the file; that is the meaning of a seed, and a module
that needs the new content ships it as a run-once step that migrates the existing file.
## How it is checked
The host's apply tests: a seed is created, grown into by hand, reconciled, and the growth survives
with the outcome `kept`. The vault bed declares one on a real node, grows it, pushes again, and
reads it back.
## References
- [issue 035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md)
- [ADR 0010](0010-delivery.md), [ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md), [ADR 0052](0052-a-step-that-runs-once-before-a-container.md)
@@ -0,0 +1,60 @@
---
topic: the mesh
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0078-the-store-and-broker-are-modules.md
---
# 88. The foundation filters before anything listens
## Context
Adopting the store and broker as ordinary modules
([ADR 0078](0078-the-store-and-broker-are-modules.md)) needs them reachable by consumers across
the mesh, so genesis raises them bound to every interface. The packet filter that decides who may
reach them is a module too, installed a dozen steps later. Between the two, a control-node facing
the network has its store and its bus open to anyone who can reach the machine, for the length
of the install ([issue 054](../04-ISSUES/054-the-adopted-store-and-broker-are-open-before-the-filter/00-report.md)).
The design's rule — what a port is reachable from is decided by the filter — is enforced by
nothing for that window.
## Considered Options
1. **Accept the window**: the machine is mid-bootstrap and the exposure matches what the
pre-adoption modules had in steady state. Rejected — the whole point of deriving the filter
was to stop accepting that.
2. **Bind narrowly at genesis and widen once the filter exists.** Rejected — a bind change is a
recreate of the mesh's store during install, and adoption in place needs the same spec.
3. **The foundation carries a filter of its own, applied before the store.** Adopted.
## Decision
The foundation bundle installs the packet filter and loads a base ruleset **before the store and
broker are raised**: drop by default; keep loopback, replies, ping and ssh; keep the mesh's own
ports a node must reach before it is on the private network — the bus it enrols over and the
registry it pulls from; and let the container runtime's own networks through the forward chain so
containers keep working. It is written into the **same table** the filter module later derives, so
that module replaces it wholesale the moment it can compute one from what the mesh knows, and
nothing of the base survives to contradict it.
## Consequences
From its first resource a machine being made into a mesh refuses what it will refuse when
finished; the window closes. What got harder: the base ruleset is static and names two ports the
mesh's derived one also names — a change to which ports the foundation needs is now made in two
places, and the bundle's own test says which.
## How it is checked
The installer's bundle test asserts the filter and its load precede the store and broker and that
the rules name ssh, the bus and the registry and not the store's or broker's client ports. The
genesis bed probes the machine from outside throughout the install: the store's port is never
reachable, while the bus becomes reachable.
## References
- [issue 054](../04-ISSUES/054-the-adopted-store-and-broker-are-open-before-the-filter/00-report.md), issue 047
- [ADR 0078](0078-the-store-and-broker-are-modules.md)
- [`03-DESIGN/01-to-be/07-the-foundation.md`](../03-DESIGN/01-to-be/07-the-foundation.md)
@@ -0,0 +1,69 @@
---
topic: checking it
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0016-the-lab.md
---
# 89. A bed reads the catalogue it proves
## Context
A lab bed installs a catalogue module and asserts what the mesh does with it; "proven in the
lab" is the standard a module must meet before it ships. The beds built the manifests they
install inline — a literal copied from the catalogue when each bed was written, and never since.
Six modules were converted to file-delivered secrets ([ADR 0086](0086-a-secret-reaches-a-process-as-a-file.md))
and not one bed ran the converted shape; each ran its copy, and the copies still delivered
secrets the way the catalogue engine now refuses
([issue 073](../04-ISSUES/073-beds-carry-copies-of-catalogue-manifests/00-report.md)). A run's
receipt named the commits of the host, the controller and the lab, and said nothing about the
catalogue, so a catalogue change and a proven catalogue change were indistinguishable.
The end-to-end design already has the rule this breaks — *the run rebuilds what it tests; an
artifact rebuilt from memory is one rebuilt sometimes* — for binaries and images. A manifest is
an artifact too.
## Considered Options
1. **Keep the copies and check them against the catalogue** — a test that diffs each literal
against the module's manifest, ignoring what the lab must rewrite. Rejected: it keeps two
sources of truth and adds a third thing that can drift, the list of what to ignore.
2. **Read the catalogue, rewriting only what the lab must.** Adopted.
## Decision
A bed that installs a catalogue module reads that module's manifest from the catalogue checkout
the run was pointed at. It may rewrite what the lab must and nothing else: a build artifact
becomes the image the machine holds, an image is pinned to what the machine holds, a host port
is remapped where one machine carries colliding modules, and an address may point at a stand-in
the bed raises in place of an upstream. Everything else is the catalogue's, verbatim.
A bed that needs less than the catalogue declares — no upstream server, a secret in the
environment, a requirement edge removed — is not testing that module. It is a mesh test, and it
carries a name of its own (see [issue 074](../04-ISSUES/074-a-mesh-test-wears-a-catalogue-modules-name/00-report.md)).
The run's receipt names the catalogue's commit alongside the other repositories', so a receipt
taken before a manifest changed says so.
## Consequences
A catalogue change is proven by the beds that install the module, or it is not proven, and the
receipt says which. What got harder: a bed can no longer trim a module to the shape it finds
convenient; it meets the module's declared requirements or gives its fixture another name.
The beds that still carry a copy are declared, each with its reason, and the declared list
only shrinks.
## How it is checked
A unit test in the lab refuses an inline manifest literal that names a catalogue module unless
the bed is declared, with its reason, in the test's own list; a declaration for a bed that no
longer carries the copy is refused too, so the list cannot outlive the debt. The receipt test
asserts the catalogue is claimed whenever the run is pointed at one.
## References
- [issue 073](../04-ISSUES/073-beds-carry-copies-of-catalogue-manifests/00-report.md), [issue 074](../04-ISSUES/074-a-mesh-test-wears-a-catalogue-modules-name/00-report.md)
- [ADR 0016](0016-the-lab.md), [ADR 0086](0086-a-secret-reaches-a-process-as-a-file.md)
- [`03-DESIGN/01-to-be/01-end-to-end-testing.md`](../03-DESIGN/01-to-be/01-end-to-end-testing.md)
@@ -0,0 +1,65 @@
---
topic: the mesh
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0010-delivery.md
---
# 90. A failure that repeats is said to be stuck
## Context
Delivery is a comparison, not a one-shot ([ADR 0010](0010-delivery.md)): a node re-applies the
declaration it holds on a steady interval and reports each time. That is right for a failure that
goes away by itself — the overlay not up yet, a registry briefly unreachable — and it makes a
failure that will never go away look exactly the same. A resource nothing can ever apply is
attempted, fails, is reported, and is attempted again every few minutes, indefinitely; the mesh
keeps one report per machine, replaced, so each attempt arrives as "failed" at a fresh time.
Nothing distinguished "failed once, will succeed when its dependency arrives" from "failed
identically for ever", and nothing escalated the second
([issue 065](../04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md)).
## Considered Options
1. **The host gives up** after some number of attempts. Rejected: the host does not know whether
a failure is permanent — that a registry has not answered three times is not evidence it never
will — and a host that stops trying is a node that must be pushed to again by hand.
2. **A duration** since the failure was first seen. Rejected as the signal: a laptop shut for a
week has had one attempt, and a week is not evidence of anything.
3. **The controller counts identical reports**, and says when there are enough of them. Adopted.
## Decision
The mesh keeps, beside each machine's last report, when the current failure was first reported
and how many reports in a row have said it — the same outcome, the same refusal, the same failed
resources by id. Not by the host's words: an error carrying a duration or a counter would read as
new on every report, and the resource looping on it is exactly what this is for. A report that
says something different starts the count again; a clean apply clears it. **Three identical reports in a row make a machine stuck**: `status` says
so beside the failure, with the count and the time it began, and the machine-readable status
carries the same three facts. The host keeps retrying; being stuck is a statement about the
mesh's knowledge, not an instruction to the machine.
The controller counts rather than the host, because only it sees every node: one stuck machine
and a mesh-wide fault are different situations, and the host cannot tell them apart.
## Consequences
A resource that will never apply is visible from `status` after three reconcile intervals, to
anyone who looks, without being asked for. What got harder: nothing on the machine changes — a
gating failure that stops what follows still stops it, and this only makes the wait visible.
Whether a stuck machine should also be raised as an event, and what a gating failure should do,
stay open in the issue's own questions.
## How it is checked
An inventory test records the same failure three times and asserts the count and the unchanged
start; the same resource failing in other words, and asserts the count went on; a different
failure, and asserts it restarted; a clean apply, and asserts both cleared. The status command's own test asserts a stuck machine is said to be one.
## References
- [issue 065](../04-ISSUES/065-a-permanently-failing-resource-is-retried-for-ever-with-no-escalation/00-report.md)
- [ADR 0010](0010-delivery.md)
- [`03-DESIGN/01-to-be/10-delivery.md`](../03-DESIGN/01-to-be/10-delivery.md)
@@ -0,0 +1,68 @@
---
topic: what runs on it
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0051-shared-data-is-the-operators.md
---
# 91. A mount is declared, and there are three things it can be
## Context
A container's bind mount whose source does not exist is created by the container runtime, as
root, with whatever mode it picks. So `owner` and `mode` — which exist so a module can say who
its data belongs to — never reach the directories that hold data, and the rule that keeps a
directory when a module goes away ([ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md))
does not cover them, because the mesh has never heard of them
([issue 026](../04-ISSUES/026-the-data-directories-are-not-declared/00-report.md)). Fourteen
such mounts were declared by hand; a check that every mount is declared was then written and
withdrawn, because it refused the builder: the builder mounts the container runtime's socket,
which is not its data, already exists, and belongs to the machine. Declaring it as the module's
own directory would be a lie the host would act on.
## Considered Options
1. **Declare the socket as a directory anyway.** Rejected: the host would create, own and
protect a path that is the machine's.
2. **A new manifest field** naming machine paths a module may mount. Rejected: the manifest
already says the module needs the container runtime, and a second field would say the same
thing in paths.
3. **Three declarations, one for each kind of path a mount can be.** Adopted.
## Decision
A container may not mount a path the module never declared, and a path is declared in one of
three ways, which are the three things a path can be:
- **the module's own** — a directory or file resource, or where a secret, a grant or a
contribution lands. Created and owned by the mesh for this module, kept when the module goes;
- **the operator's** — an `accesses` entry ([ADR 0051](0051-shared-data-is-the-operators.md)):
pre-existing, shared, granted for use, never owned;
- **the machine's** — a facility a declared capability grants. `container-runtime` grants its
socket. The path exists, the machine owns it, and the capability is the declaration.
A mount under a declared directory is declared. The check runs where the manifest is parsed,
and names the path and the three remedies.
## Consequences
Every directory that holds a module's data is one the mesh created with the module's owner and
mode, and one ADR 0030 protects. What got harder: a manifest borrowed from a compose file no
longer passes on the strength of its volume lines; each must say what kind of path it mounts.
The table of what a capability grants is small and in the catalogue's parser; a new capability
that grants a path adds a row.
## How it is checked
Manifest tests refuse an undeclared mount, accept one under a declared directory, accept one
the module accesses, and accept the runtime's socket with the capability and refuse it without.
A test parses every manifest in the catalogue beside the checkout and fails on any that breaks
the rule, so the catalogue cannot drift back.
## References
- [issue 026](../04-ISSUES/026-the-data-directories-are-not-declared/00-report.md)
- [ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md), [ADR 0051](0051-shared-data-is-the-operators.md)
- [`03-DESIGN/01-to-be/18-building-a-module.md`](../03-DESIGN/01-to-be/18-building-a-module.md)
@@ -0,0 +1,61 @@
---
topic: the tiers
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0085-a-secret-is-a-provision.md
---
# 92. An operator delivers a pair credential, and the mesh never replaces it
## Context
[ADR 0085](0085-a-secret-is-a-provision.md) names three species of secret and gives the vault
two of them: a module's own secret, which the mesh mints, and an operator-delivered secret — a
credential for something outside the mesh, which only a person can supply. Under 0085 a secret
from the vault is a pair credential between the consumer and the vault. The controller's one
command that takes a value from a person wrote only a module's own secret, sealed to one node.
Nothing could put a value into a pair, so the third species had no entry
([issue 070](../04-ISSUES/070-an-operator-cannot-deliver-a-pair-credential/00-report.md)), and
an operator's credential could be held only as an own secret — un-audited, un-rotatable, the
gap 0085 opened to close.
## Considered Options
1. **A new verb on the pair.** Rejected: `secret accept` already means "a value a person
supplied, sealed on the way in, plaintext discarded"; a second verb would mean the same.
2. **`secret accept` grows a provider end.** Adopted.
## Decision
`secret accept <node> <module> <name> --provider <node>` seals the supplied value to the
consumer's node, to the provider's node, and to the operator's key when the mesh has one, and
records the pair as `accepted`. Every pair credential now says where it came from: `made` or
`accepted`.
An accepted pair is never replaced by a made one. When a sealing key at either end changes, the
mesh cannot re-seal a value it does not hold, so the read is refused and names the remedy —
accept it again. `rotate` refuses an accepted pair for the same reason: the mesh cannot make its
replacement, and deleting it would have the next read mint one, delivered and reported as
applied while failing to authenticate somewhere else entirely. Rotating an accepted credential
is accepting a new value.
## Consequences
A credential for something outside the mesh lives in the vault's ledger with the others, sealed
to both ends and recoverable by the operator. What got harder: a mesh whose node keys change
cannot heal an accepted pair by itself; a person is asked. That is the honest shape — the value
was never the mesh's to make.
## How it is checked
An inventory test accepts a value into a pair, reads it back twice unchanged with origin
`accepted`, asserts `rotate` refuses it naming the remedy while a made pair still rotates;
another changes a node's key and asserts the read is refused, then accepts again and reads.
## References
- [issue 070](../04-ISSUES/070-an-operator-cannot-deliver-a-pair-credential/00-report.md), [issue 069](../04-ISSUES/069-one-secret-provision-yields-one-value/00-report.md)
- [ADR 0085](0085-a-secret-is-a-provision.md)
- [`03-DESIGN/01-to-be/24-the-secrets-vault.md`](../03-DESIGN/01-to-be/24-the-secrets-vault.md)
@@ -0,0 +1,53 @@
---
topic: checking it
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0089-a-bed-reads-the-catalogue-it-proves.md
---
# 93. A fixture that runs a module's runtime carries the module's name
## Context
[ADR 0089](0089-a-bed-reads-the-catalogue-it-proves.md) said a bed that needs less than a module
declares is a mesh test, and carries a name of its own. Twelve beds were to be renamed on that
basis ([issue 074](../04-ISSUES/074-a-mesh-test-wears-a-catalogue-modules-name/00-report.md)).
Reading how a module's tools are reached showed why they cannot be: the account the mesh issues
a module is scoped to `serve.<module>.*` from the manifest's name, and the runtime binds its tool
queues under the name baked into its image. A fixture named otherwise but running the real
runtime would be refused its own queues. The name is not a label; it is the tool namespace and
the broker scope.
## Considered Options
1. **Rename anyway, and rebuild each runtime under the fixture's name.** Rejected: a runtime built
under a false name proves nothing about the module and costs a build per bed.
2. **A fixture that runs a module's runtime carries the module's name, and therefore reads the
catalogue.** Adopted.
## Decision
A bed that runs a module's runtime installs the catalogue's manifest for that module and what it
requires — the vault for a `secret`, the route module for a route — and proves the mechanism
against the real module. A name of its own is for a fixture that runs no real runtime: a
declaration-only stub, a bare upstream image.
## Consequences
The mechanism beds become module beds with a mechanism inside them, which is more than they
were. What got harder: a bed that wanted a cut-down redis now raises the vault beside it; one that
wanted a sidecar without its server raises the server. Each conversion is a lab run, and the beds
still carrying a copy are declared with this reason until converted.
## How it is checked
The lab's inline-copy check refuses undeclared copies as before; the declared list's reason for
these beds names this record. Two beds converted with the vault beside them ran green.
## References
- [issue 074](../04-ISSUES/074-a-mesh-test-wears-a-catalogue-modules-name/00-report.md)
- [ADR 0089](0089-a-bed-reads-the-catalogue-it-proves.md), [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)
- [`03-DESIGN/01-to-be/01-end-to-end-testing.md`](../03-DESIGN/01-to-be/01-end-to-end-testing.md)
@@ -0,0 +1,63 @@
---
topic: the tiers
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0085-a-secret-is-a-provision.md
---
# 94. A module may hold several secrets from one provider, each a pair of its own
## Context
[ADR 0085](0085-a-secret-is-a-provision.md) makes a module's own secret a provision: the module
requires `secret` from the vault and reads the pair credential minted for that consumer↔vault
pair. A pair has one credential, a module requires a provision once, and so a module received
one value. Read against the catalogue, nine modules hold two or more secrets besides their
broker account; seven of them hold genuinely independent values with independent lifetimes — a
root certificate, its key and that key's password; an admin password beside an API token. None
derives from another, so "one value, derivation the module's business" answers nothing
([issue 069](../04-ISSUES/069-one-secret-provision-yields-one-value/00-report.md)).
## Considered Options
1. **One value per module; the module derives the rest.** Rejected: the values are independent.
2. **Require the provision several times.** Rejected: `requires` is a list of names, and a
requirement is matched by name everywhere.
3. **The `secrets` map names several files under local names, and each local name is a pair
credential of its own.** Adopted.
## Decision
A module's `secrets` entry for a requirement may be a path, as before, or an object of local
names to paths. Each local name is its own pair credential, keyed on it beside the provision,
the consumer node, the consumer module and the provider; its own file on the consumer, referred
to as `${secret:<local name>}`; its own holder at the provider, named the consumer's identity
with the local name after it; and rotated apart from the others. A local name may not be one of
the module's own secrets or something it requires, so what a placeholder means is never
ambiguous. The plain shape is unchanged, and every credential that exists is the one it was.
The holder's suffix is not a login any backend checks — a secret is not a login — so the
identity limit that binds a database role or an access key does not apply to it.
## Consequences
The ten modules that could not move onto the vault can. What got harder: `rotate secret` for a
consumer rotates every local name it holds from that provider; rotating one of several is a
finer command than the mesh has, and waits for a case that needs it.
## How it is checked
Manifest tests read both shapes, write them back, and refuse a colliding or unusable local
name. A resolver test asserts two local names are two needs, two files with two credentials,
and two holders at the provider. An inventory test asserts two local names are two rows, that
rotating one leaves the other, and that the provider is told both. The vault bed installs a
consumer that keeps two secrets and asserts two values delivered, two holders in the vault's
ledger, and both rotated by one command.
## References
- [issue 069](../04-ISSUES/069-one-secret-provision-yields-one-value/00-report.md)
- [ADR 0085](0085-a-secret-is-a-provision.md), [ADR 0049](0049-a-consumers-identity-fits-the-tightest-backend.md)
- [`03-DESIGN/01-to-be/24-the-secrets-vault.md`](../03-DESIGN/01-to-be/24-the-secrets-vault.md)
@@ -0,0 +1,60 @@
---
topic: the tiers
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md
---
# 95. The control plane is the way to ask a module
## Context
A module serves tools over the broker under an account scoped to what it emits, consumes and
serves ([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). A
tool call is a request and a reply: the caller creates a reply queue and publishes to the serving
module's request key, and no module's scope grants either — nor should it, since a module that
only publishes events has no business declaring queues. So a module could serve tools and nothing
in the mesh could call them
([issue 049](../04-ISSUES/049-a-module-can-serve-tools-and-nothing-can-call-them/00-report.md)):
not an operator at a terminal, not an agent acting for one.
## Considered Options
1. **Calling is a grant**: a module declares it may be asked, and a consumer is issued an
account that may create a reply queue and publish to that module's request key. Deferred: a
module-to-module call is the only caller that resembles what the mesh mints today, and none
asks for one yet.
2. **The control plane is the way in.** Adopted. It holds a connection that may already, so a
person or an agent asks through it, and every question passes one process — which is where
an audit of who asked what belongs.
## Decision
`mesh-controller ask <module> <tool> [json]` publishes the request on the tool exchange under
`<module>.<tool>`, with a private reply queue bound under its own name, waits for the answer
whose correlation matches, and prints it as the module gave it. A tool that answered with an
error has answered: the answer is printed and the exit status says so. A module that never
answers is said to have not answered, with where to look.
A module declares nothing about being asked: serving a tool is being askable through the control
plane. A module-to-module call, if one is wanted, is a grant like any other and a later decision.
## Consequences
Anything with the control plane in reach can ask any module anything it serves. What got
harder: nothing outside the control plane can, and the control plane's connection is one more
thing on the path of every question — a cost accepted for the audit it buys.
## How it is checked
A tools-only bed asks a served tool through the control plane and asserts an answer arrived —
an error, since the lab has no upstream and no token, which is an answer where a timeout would
not be.
## References
- [issue 049](../04-ISSUES/049-a-module-can-serve-tools-and-nothing-can-call-them/00-report.md)
- [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)
- [`03-DESIGN/01-to-be/19-the-module-protocol.md`](../03-DESIGN/01-to-be/19-the-module-protocol.md)
@@ -0,0 +1,66 @@
---
topic: building it
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0006-the-substrate-and-the-control-plane.md
---
# 96. An upstream image is copied between registries, never through a machine's image store
## Context
A module may declare that an artifact is an image published elsewhere, to be copied into the
mesh's own registry so machines fetch it by a digest this mesh assigned rather than by a name
somebody else controls. The builder pulled it into the build machine's image store and pushed it
under the mesh's name, and the push was refused: a published image is an index over several
architectures, the runtime's store keeps the index, and pushing one platform out of it fails
however the platform is asked for
([issue 046](../04-ISSUES/046-an-upstream-image-cannot-be-mirrored-into-the-mesh/00-report.md)).
Every variant of pull-then-push was tried and failed the same way.
## Considered Options
1. **Resolve the index to one platform and push that.** Tried, reverted: it did not make the
push work, and a workaround for a store's behaviour is a thing nobody removes later.
2. **Tooling that copies between registries**, installed on the build machine. Rejected: one
more thing the builder's image carries, for a protocol the builder already speaks for blobs.
3. **Copy over the registry API**, in the builder. Adopted.
## Decision
The builder copies an upstream image between registries and never through a machine's image
store: it reads the index and every manifest it names, moves each blob by digest into the mesh's
registry — skipping what is already there, since blobs are content-named — puts the manifests
and then the index under the module's repository, and pins the index's digest. Public images are
read with the anonymous bearer token the registry hands out on challenge, which is how the
public hub and the others the catalogue names serve them. The mesh mirrors the whole index, so
what a machine fetches is the image for its own architecture; that every machine on one mesh is
the same architecture is an assumption this mesh makes and had not written down until now.
Genesis has no registry to copy into and keeps the pull: the image stays in the first machine's
store, named by its own id, as every artifact does before there is anywhere to publish.
## Consequences
An upstream artifact builds. What got harder: the builder now holds a registry client of its
own, some two hundred lines, where a runtime command used to do; and a private upstream that
demands a credential is refused, since the copy is anonymous by design.
## How it is checked
A test raises a fake upstream registry serving an index over two platforms behind a bearer
challenge, and a fake mesh registry that records what arrives: every blob of both platforms
arrives once, two manifests and the index are put under their digests, the reference returned
pins the index under the module's repository, and a second copy uploads nothing. A reference
test reads names the way a runtime does. *Proven against the real thing the same day:* the genesis
bed built the tool runtime through the mesh's builder with its node base copied out of the public
hub into the mesh's registry by this code — after one finding the fake could not give: the builder
ran on the default bridge, where loopback is not the machine, and now runs on the host network.
## References
- [issue 046](../04-ISSUES/046-an-upstream-image-cannot-be-mirrored-into-the-mesh/00-report.md)
- [ADR 0006](0006-the-substrate-and-the-control-plane.md)
- [`03-DESIGN/01-to-be/18-building-a-module.md`](../03-DESIGN/01-to-be/18-building-a-module.md)
@@ -0,0 +1,65 @@
---
topic: building it
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0096-an-upstream-image-is-copied-between-registries.md
---
# 97. A vendor image is a declared build input, and a recipe fetches nothing undeclared
## Context
Three modules could not be built by the mesh's builder because their recipes reached for what
no manifest named: a public package, or a binary copied out of a public image
([issue 064](../04-ISSUES/064-a-mesh-build-cannot-fetch-a-modules-external-dependencies/00-report.md)).
A module already names the bases it stands on — another module's artifact, by name — and the
builder answers with what the mesh holds; a vendor's image had no such declaration, so a recipe
named it directly and the build worked when the public registry answered, which is sometimes.
## Considered Options
1. **Let the build environment reach public registries.** Rejected: a build that fetches from
somebody else's registry on its own is one the mesh cannot rebuild the same way twice.
2. **A vendor image is a base like any other**, declared under `build.on` with the argument the
recipe reads it from, pinned by digest, copied into the mesh's registry before the build
([ADR 0096](0096-an-upstream-image-is-copied-between-registries.md)). Adopted.
## Decision
A build's `on` entry is either a module's artifact or an image published elsewhere, pinned by
digest, read from one build argument. Before the build the image is copied into the mesh's
registry under the module's repository and the recipe is handed the copy; genesis, with no
registry, pulls it into the first machine's store. A recipe whose `COPY --from` names a registry
image the manifest did not declare is refused before the build, naming the image and the remedy;
its own stages, declared arguments and `scratch` are not fetches. An unpinned vendor image is
refused: a tag is what somebody else can move.
A recipe whose `FROM` names an undeclared base was at first said, not refused: the mesh's own
images — the control plane's, the tool runtime's, the route proxy's — started from a public base
and declared none, and refusing those refuses genesis. *Amended the same day:* those three declare
their bases now, and an undeclared `FROM` is refused like an undeclared copy. The builder's own
image and the examples are built by `make`, not by the mesh, and take arguments with defaults.
The package half of the issue is not decided here: the mesh's package registry already proxies
the public one, and the failure the report saw has to be run again to be placed.
## Consequences
A module's build inputs are all in its manifest, and every one of them is something the mesh
holds a copy of. What got harder: a recipe that used to name a base image on its first line now
names an argument, and the manifest names the image.
## How it is checked
A builder test declares a pinned vendor image, asserts it is copied under the module's
repository and handed to the recipe as the argument, and asserts an unpinned one is refused. A
recipe test asserts an undeclared `FROM`, an undeclared `COPY --from` and an undeclared argument
are named, and that stages, declared arguments and `scratch` are not.
## References
- [issue 064](../04-ISSUES/064-a-mesh-build-cannot-fetch-a-modules-external-dependencies/00-report.md)
- [ADR 0096](0096-an-upstream-image-is-copied-between-registries.md)
- [`03-DESIGN/01-to-be/18-building-a-module.md`](../03-DESIGN/01-to-be/18-building-a-module.md)
@@ -0,0 +1,61 @@
---
topic: the tiers
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0085-a-secret-is-a-provision.md
---
# 98. A fact a provider makes at first start is fetched from it, not carried in its manifest
## Context
The catalogue's certificate authority declared its root certificate, its root key and that key's
password as its own secrets, and told the container to initialise from them. The mesh mints an
own secret nobody delivers as random bytes, and random bytes are not a certificate: issued, the
authority could not start; only an operator hand-making its root could raise it
([issue 076](../04-ISSUES/076-a-served-fact-made-at-first-start-cannot-be-served/00-report.md)).
The authority can make its own root at first start. What it could not do then was tell the mesh
what that root is: a consumer was given `${bound:acme-ca:root}` from the provider's `serves`,
which is written in the manifest before anything runs.
## Considered Options
1. **A secret the module makes**, with the mesh taking custody once the file exists. Rejected
for now: a node would have to send a value up to the mesh, which no channel does today, and
a root key is the one thing the mesh has no reason to hold.
2. **A served fact the provider contributes at run time.** Rejected for now: the same new
channel, for a fact that is not secret at all.
3. **The consumer fetches it from the provider**, over the mesh network, through a gate before
the thing that needs it starts. Adopted.
## Decision
A provider's `serves` names where a fact made at first start can be fetched — the authority
serves its root at a path beside its ACME directory — and a consumer fetches it in a `run-once`
step declared before the resource that needs it, from the provider's bound address. The mesh
network is where the fetch happens, which is what makes fetching without a prior trust
acceptable: it is the network the mesh itself authenticates. The mesh mints only what it can
make: the authority's password. The root key stays where it was made.
## Consequences
The catalogue's authority starts, and the proxy that requires it trusts what it fetched. What
got harder: a consumer of such a fact carries one more resource, the gate that fetches it, and
a fact that changes after first start is refetched only when the declaration changes.
## How it is checked
The route-forwarding bed installs the authority, the proxy and a consumer from the catalogue and
asserts a routed name is served through the proxy. The proxy refuses to start on a bundle that is
not a certificate, so the name being served proves the gate fetched one; the gate itself refuses
a body that is not a certificate. That the proxy obtains a certificate from this authority through
that root is the certificate bed's proof, against the same authority with the same proxy. The
catalogue-wide manifest test parses both manifests.
## References
- [issue 076](../04-ISSUES/076-a-served-fact-made-at-first-start-cannot-be-served/00-report.md)
- [ADR 0053](0053-a-step-that-runs-on-a-schedule.md), [ADR 0085](0085-a-secret-is-a-provision.md)
- [`03-DESIGN/01-to-be/08-connectivity.md`](../03-DESIGN/01-to-be/08-connectivity.md)
@@ -0,0 +1,94 @@
---
topic: what runs on it
status: accepted
date: 2026-09-21
deciders: jochen
reconstructed: false
extends: 0052-a-step-that-runs-once-before-a-container.md
---
# 99. A step that runs once names what it reads, and runs again when it changed
## Context
[ADR 0098](0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md) has a consumer
fetch a fact its provider made at first start through a run-once step: the route proxy fetches
the certificate authority's root before it starts. A run-once step runs once per declaration
([ADR 0052](0052-a-step-that-runs-once-before-a-container.md)): its marker is the digest of its
own declaration, and a re-apply that finds the marker does nothing.
The provider can move. When the authority is assigned to another node it makes a new root there,
and the mesh rewrites the consumer's binding file with the new address — but the step's own
declaration has not changed, so the step does not run again, the proxy keeps the old root, and it
refuses every certificate the new authority issues
([issue 077](../04-ISSUES/077-a-fact-fetched-at-first-start-is-fetched-once/00-report.md)). A
restart trigger was the natural remedy and was refused on a run-once step, on the ground that a
step does not stay running to be restarted.
Two things the host already does point at the answer. What a container reads is part of what it
is: a container's digest includes the digest of every resource it names under `restart-on`, so a
rewritten file it reads is a changed container
([issue 045](../04-ISSUES/045-a-container-keeps-the-values-it-started-with/00-report.md)).
And a run-once step's marker *is* its digest. Nothing new is needed for the step to run again when
what it reads changed; only the refusal stands in the way.
## Decision
**A run-once step may name what it reads under `restart-on`. For a step the word means *run
again*: when a named resource changed in this apply, the step's digest has moved, its marker no
longer matches, and it runs again — gating what follows, as it did the first time.** Nothing
about the marker changes; the refusal of the pair is lifted, in the control plane and on the host.
**The container that consumes what a step made names the step.** A step that ran counts as a
change, so a service that names it under `restart-on` is recreated after it, holding what the step
fetched. Without this the step fetches a new root and the service keeps serving with the old one.
The route proxy's gate names the binding file it reads; the proxy's server names the gate and the
binding. When the authority moves, the binding is rewritten, the gate fetches the new root, and the
server is recreated with it — in one apply.
## Considered Options
1. **A provider epoch in the binding — the mesh raises a number when a provider is re-issued or
moved, and the consumer's file carries it.** Rejected: the binding already changes when the
provider moves (its address does), and a re-issue does not change what the authority serves —
its state persists. An epoch would be a second signal for a change the file already shows.
2. **The step runs before every start of the service, with no marker.** Rejected: every reconcile
would run it, and a step that runs on every apply reads as a change on every apply, so the
service naming it would be recreated every few minutes.
3. **The proxy fetches the root itself, at start.** Rejected as the general answer: it fixes the
proxy and leaves the next consumer of a fact made at first start to fix itself. The step is the
general shape ([ADR 0098](0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md)).
4. **Lift the refusal and read `restart-on` as *again* on a step.** Adopted: it is what the
digest already does, and it needs no new word.
## Consequences
A fact fetched at first start follows its provider when the provider moves. What is not covered:
a provider whose state is wiped behind the mesh's back, on the same node, makes a new fact that
nothing the mesh knows reflects. That is not a change the mesh can see, and it is not claimed.
The one contradiction the refusal named is real and is now a documented reading: on a running
container `restart-on` means recreate, on a step it means run again. Both are "this must reflect
what it reads". This record is about a run-once *container*, the step ADR 0052 defined. A
run-once process is a different shape whose marker does not carry what it reads; it is not
covered here.
## How it is checked
- mesh-host: a unit test declares a run-once step naming a file, records its marker against the
file's old content, applies with the new content and asserts the step ran; applies again with
nothing changed and asserts it did not. A second test declares a container naming a run-once
step and asserts the container is recreated after the step ran, with the step as the stated
reason.
- mesh-controller: the manifest parser accepts a run-once step with `restart-on`; the
catalogue-wide manifest test parses the route proxy's manifest, whose gate and server name what
they read.
- The route-forwarding bed still passes with the host that accepts the pair. No bed moves the
authority: the mechanism is proven by the unit tests, the declaration by the manifest test.
## References
- [issue 077](../04-ISSUES/077-a-fact-fetched-at-first-start-is-fetched-once/00-report.md)
- [ADR 0052](0052-a-step-that-runs-once-before-a-container.md), [ADR 0053](0053-a-step-that-runs-on-a-schedule.md), [ADR 0098](0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md)
- [`03-DESIGN/01-to-be/08-connectivity.md`](../03-DESIGN/01-to-be/08-connectivity.md), [`03-DESIGN/01-to-be/20-writing-a-module.md`](../03-DESIGN/01-to-be/20-writing-a-module.md)
@@ -0,0 +1,221 @@
---
topic: the mesh
status: accepted
date: 2026-09-22
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0078-the-store-and-broker-are-modules.md
---
# 100. A node in use is adopted before it is converged
## Context
The mesh replaces a predecessor mesh that is running, on the same machines, with the services
people use. The control-node is the machine that already carries the predecessor's broker and
build pipeline. The operator proposed the migration's shape: stop the predecessor's control on a
machine, bring the mesh up there in an adoption mode that keeps the machine's configuration in
force, migrate its modules one at a time, move to the next machine, and flip adoption off when every
machine is done.
Measured on the control-node ([research 012,
*migrating a node that is in use*](../01-RESEARCH/012-the-minimum-viable-node/migrating-a-node-in-use.md)):
60 predecessor containers; the predecessor's control in user-level daemons separate from its
services; its firewall active, with 54 incoming and 52 forwarding rules allowing each served port
explicitly; 12 files its configuration sync writes, 3 of them system files.
Four things in the mesh as it stands break that shape:
1. **The base ruleset closes the machine.** Genesis loads a drop-by-default table
([ADR 0088](0088-the-foundation-filters-before-anything-listens.md)). Every base chain at a hook
runs in priority order; an accept ends only its own chain and a drop in any is final — whether
the other firewall's chains are nftables or legacy iptables. So the predecessor's allowed ports
would close at genesis.
2. **The host replaces what it finds.** A declared file is written whatever is at its path, except
a file declared create-once. A declared container replaces a running one of the same name. So
assigning a module the predecessor also runs replaces the predecessor's service and its files at
once.
3. **Some foundation ports are held.** The foundation's store and bus ports are free — the
predecessor publishes its own elsewhere — but the registry's port and the broker's management
port are held, and on a control-node that is the private network's hub, so is the private
network's port. The foundation's ports are fixed in the installer's bundle and the catalogue's
manifests, so a collision surfaces as a container that fails to bind, and a port changed at
genesis would be changed back when the foundation is adopted as modules.
4. **Published container ports are forwarded, not received.** The foundation publishes its ports
on every interface; a predecessor firewall that filters only incoming traffic never sees them.
The base ruleset is what keeps the store unreachable from outside, and it is the thing that
cannot be loaded.
The mesh already adopts in one place: the foundation's store and broker are taken over in place as
modules, keyed on the container that is already running
([ADR 0078](0078-the-store-and-broker-are-modules.md)). The research this record draws on already
settled the conflict rule for adoption: *on conflict, what is on the machine stays*.
## Considered Options
1. **Cut each machine over in one go** — stop the predecessor's store, broker, registry and proxy,
raise the foundation in their place. Rejected: every predecessor service on the machine is down
until it has migrated, the predecessor's other machines lose their broker, and the rollback is
restarting the predecessor — a recovery, not a step.
2. **Put the control-node on a separate machine.** Rejected: the control-node is decided.
3. **Converge on joining, as today.** Rejected: the base ruleset closes the machine at genesis, and
the predecessor's services and files are replaced as soon as any module naming them is assigned.
4. **Only make the foundation's ports configurable.** Rejected as insufficient: it answers the bind
collisions and neither the firewall nor the files.
5. **Treat assigning a module as migrating it.** Considered and rejected on review: it makes the
rule that keeps found files never fire — the host only ever sees files of assigned modules — and
it makes an assignment on an adopted node an outage rather than a preparation.
6. **Adoption as a mode per node; each module taken explicitly; the node converged by an explicit
flip.** Adopted.
## Decision
**A node is adopted or converged, and the controller records which.** The operator says so: at
genesis for the control-node, and in the enrolment token for the others. The controller is
authoritative, and every declaration it sends says whether the node is adopted and which modules
have been taken on it. A node stays adopted until the operator converges it. An adopted node is
said to be adopted wherever the mesh reports a node's state.
**A converged genesis refuses a machine in use.** A machine is in use when a container is running
on it, or a port is listening on an address other than loopback that is not ssh's. Raised without
saying adopted on such a machine, genesis refuses and names every container and listener it
counted — a forgotten flag must not close a working machine.
**Before a node is adopted, its predecessor's control is stopped by the operator** — the daemons
that write its configuration. Its services keep running on what they have.
**Found means present with no record.** A file at a declared path, or a container at a declared
name, that the host's store has no record of writing is *found*. A file the host wrote in an
earlier life of the node is not found; its record says so.
**On an adopted node, what is found is kept until its module is taken.** The host keeps a found
file and a found container as they are, records the file's original content before anything else
happens to it, and reports each as held. Assigning a module on an adopted node prepares it: what
the module declares that is not found is created; what is found is held. **Taking a module** on a
node is its cutover — the operator's act, done when that module's data has moved — and from then
on the module's resources converge on that node like any other. What is held is never removed,
even when its module is unassigned, and a held file or container that changes while held — a file
rewritten, a container stopped or replaced — is reported as changed by something else, not reverted
or restarted: that is how a predecessor still writing is caught.
**The firewall found on the machine stays in force.** The mesh loads no table on an adopted node
that drops by default or holds an accept — neither genesis's base ruleset nor the filter module's
derived one. What the mesh needs reachable is declared as **openings**: a resource that says a port
is reachable, from where, on the incoming path or the forwarded path — a published container port is
forwarded. The controller derives them from the same inputs as the filter, each from where the
filter would admit it: the `listens` of the modules assigned there, the private network's hub port
and the bus and the registry from anywhere, the store's port and the broker's management port from
the private network. The host
converges an opening through the found firewall in that firewall's own terms, marks it as the
mesh's, and removes only what it marked; it re-checks each opening on every reconcile, so a reload
or a reboot of the found firewall does not lose it for longer than one reconcile. An opening is a
state, not a command, which is what lets it travel over the link. The host reports which firewall
it found. A machine with no firewall needs no openings; a machine with a kind no host speaks is
refused adoption.
**The mesh guards its own ports itself, in a table of its own that only refuses.** It passes
everything by default and holds nothing but refusals, and the two ports it refuses are the
foundation's own, checked free at genesis, so it cannot close anything the machine serves; it is
the mesh's, so the found firewall reloading does not touch it. It refuses the store's port and the
broker's management port except from the private network and from the machine itself — its
loopback and the container runtime's own networks, known by the interface a packet arrives on and
never by its source address alone — at the prerouting hook, ahead of the runtime's
destination translation, so it matches the port the packet was sent to, for both address
families. The bus and the registry stay reachable from anywhere, as a node
enrols over the bus and pulls from the registry before it has a private-network address
([ADR 0088](0088-the-foundation-filters-before-anything-listens.md)); so does the private network's
hub port. The store is unreachable from outside whatever the found firewall does, and on a machine
with none.
**The foundation's ports are the node's.** Every port the foundation binds is an input to genesis,
checked free before anything is raised, refused with the name of what holds it. The ports given
become that node's settings for the foundation's modules — the catalogue's numbers are only their
defaults — and every place that uses them reads them from there: the modules' containers, the
filter, the base ruleset, the private network's endpoint, and the addresses consumers are given.
The private network's address range must not overlap a tunnel the predecessor still runs; genesis
checks that too.
**Converging a node is one act, previewed.** It refuses while an assigned module still holds a
found container: each service is taken on its own, when its data has moved, never by the flip. The
preview lists what is reachable on the machine now — every listening socket and every published
container port — and for each whether an assigned module declares it or it will close, and every
module the flip will take, with the held files each will replace. The flip then takes those modules, loads the
mesh's derived filter in place of its refusal-only table, and retires the found firewall by
disabling it, never by flushing: the container runtime's rules and the found firewall's own
configuration stay on disk. Returning a converged node to adopted unloads the derived filter,
restores the refusal-only table, enables the found firewall again and converges the openings
through it once more; what was taken stays taken. A
node converges when its migration is done; the mesh is migrated when every node has converged.
**The order is the operator's:** the control-node first, adopted, its modules assigned and taken
one at a time; then each other machine, adopted, migrated, converged in turn. The predecessor's
pipeline runs on the control-node, so its updates stop for every machine while the migration runs;
that is accepted.
## Consequences
Each step says what it changes before it changes it. Adopting a node changes nothing that serves;
assigning a module adds what is not there; taking a module replaces one service; the flip replaces
the firewall, after naming every port it will close. Two steps are not undone by the mesh: taking a
module replaces the predecessor's container, and the kept original of a file is recorded but not
yet restored by any act of the mesh
([research 012](../01-RESEARCH/012-the-minimum-viable-node/00-overview.md) leaves where it lives
open).
What got harder:
- **The mesh must speak a firewall it did not install**, on both the incoming and the forwarded
path. One kind is found on the machines measured; another is refused until a host speaks it. The
mesh's own guard does not depend on it: that table is the mesh's.
- **The host gains a guard it did not have** — keep what you found — and its report must say
which files and containers it holds, or an adopted node reads as converged.
- **The declaration gains a node's mode and its taken modules**, and a resource, the opening.
- **Genesis grows inputs, and they outlive genesis.** The foundation's ports stop being constants;
every reader of them reads the node's settings.
- **A node can sit adopted indefinitely.** Nothing forces the flip; the mesh's status says which
nodes are adopted, so one left behind is visible.
- **This narrows [ADR 0088](0088-the-foundation-filters-before-anything-listens.md) for adopted
nodes**: the base ruleset is not loaded on a node raised adopted, and its duty — the store never
reachable from outside — passes to a table of the mesh's that only refuses.
## How it is checked
A lab bed prepares a machine the way the predecessor leaves one: its firewall allowing a served
port and denying the rest, a service container listening on that port under a name a catalogue
module also uses, a file at a path that module declares, a stand-in for the predecessor's control
that would rewrite that file, and a container holding the registry's port. Then:
- **Genesis converged** on it refuses and names every container and listener it counted.
- **Genesis adopted, with the registry's port held**, refuses and names the holder; with another
port given, the foundation comes up — and adopting the foundation as modules leaves it on that
port.
- **Nothing that serves changed**: the service is reachable from a second machine, the file is byte
for byte what it was, and the found firewall's rules differ only by rules marked as the mesh's.
- **The store is unreachable from outside** — probed from a machine off the private network, and
again after the found firewall is reloaded — and reachable over it and from a container on the
node itself; the bus is reachable from a machine that has not yet enrolled.
- **The mesh works through the found firewall, and keeps working after it is reloaded and after the
machine reboots**: the second machine enrols, and the openings are there again.
- **A predecessor still writing is caught**: with the stand-in left running, the held file's change
is reported and not reverted.
- **Assigning prepares, taking cuts over**: the module assigned holds the found container and file;
taken, it replaces them and its port is opened.
- **Converging previews, then changes**: it refuses while the service's module holds its found
container; once that module is taken, the preview names the service's port and a published port no
firewall rule mentions, and the modules it will take; after the flip the mesh's derived filter is
loaded, the found firewall is disabled with its configuration still on disk, the declared port is
open and the undeclared one closed. Returned to adopted, the found firewall is enabled again and
the derived filter is gone.
Unit tests hold the host to keeping a found file and container on an adopted node, converging them
once taken, never removing what it holds, and reporting a held file or container that changed;
genesis to refusing a held port and a converged raise on a machine in use; the controller to
carrying the mode and the taken modules in every declaration, deriving the openings, and refusing
a flip while a found container is held.
## References
- [research 012 — the minimum viable node, and adopting what is already there](../01-RESEARCH/012-the-minimum-viable-node/00-overview.md),
and its document [*migrating a node that is in use*](../01-RESEARCH/012-the-minimum-viable-node/migrating-a-node-in-use.md)
- [ADR 0005](0005-the-node-host.md), [ADR 0011](0011-managed-files-are-generated-never-edited.md),
[ADR 0078](0078-the-store-and-broker-are-modules.md), [ADR 0088](0088-the-foundation-filters-before-anything-listens.md)
@@ -0,0 +1,70 @@
---
topic: the mesh
status: accepted
date: 2026-09-22
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
---
# 101. A machine's own resolver does not make it in use
## Context
[ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md) has a converged genesis
refuse a machine in use. It defines *in use* as a container running, or a port listening on an
address other than loopback that is not ssh's. That definition was written before anything was
measured.
Measured on a freshly installed lab machine, running nothing but its operating system:
| Listening | Held by |
|---|---|
| TCP and UDP on every address, the link-local name resolution port | the system's name resolver |
| UDP on every address, the multicast name resolution port | the system's name resolver |
| UDP on a link-local address, the address-configuration client port | the system's network manager |
| TCP and UDP on loopback | the resolver's stub and the container runtime |
By ADR 0100's words, the resolver's TCP listener on every address makes **every** freshly
installed machine a machine in use. A converged genesis would refuse them all, and `--adopted`
would become the only way to raise anything. The refusal exists to catch a forgotten flag on a
working machine. Refusing an empty one defeats it, and teaches operators to pass the flag by
habit.
## Considered Options
1. **Keep the words.** Rejected: every fresh machine is refused.
2. **Count TCP only, ignore UDP.** Rejected: the resolver listens on TCP too, and a machine that
serves over UDP alone, a resolver or a tunnel, is in use.
3. **Ignore listeners held by the operating system's own network daemons**, a short named list,
on both protocols. Adopted.
## Decision
**A listener held by one of the operating system's own network daemons does not make a machine
in use.** The daemons are the ones the measurement found: the name resolver and the network
manager, named in the installer's code beside that measurement. Everything else in ADR 0100's definition stands: a running container, or any other
listener on an address other than loopback that is not ssh's, makes the machine in use, and
genesis still names every one it counted.
A daemon is added to the list only with a measurement of a fresh machine that holds it.
## Consequences
- A converged genesis on a fresh machine goes ahead, as it did before ADR 0100.
- A machine whose resolver is also serving other machines is not counted as in use by its
resolver alone. It is one of the daemons that serves nobody on a fresh machine, and the one
kind of service this lets through.
- The list is code, not configuration, so it changes by review.
## How it is checked
A unit test in the installer feeds the listeners captured from the fresh machine, as the
listening-socket tool printed them, and asserts the machine is not in use. The existing tests
still assert that a serving machine is in use and that every container and listener is named.
The adoption lab bed raises a converged genesis on a fresh machine and asserts it is not refused.
## References
- [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md)
- [research 012, *migrating a node that is in use*](../01-RESEARCH/012-the-minimum-viable-node/migrating-a-node-in-use.md)
@@ -0,0 +1,93 @@
---
topic: the mesh
status: accepted
date: 2026-09-22
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
---
# 102. The mesh writes into a shared file, never over it
## Context
[Issue 084](../04-ISSUES/084-taking-networking-on-an-adopted-node-restarts-every-container/00-report.md)
found that the networking module declares the container runtime's configuration file to state
one fact in it: the mesh's registry is trusted over the private network. Reading the code
closer showed it worse than reported. The controller merges an operator's settings into the
module's own content, but the host writes the result **whole**. Whatever the machine had in that
file is replaced, including where the runtime keeps its data. On a machine in use, that is every
image and container gone from the runtime's view at its next start. The runtime's service is then
restarted, which stops every container on the machine.
Measured on a lab machine: the runtime takes a new list of trusted registries on a reload, with
no restart, and a running container with no restart policy keeps running through it. The
runtime's log says it reloaded its configuration, and the registry reads as trusted afterwards.
Under [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md), the file is found
on an adopted node and held until networking is taken. Holding it is safe, but it means an
adopted machine cannot pull the mesh's images until then, and taking networking would restart
the runtime.
## Considered Options
1. **Keep writing the whole file; take networking as a cutover of its own.** Rejected: it
still replaces the machine's settings, and the cutover stops everything on the machine.
2. **A drop-in the runtime reads beside its main file.** Rejected: the runtime has no such
directory for its daemon settings.
3. **Write into the file: set the mesh's keys, keep the rest; reload, don't restart.** Adopted.
## Decision
**A file the mesh shares with software it did not install is written into, never over.** A
file resource may say it is written *into* a structured file. The host then reads what is there,
sets only the keys the mesh declares, keeps every other key as it found it, and records what each
of its keys held before. **A list is added to, never replaced**: where the machine already has a
list under a key the mesh declares, the mesh's members are added to it and the host records
exactly which members it added — setting the key would replace the operator's own list, the harm
this record exists to prevent. Undeclared later, each key goes back to what it held and each
added member is taken out again, and a file the mesh created is removed only if nothing but its
own keys is left. A file written into replaces nothing, so on an adopted node it is never held:
it is written whether or not its module has been taken.
**Whatever the host writes over without a record of it, it keeps first.** On any node, adopted
or converged, before the host writes a file over one it has no record of making, it keeps the
original once and says where; if it cannot keep it, it does not write. A file the mesh takes over
is then never lost, whatever put it there.
**A service that re-reads its configuration on a reload is reloaded, not restarted.** A service
resource may name what it must be *reloaded* on, beside what it must be restarted on. The
container runtime is reloaded for the registry's trust.
**The networking module writes the runtime's trust into its file and reloads it.** An adopted
node therefore trusts the mesh's registry as soon as it is on the private network, and taking
networking no longer touches the runtime. The hosts file networking writes is still written whole
and stays held until networking is taken; a converge preview names it among the files it replaces.
## Consequences
- The runtime's file on a machine in use keeps its data directory, its logging settings and
everything else the predecessor set.
- A host that does not know *into* or *reload-on* refuses a declaration carrying them, so hosts
are upgraded before the controller that emits them — the same order ADR 0100 needs.
- One shape more for every host: a file written into a structured document. Only JSON is spoken;
another format is refused until written.
- The hosts file remains a whole file. Writing a marked block into it is the same idea for a text
file and is not decided here.
## How it is checked
Unit tests hold the host to setting only the declared keys and keeping the rest, restoring each
key and removing only a file it created when the resource is undeclared, adding to a list and
removing only the members it added, keeping the original of a file it writes over without a
record, refusing a file that is not a JSON object rather than overwriting it, never holding a file written into on an adopted
node, and reloading rather than restarting a service whose reload-on resource changed. A test in
the controller holds the networking module to declaring the runtime's file written into and the
runtime reloaded. The adoption lab bed gives the machine a runtime file with a setting of its own
and asserts it survives adoption with the registry trusted added, and that a container without a
restart policy is still running afterwards.
## References
- [Issue 084](../04-ISSUES/084-taking-networking-on-an-adopted-node-restarts-every-container/00-report.md)
- [ADR 0082](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md), [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md)
@@ -0,0 +1,112 @@
---
topic: the mesh
status: accepted
date: 2026-09-22
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
---
# 103. What an adopted node holds, and what its guard refuses
## Context
[ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md) holds a found file or
container until its module is taken. It also has the mesh guard the store's port and the
broker's management port with a table that only refuses. Two independent reviews of the build
found both rules drawn too narrowly and, for the guard, in the wrong place:
1. **Other kinds of resource reach what was found.** An untaken module's directory is re-owned
and re-moded if a predecessor's directory is already at its path, and a database refuses to
start on a data directory whose mode changed. A service of the same name as a predecessor's
unit is started, stopped or re-enabled. An action run *in* a held container runs inside the
predecessor's service. A container that is not found under its own name is created, and can
mount the predecessor's data beside the predecessor's own container.
2. **The guard's two ports are not the only ones a found firewall misses.** A published
container port is forwarded, not received, and a firewall that filters only incoming traffic
never sees it (ADR 0100's own Context). The broker's plaintext port is published on every
interface and admitted by the filter from the private network only. So on an adopted node it
is reachable from anywhere. The same holds for any published port a module restricts to the
private network.
3. **The guard refuses by port alone.** On a machine that routes for others, such as a
predecessor's private-network hub, a packet for another machine's database port is refused as
well. And a port the guard refuses for a module that is assigned but not taken may still be
the predecessor's own, serving the predecessor's other machines.
## Decision
**Found covers every kind that can reach what the machine already has.** On an adopted node,
for a module not yet taken:
- a **directory** present with no record is held: its mode and owner are left, and nothing in it
is touched;
- a **service** with no record is held when an administrator installed its unit — the service
manager reads the unit from outside the packages' own directory — or when the machine uses it,
running or started at boot. Its state and whether it starts at boot are then left. A reload
named by the module still happens, since a reload stops nothing
([ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md)). A unit a package merely
ships that nothing runs and nothing enables is not found: the private network's own tunnel unit
is one, and holding it kept the private network from ever coming up;
- an **archive** whose target is present with no record is not unpacked or re-owned; a
**process** whose unit file is present with no record is not written over or restarted; a
**user** that exists with no record keeps its shell and groups;
- a **container** not found by its name is still held if it would mount a path or a volume that
is present with no record, because it would share the predecessor's data;
- an **action** or one-off step run *in* a held container is held until the module is taken.
What is held is reported as held and is never removed, as ADR 0100 decides for files and
containers.
**The guard refuses what the found firewall may miss, for taken modules only.** Its ports are
derived from each *taken* module: every machine port it publishes that the filter would admit
from the private network only, and the ports its manifest names under `guards` — which is how
the store's port and the broker's management port are included. Only TCP is guarded. A port of a module that is assigned but not taken is not guarded: it may still
be the predecessor's. The guard matches only packets addressed to this machine. Traffic the
machine routes for others is never its business.
**An opening already answered by a found rule is not added.** If the found firewall already
admits what an opening says, the opening is reported as satisfied by the found rule, and the
mesh adds nothing and later removes nothing. The found firewall treats two rules differing only
in their action, log setting or comment as one rule, so adding the mesh's would take over the
operator's. Where the found rule is such a twin but is not a plain allow — a deny, a limit, a
logged allow — the opening is refused, naming that rule, and nothing is added.
**A machine raised adopted stays adopted if genesis is run again.** The installer reads the
node's mode from what the machine records, not only from the operator's flag. A run without
the flag on an adopted machine is refused.
**What of ADR 0100 this replaces.** ADR 0100's guard refused two ports, the foundation's own,
checked free at genesis, and could therefore close nothing the machine served. That guard is
replaced by the one above: a guarded port of a taken module may be one the predecessor served
more widely, and taking the module narrows it to the private network. Everything else in ADR 0100
stands, and its rules for files and containers now cover the other kinds listed here.
## Consequences
- Taking a module can narrow a port the predecessor served to anyone: the guard then refuses it
from outside the private network, as the module declares. Nothing says so yet
([issue 086](../04-ISSUES/086-taking-a-module-narrows-a-port-without-saying-so/00-report.md)).
- An untaken module on an adopted node can come up only beside what was found, never on top of
it: whatever would share the predecessor's data waits for the cutover.
- The guard grows with the modules taken, and a module's published private-network port is
protected on an adopted node the way the derived filter protects it on a converged one.
- Guarding only taken modules means a foundation port on a node joining adopted is guarded once
its module is taken, not when it is assigned. Genesis takes the foundation's modules, so the
control-node's store is guarded from the first push.
## How it is checked
Unit tests hold the host to holding a found directory, a found service, a found archive, process
and user, and a container that would mount found data; to not holding a unit nothing runs; to
deferring an action run in a held container; to adding no opening a found rule answers, and to
refusing one a conflicting found rule would absorb. They hold the controller to deriving the guard from taken modules'
private-network published ports, and the guard's text to matching only this machine's
addresses. The installer's tests hold a re-run without the flag on an adopted machine to a
refusal. The adoption lab bed asserts that the broker's plaintext port is unreachable from
outside the private network even with the found firewall admitting it, and that an operator's
own rule for a port the mesh opens survives the opening being removed.
## References
- [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md),
[ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md)
@@ -0,0 +1,93 @@
---
topic: the mesh
status: accepted
date: 2026-09-22
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
---
# 104. A provision may be answered by an adapter to the predecessor
## Context
[Issue 093](../04-ISSUES/093-the-successor-proxy-cannot-serve-what-the-predecessor-still-serves/00-report.md),
found on the first adopted node hours after it was raised. Every module reachable by name requires
a **route**. The mesh has one provider of it, its own proxy, which binds the two public ports on
the machine's own network. The predecessor's proxy holds those ports and serves every public name
there. So:
- a web module cannot be taken until the mesh's proxy runs;
- the mesh's proxy cannot run until the predecessor's stops;
- and when it stops, every name the predecessor served goes dark, because the mesh's proxy serves
only what mesh modules have contributed.
Adoption ([ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md)) keeps a *file* as it
was and a *container* as it was. It has nothing for a service whose successor is **different
software**: nothing shares a name to hold, so there is only replacement, of every name at once.
The certificates decide the rest. The predecessor's proxy holds them and renews them. A successor
taking over first would have to obtain a certificate for every public name in the same window in
which it took the ports — with nothing else migrated yet, and the only way back being to start the
predecessor again.
## Considered Options
1. **The successor proxy serves routes it did not derive**, given as the operator's configuration,
each disappearing as the module owning that name contributes its own. Rejected for the
migration: it puts the riskiest step first, and it puts configuration the mesh does not own
into the one component whose design is that the mesh derives its configuration.
2. **Take the proxy first and accept the window.** Rejected: every public name at once, before
anything else has moved, with certificates to obtain in the same window.
3. **Make the requirement optional**, so a module can be taken while nothing provides route.
Rejected: it is not optional. A module that says it must be reachable and is not is a module
reporting success while serving nobody.
4. **An adapter module provides the provision by writing into the predecessor's service.** Adopted.
## Decision
**A provision whose only provider owns a scarce machine-wide resource may be answered, during a
migration, by an adapter module that writes into the predecessor's own configuration.**
For the route: a module that **provides `route`**, receives every contribution as the proxy does,
and writes each one as a route file where the predecessor's proxy reads its configuration —
pointing at the machine port the contributing module now publishes. The predecessor's proxy keeps
serving every name it already serves, keeps its certificates and keeps renewing them; a name whose
module has migrated is served by the same proxy, pointing at the mesh's container instead of the
predecessor's.
**It is migration scaffolding and says so.** It is assigned only on an adopted node, it writes only
files it can name as its own, and it is removed when the predecessor's proxy retires — at which
point the mesh's own proxy takes the ports and already knows every route, because by then every
name belongs to a module that contributes it.
**This does not make the mesh the predecessor's manager.** The adapter writes route files and
nothing else, and only because the predecessor's control has been stopped by the operator, which
ADR 0100 requires before a node is adopted.
## Consequences
- The migration keeps its shape: one service at a time, each reversible on its own. The proxy is
the **last** cutover again, not the first.
- There is no certificate event until the end, and at the end there is one, for a proxy that
already has every route from the mesh.
- The mesh writes into a directory the predecessor owns. That is safe only while the predecessor's
control is stopped, and it is why the adapter is refused on a converged node.
- The adapter is throwaway code with a stated end. Something has to delete it; that something is
the operator, and the flip is the moment.
- A second provision with this shape — a resolver on 53, a mail relay on 25 — now has a pattern to
follow rather than a new argument to have.
## How it is checked
Unit tests hold the adapter to writing one route file per contribution, naming each file as its
own, removing a file when its contribution goes, and leaving every file it did not write. The
controller is held to resolving `route` from it exactly as from the proxy. The adoption lab bed
gains a step: with the predecessor's proxy serving a name, a module taken behind the adapter is
reachable under that name, and the predecessor's own names keep answering. Assigning the adapter
to a converged node is refused.
## References
- [Issue 093](../04-ISSUES/093-the-successor-proxy-cannot-serve-what-the-predecessor-still-serves/00-report.md)
- [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md), [ADR 0007](0007-connectivity.md)
@@ -0,0 +1,146 @@
---
topic: the mesh
status: accepted
date: 2026-09-23
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
---
# 105. The mesh adopts the predecessor's tunnel in place
## Context
The control-node is adopted and its store and broker have been moved onto the ports the
predecessor served them on, so the predecessor's other machines keep reaching them. They reach
them **over the predecessor's tunnel**: a WireGuard interface on the control-node with three
peers, an address range, and a port the hosting provider already lets through. The mesh's own
private network runs beside it on a second interface, a second range and a second port — one the
provider does not let through, so no other machine can join the mesh
([research 012](../01-RESEARCH/012-the-minimum-viable-node/00-overview.md), the runbook's
measurement of the upstream filter).
[ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md) says the mesh's range *must not
overlap a tunnel the predecessor still runs*. That was written for coexistence. It leaves the
migration with two tunnels for as long as any predecessor machine exists, and the second one
unreachable.
The mesh already adopts in place where the thing found is the thing it would have raised: the
store and the broker were taken over as modules, keyed on the container that was running
([ADR 0078](0078-the-store-and-broker-are-modules.md)). A WireGuard interface is the same shape: a
private key, a listening port, a list of peers by public key, an address. The mesh's is not
different in kind from the predecessor's; it is a second one.
The operator's instruction: take over the tunnel interface, same range — everything stays the same.
## Considered Options
1. **Two tunnels until the last predecessor machine is gone.** Rejected: the mesh's stays
unreachable from outside, so no machine can join, so the last predecessor machine is never gone.
2. **Move the mesh's tunnel onto the predecessor's port with the mesh's own key and range.** The
peers' packets arrive and are dropped — WireGuard authenticates by key, and the mesh's key is not
the one they know. Every other machine loses its tunnel until it enrols, and it enrols over a bus
it reaches through that tunnel. Rejected.
3. **Adopt the predecessor's tunnel in place: its private key, its peers, its range, its port.**
Adopted.
## Decision
**On an adopted node that is the hub, the private network takes over the tunnel it finds.** The
mesh's interface is raised with the found interface's **private key**, on its **port**, with its
**address and range**, and every **peer** the found interface had — public key, allowed address —
carried into the mesh's peer list as a peer not yet enrolled. The found interface is stopped, never
flushed; its configuration stays on disk, kept like any held file.
**Nothing a peer knows changes.** A predecessor machine keeps the same server key, the same
endpoint, the same address and the same route; it cannot tell the tunnel changed hands. When that
machine enrols, it keeps its address: the mesh assigns an enrolling node the address the tunnel
already had for its key, and only a node with no such address is given a fresh one from the range.
**The mesh's own addresses are the range's.** The controller composes every node's private address
from the tunnel it holds, so adopting the predecessor's range moves the mesh's addresses with it —
the hub's, and every binding, hosts-file entry and endpoint derived from it. Those are readers of
the setting; they follow it, per ADR 0100's rule for ports. A reader that does not follow is
[issue 102](../04-ISSUES/102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md).
**This narrows ADR 0100.** Its rule that the range must not overlap a tunnel the predecessor still
runs applies to a node that is *not* adopting the tunnel: where the found tunnel is left running
beside the mesh's, the ranges must differ. Where it is adopted, there is one tunnel and one range.
**The guard's question answers itself.** [ADR 0103](0103-what-an-adopted-node-holds-and-what-its-guard-refuses.md)
admits the mesh's own ports only from the private network's interface. With one tunnel, the
predecessor's peers arrive on it, and nothing needs to be admitted from an interface the mesh does
not own.
## Consequences
- **The order of a migration changes.** The hub's tunnel can be taken as soon as the node is
adopted, before any service — and should be, because it is what lets other machines join. The
runbook's step order is amended.
- **The hub takes the provider's open port for free.** The port the predecessor's tunnel used is
by definition one the provider passes.
- **A peer's identity precedes its enrolment.** The mesh holds public keys and addresses for
machines it has no record of. They are peers of the tunnel, not nodes of the mesh, until they
enrol; the registry must be able to say both.
- **What got harder:** the private key of the found interface is read from the machine and becomes
the mesh's — the one case where the mesh takes a credential it did not mint. It is sealed like
any own secret from then on, and the found configuration file is kept, not copied further.
- Two documents currently say the opposite: ADR 0100's non-overlap rule (narrowed above) and the
runbook's port plan, which is amended with this record.
## How it is checked
A lab bed prepares a hub the way the predecessor leaves one: a WireGuard interface with a key, a
port, a range and two peers, each peer a second machine that reaches a service on the hub through
the tunnel. Then:
- **Adopted, the tunnel changes hands and the peers notice nothing**: the found interface is down
and its file is on disk; the mesh's interface is up with the found key, port and address; each
peer's service call succeeds before, during and after, with no reconfiguration on the peer.
- **A peer enrols and keeps its address**: the machine joins the mesh over the tunnel it already
has, and its node address is the one the tunnel held for it.
- **A new machine gets a fresh address from the same range**, and reaches both the hub and the
enrolled peer.
- **Nothing derived from the address is stale**: every binding, hosts entry and endpoint the
controller composes says the adopted range, before and after a push.
Unit tests hold the controller to reading the hub's address and range from the adopted tunnel,
assigning an enrolling node the address its key already had, and refusing to hand out an address
the tunnel already holds; and the host to raising the mesh's interface with the found key and
peers and stopping the found interface without flushing it.
## What review settled that this record did not
*Added 2026-09-24, from the review of the implementation. A decision record is not edited to change
its meaning; this says what was decided under it.*
- **A spoke's view of its hub is not a peer the mesh carries.** A predecessor gives a spoke the
whole subnet through the hub, so the spoke's found tunnel names one peer routed a range rather
than an address. Only the hub's peers are ever carried; a spoke presenting its own is skipped,
not refused — a machine enrols with what it found, and what it found is its route home.
- **The range and the carried peers outlive the flip.** They follow from the hub having taken the
tunnel over — its key being the tunnel's — and not from the node being adopted. A converged hub
keeps the range it adopted and the addresses it is holding, and `AssignAddress` keeps excluding
them.
- **Converging the hub is not refused while a carried peer has not enrolled.** Proposed in review
and rejected on the record's own terms: *a node converges when its migration is done*, and the
other machines' migrations are not this node's. With the range and the peers surviving the flip
there is nothing left for the refusal to protect, and it would have made one machine's converge
wait on every other machine.
- **A takeover whose placement disagrees with the tunnel is refused before it is composed** — an
address or a port that is not the tunnel's would stop the found interface and raise the mesh's
somewhere the peers are not, while reporting success.
- **The host says three things, not two**: the found interface still up, the mesh's up in its place,
or — the state worth naming — the found one down and the mesh's not up, which is the only one
where the peers reach nothing.
- **A hub that enrolled before this existed keeps its identity.** Re-enrolling would have remade
every credential in the mesh, because the hub provides the store and the broker. Instead the
machine takes the tunnel's key as its overlay key and says so in a message signed with the
identity it already has, so a forged report cannot move a node's key.
## References
- [ADR 0078](0078-the-store-and-broker-are-modules.md), [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md),
[ADR 0103](0103-what-an-adopted-node-holds-and-what-its-guard-refuses.md)
- [issue 102](../04-ISSUES/102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md)
- [research 012 — the minimum viable node](../01-RESEARCH/012-the-minimum-viable-node/00-overview.md)
+83
View File
@@ -0,0 +1,83 @@
---
topic: the mesh
status: accepted
date: 2026-09-23
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0002-nodes-communicate-over-a-broker.md
---
# 106. The bus is NATS
## Context
The mesh's bus is an AMQP broker. [Research 014](../01-RESEARCH/014-the-bus-on-nats/00-overview.md)
measured what that costs and what it would take to change: AMQP is spoken in three places of the
mesh's own code — the controller's link, the host's link, the tool runtime's client — and in none of
the sdk or the modules, because [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) kept the client
out of the sdk for exactly this. The predecessor's world is AMQP and is being retired module by
module; it cannot move and does not need to.
The operator's direction: NATS, native — the mesh built on it, not bridged to it — and the tools a
person reaches from a workstation designed on it rather than built quickly on what exists.
## Considered Options
1. **Stay on AMQP.** Rejected by the operator.
2. **NATS behind a bridge**, the mesh unchanged. Rejected: two buses, every guarantee crossing a
seam, and the reason for changing — one native, simple, subject-addressed bus with request/reply
and accounts built in — lost at the seam.
3. **NATS native, built in the lab in parallel with the migration, cut over in one rehearsed
rollout after the migration's core is done.** Adopted.
## Decision
**The mesh's bus is NATS.** Every link the mesh has — control, node queues, builds, enrolment,
reports, events, tool invocation — is carried on NATS subjects; durability, catch-up and
hold-unacknowledged-and-retry ([ADR 0083](0083-one-push-leaves-the-mesh-consistent.md)) are
JetStream's; a module's account is a NATS account with per-subject permissions derived from its
`emits` and `consumes` ([ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md)),
written as configuration the host declares and the server reloads — never through a management
API. The broker module changes; the seat `mesh-broker` does not
([ADR 0079](0079-the-foundation-seats-are-named-after-their-servers.md)).
**The sdk's broker contract does not change.** Modules are written against `request`, `handle`,
`publish`, `subscribe`; the runtime implements them on NATS; no module speaks a protocol.
**The AMQP broker the mesh adopted becomes the predecessor's compatibility broker**, kept as a
module with one purpose — the predecessor's clients — and retired with the last of them.
**It is built beside the migration and cut over after its core.** The bus is not changed under a
half-migrated node. A lab bed proves the whole path — enrolment, the store window, upgrades, tool
invocation — on NATS before any node's bus moves, and the move is one rollout: controller, every
host, every runtime together.
**What a person reaches the mesh's tools with is designed on NATS**, as part of the same design: a
person's account, scoped like a module's, and a client that speaks the bus directly.
## Consequences
- The architecture — subjects, streams, accounts, the enrolment handshake, the person's client — is
written as a to-be design before code, and reviewed.
- Eight records are touched: 0002 and 0033 (the substrate names a broker — now NATS), 0041 and
0042 (events keep their shape; the exchange becomes a subject prefix), 0043 (accounts as
configuration), 0078 (the broker module is `nats`), 0083 (JetStream carries the guarantee), 0039
(unchanged, and the reason this is possible).
- The lab beds that prove the bus are re-run on NATS; none is skipped.
- Until the cutover, nothing changes on any node.
## How it is checked
A lab bed raises a mesh on NATS from genesis and proves: a node enrols over TLS with a claimed
token; a push composes, is held while the store restarts and applies after; an upgrade rolls out;
a module's tools are invoked from another node and from a person's client; an event dead-letters
after its deliveries are exhausted; a module's account cannot publish outside its `emits` nor
subscribe outside its `consumes`. Then the cutover bed: a mesh on AMQP with the predecessor's
compatibility broker beside it moves its bus in one rollout with every node reporting afterwards.
## References
- [research 014](../01-RESEARCH/014-the-bus-on-nats/00-overview.md)
- [ADR 0002](0002-nodes-communicate-over-a-broker.md), [0033](0033-the-substrate-is-a-store-and-a-broker.md),
[0039](0039-what-the-sdk-holds-and-refuses.md), [0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md),
[0078](0078-the-store-and-broker-are-modules.md), [0083](0083-one-push-leaves-the-mesh-consistent.md)
+73
View File
@@ -84,6 +84,17 @@ python3 00-META/checks/index.py fail if stale
- **0001** — [The mesh brokers capabilities; nodes host; agents think](0001-mesh-brokers-nodes-host-agents-think.md)
- **0002** — [Nodes communicate over a message broker, not over HTTP](0002-nodes-communicate-over-a-broker.md)
- **0003** — [An agent is a persistent employee, not an instance of a pool](0003-agents-are-persistent-employees.md)
- **0077** — [The parts are named controller, foundation, node — not control plane, substrate, master](0077-the-controller-and-the-foundation.md)
- **0083** — [One push leaves the mesh consistent](0083-one-push-leaves-the-mesh-consistent.md)
- **0088** — [The foundation filters before anything listens](0088-the-foundation-filters-before-anything-listens.md)
- **0090** — [A failure that repeats is said to be stuck](0090-a-failure-that-repeats-is-said-to-be-stuck.md)
- **0100** — [A node in use is adopted before it is converged](0100-a-node-in-use-is-adopted-before-it-is-converged.md)
- **0101** — [A machine's own resolver does not make it in use](0101-a-machines-own-resolver-does-not-make-it-in-use.md)
- **0102** — [The mesh writes into a shared file, never over it](0102-the-mesh-writes-into-a-shared-file-never-over-it.md)
- **0103** — [What an adopted node holds, and what its guard refuses](0103-what-an-adopted-node-holds-and-what-its-guard-refuses.md)
- **0104** — [A provision may be answered by an adapter to the predecessor](0104-a-provision-may-be-answered-by-an-adapter-to-the-predecessor.md)
- **0105** — [The mesh adopts the predecessor's tunnel in place](0105-the-mesh-adopts-the-predecessors-tunnel-in-place.md)
- **0106** — [The bus is NATS](0106-the-bus-is-nats.md)
### Its tiers, from the bottom up
@@ -92,11 +103,57 @@ python3 00-META/checks/index.py fail if stale
- **0006** — [The substrate and the control plane](0006-the-substrate-and-the-control-plane.md)
- **0007** — [Connectivity](0007-connectivity.md)
- **0008** — [A context owns its store, exclusively](0008-a-context-owns-its-store.md)
- **0028** — [The substrate supplies the control plane and nothing else](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md)
- **0029** — [A network is a shape, because an action cannot be undone](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md)
- **0030** — [Data outlives the mesh that declared it](0030-data-outlives-the-mesh-that-declared-it.md)
- **0031** — [The control plane authenticates nobody, so identity is a module](0031-the-control-plane-authenticates-nobody.md)
- **0033** — [The substrate is a store and a broker](0033-the-substrate-is-a-store-and-a-broker.md)
- **0036** — [Bootstrap ends at a usable mesh, and the first credential comes from a person](0036-bootstrap-ends-at-a-usable-mesh.md)
- **0066** — [Public routing is name-agnostic, its names are resolved inside the mesh, and an internal authority can certify them](0066-public-routing-is-name-agnostic.md)
- **0067** — [Genesis is a pivot: a temporary control plane installs the registry that makes it permanent](0067-genesis-is-a-pivot.md)
- **0070** — [The catalogue owns the module graph, and genesis builds rather than carries](0070-the-catalogue-owns-the-module-graph.md)
- **0071** — [Genesis clones from a mesh, and checks what it got](0071-where-genesis-gets-its-source.md)
- **0072** — [Two graphs, and a build chain that orders itself](0072-two-graphs-and-the-build-chain.md)
- **0073** — [The installer carries a builder, and the registry stays where it is](0073-the-installer-carries-a-builder.md)
- **0074** — [The mesh defines a module protocol; an SDK is an implementation of it](0074-the-wire-is-specified-not-the-types.md)
- **0075** — [An artifact store is a provision; a package registry is a different one](0075-two-stores-and-which-provides-what.md)
- **0078** — [The store and the broker are ordinary modules](0078-the-store-and-broker-are-modules.md)
- **0079** — [The foundation seats are named after their servers](0079-the-foundation-seats-are-named-after-their-servers.md)
- **0092** — [An operator delivers a pair credential, and the mesh never replaces it](0092-an-operator-delivers-a-pair-credential.md)
- **0094** — [A module may hold several secrets from one provider, each a pair of its own](0094-a-module-may-hold-several-secrets-from-one-provider.md)
- **0095** — [The control plane is the way to ask a module](0095-the-control-plane-is-the-way-to-ask-a-module.md)
- **0098** — [A fact a provider makes at first start is fetched from it, not carried in its manifest](0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md)
### What runs on them, and how it gets there
- **0009** — [Modules and the graph](0009-modules-and-the-graph.md)
- **0010** — [Delivery](0010-delivery.md)
- **0024** — [Model access is a provision, and a licence is a thing with a name](0024-model-access-is-a-provision.md)
- **0026** — [The mesh has a session of its own, and it is the node session's mechanism](0026-the-mesh-has-a-session-of-its-own.md)
- **0027** — [A provision names what the consumer is coupled to, not the role it plays](0027-a-provision-names-what-the-consumer-is-coupled-to.md)
- **0035** — [One implementation, several surfaces, and what that costs](0035-one-implementation-several-surfaces.md)
- **0038** — [The mesh assigns the port, and a module does not care](0038-the-mesh-assigns-the-port.md) *(proposed)*
- **0040** — [What a module is](0040-what-a-module-is.md)
- **0041** — [Events are a relationship, the lighter sibling of provisioning](0041-events-are-a-relationship.md)
- **0042** — [The shape of an event on the wire](0042-the-shape-of-an-event-on-the-wire.md)
- **0043** — [A module's broker account is scoped by what it emits and consumes](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md)
- **0044** — [A public name is provisioned, not registered by hand](0044-a-public-name-is-provisioned-like-any-capability.md)
- **0045** — [A machine's firewall is the sum of what its modules listen on](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)
- **0046** — [A module's configuration is its assignment's, not its manifest's](0046-a-module-configuration-is-its-assignments-not-its-manifest.md)
- **0047** — [A module runs its code as its own process, with its own account](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)
- **0048** — [A provider creates the credential the mesh minted, and seals nothing](0048-a-provider-creates-the-credential-the-mesh-minted.md)
- **0049** — [A consumer's identity is bounded by the tightest backend that must accept it](0049-a-consumers-identity-fits-the-tightest-backend.md)
- **0050** — [Model access is vendor-agnostic, and a vendor is an adapter](0050-model-access-is-vendor-agnostic.md)
- **0051** — [Shared data is the operator's, and a module is granted access to it](0051-shared-data-is-the-operators.md)
- **0052** — [An init step is a container run once to completion, gating what follows](0052-a-step-that-runs-once-before-a-container.md)
- **0053** — [A scheduled step is a container run on a recurring schedule](0053-a-step-that-runs-on-a-schedule.md)
- **0054** — [Model usage is a vendor-neutral record, produced by the adapter, at two grains](0054-model-usage-is-recorded-at-two-grains.md)
- **0055** — [Model access is answered by a licence, or by a node that hosts the model](0055-model-access-is-answered-by-a-licence-or-a-node.md)
- **0084** — [Which provider serves a consumer, when the mesh runs more than one](0084-which-provider-serves-a-consumer.md)
- **0085** — [A secret is a provision, and the vault is the module that provides it](0085-a-secret-is-a-provision.md)
- **0087** — [A seeded file is created once, and what grows in it is not the mesh's](0087-a-seeded-file-is-created-once.md)
- **0091** — [A mount is declared, and there are three things it can be](0091-a-mount-is-declared-three-ways.md)
- **0099** — [A step that runs once names what it reads, and runs again when it changed](0099-a-step-that-runs-once-names-what-it-reads.md)
### How it is built
@@ -106,11 +163,22 @@ python3 00-META/checks/index.py fail if stale
- **0014** — [No workspace — each module is a standalone package consuming published dependencies](0014-no-npm-workspace.md)
- **0015** — [Applications live in their own repository; the monorepo is for the mesh](0015-applications-live-in-their-own-repository.md)
- **0016** — [The lab](0016-the-lab.md)
- **0037** — [Where a module lives](0037-where-a-module-lives.md) *(proposed)*
- **0039** — [What the SDK holds, and what it refuses](0039-what-the-sdk-holds-and-refuses.md)
- **0068** — [The lab takes requests, one at a time, and runs each from its own copy](0068-the-lab-takes-requests.md) *(proposed)*
- **0069** — [A module is a repository and a path within it](0069-a-module-is-a-repository-and-a-path.md)
- **0076** — [The SDK is a published package, and the toolchain resolves it by version](0076-the-sdk-is-a-published-package.md)
- **0082** — [The registry is reached by name, and the overlay is its security](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)
- **0086** — [A secret reaches a process as a file, and an exception is declared](0086-a-secret-reaches-a-process-as-a-file.md)
- **0096** — [An upstream image is copied between registries, never through a machine's image store](0096-an-upstream-image-is-copied-between-registries.md)
- **0097** — [A vendor image is a declared build input, and a recipe fetches nothing undeclared](0097-a-vendor-image-is-a-declared-build-input.md)
### How it is checked
- **0017** — [A test defends a decision](0017-a-test-defends-a-decision.md)
- **0018** — [A picture of a system is read from the system, never from what asked for it](0018-a-picture-is-read-from-what-runs.md)
- **0089** — [A bed reads the catalogue it proves](0089-a-bed-reads-the-catalogue-it-proves.md)
- **0093** — [A fixture that runs a module's runtime carries the module's name](0093-a-fixture-that-runs-a-modules-runtime-carries-its-name.md)
### How we work
@@ -119,5 +187,10 @@ python3 00-META/checks/index.py fail if stale
- **0021** — [HQ is the source of the mesh constitution](0021-hq-is-the-source-of-the-constitution.md)
- **0022** — [The constitution absorbs what is already enforced](0022-the-constitution-absorbs-what-is-enforced.md)
- **0023** — [The approval is the checkpoint, not the second pair of hands](0023-approval-is-the-checkpoint.md)
- **0025** — [The design record is read where it is written, never copied to be found](0025-the-design-record-is-read-not-copied.md)
- **0032** — [The local account owns the mesh; a surface delegates to a module](0032-the-local-account-owns-the-mesh.md) *(superseded)*
- **0034** — [The local account owns the mesh, and a web application's login is not that](0034-the-local-account-owns-the-mesh.md)
- **0080** — [The development cycle is checked, not trusted](0080-the-development-cycle-is-checked.md)
- **0081** — [A decision nothing cites is not yet in the chain](0081-a-decision-nothing-cites-is-not-yet-in-the-chain.md)
<!-- index:end -->
@@ -1,11 +1,12 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
code: [hal, mesh-controller, mesh-catalog, mesh-host]
updated: 2026-09-21
decisions:
- 02-DECISIONS/0011-managed-files-are-generated-never-edited.md
- 02-DECISIONS/0009-modules-and-the-graph.md
- 02-DECISIONS/0085-a-secret-is-a-provision.md
---
# Configuration and secrets
@@ -73,9 +74,31 @@ when the file is created, so regeneration left the previous mode in place. The c
worth remembering beyond the instance: a permission set at creation is not a permission
maintained.
**Rotation is not a mesh operation.** Secrets can be generated and granted; there is no
mechanism that rotates one and informs everything holding it. Where a rotation has been done,
it has been done by hand, and doing it wrong has taken services down.
**Rotation was not a mesh operation** in the mesh being replaced, and doing it by hand took
services down. On the mesh that exists now it is: `rotate <provision>` discards a pair credential
and delivers both ends in one send, and a module's own secret is such a pair credential when the
module takes it from the vault — which, at the time of writing, one module does (redis).
## The vault, as it runs
Since 2026-09-21 the mesh runs `mesh-vault`, a foundation module installed at genesis beside the
adopted store and broker. It provides `secret`: a module that requires one receives a pair
credential the controller minted, and the vault's ledger records the holder and the value's
fingerprint, notices a rotation, and answers over the mesh by fingerprint only. It holds no value.
Genesis makes the store's superuser and the broker's administrator rather than copying the
template's, keeps them at the paths the store and broker modules declare as their own secrets, and
makes an **operator sealing key** before the first secret is accepted: its private half is a file
beside the produced bundle, which the operator carries off the machine, and its public half is
what the mesh records. Every secret a module holds for itself and every pair credential is sealed
to that key as well as to its node. The export of those copies is written beside the key at the
end of genesis and kept by the vault on its own disk; the operator recovers any secret from it, off
the mesh, with `secret recover`. Secrets made before the key existed, or sealed to a replaced key,
are listed as such rather than passed off as recoverable. The produced bundle and the installer's
transcript carry no credential in the clear.
Two things this does not yet do: a module with several own secrets cannot take them all from the
vault (issue 069), and an operator cannot hand a value into a pair credential (issue 070).
## Node-level and mesh-level values
+402 -130
View File
@@ -1,158 +1,430 @@
---
layer: to-be
status: designed
code: [hal]
updated: 2026-08-23
decisions: [02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md]
code: []
updated: 2026-09-01
decisions:
- 02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md
- 02-DECISIONS/0005-the-node-host.md
- 02-DECISIONS/0009-modules-and-the-graph.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md
---
# Work breakdown — the decomposition
# Work breakdown — replacing what provisions the mesh
How [ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) gets built, in what order, and where a human must look.
*Rewritten 2026-08-31. The previous version planned a decomposition of the existing system in
place: extract contexts, convert modules to declared features, shrink its shared library. That is
not what is being done — a replacement is being built beside it, and the old plan's Phase 0 was
the only part that survived contact with it. So the document that was supposed to say what happens
next had been describing work on a system being retired.*
Ordering is not preference. Each phase removes a constraint the next one needs gone.
## The goal, in one sentence
---
**Modules move to the new mesh one at a time, until the old registry can be switched off.**
Everything below is ordered by what that requires. Nothing here is a rewrite of the old system;
its modules are the input.
## Phase 0 — a mesh that runs — **done**
Not *the code exists*. Twenty-two assertions on real machines in the lab, each confirmed to fail
when the behaviour is removed ([ADR 0016](../../02-DECISIONS/0016-the-lab.md),
[ADR 0017](../../02-DECISIONS/0017-a-test-defends-a-decision.md)).
| what is proven | |
|---|---|
| **a mesh comes into being** | a bare machine becomes one; others join with nothing but a token |
| **credentials** | delivered to both ends with the mesh holding neither; rotated so the old one stops working |
| **declarations survive reality** | a stopped machine is waited for; one that fell behind catches up unnamed; unassigning takes away exactly what it should; what the mesh says nothing about is left alone |
| **failure is legible** | a machine that cannot do what it was told is named, with why |
| **the mesh runs itself** | its own artifact store, and a builder that is a module the mesh assigns |
| **names and reachability** | internal names, wildcards under a machine, containers reaching other machines, certificates the mesh issued, filtering that matches exactly what was declared |
| **delivery** | a new commit reaches a machine already running the old one |
| **model access** | answered by a record, with a key the mesh cannot read |
**What Phase 0 does not prove, and it is the important sentence in this document:** every module
exercised above was written to test the mechanism. **No module from the existing system has ever
run on this.** The vocabulary was shaped by the things used to test it — the same fault as a
fixture agreeing with the code it checks
([`04-ISSUES/005`](../../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)), at the
scale of a design.
## Phase 1 — the vocabulary a real module needs
Found by taking real modules and asking what they would require. Each is a gap in what can be
*expressed*, not a defect in what is built.
| # | task | done when |
|---|---|---|
| ~~1.1~~ | ~~An **object-store provision**~~ — **done 2026-08-31**, and it needed no change to the mesh: see below | seven assertions against a real store |
| ~~1.2~~ | ~~**A session as a consumer of a licence**~~ — **done 2026-08-31**, and it also needed no change: see below | two sessions on one machine, different licences, each its own key |
| ~~1.3~~ | ~~A **network** shape, and ordering within a module~~ — **done 2026-08-31** | the shape is created and removed; ordering was already there, and is now asserted |
| 1.4 | **Public certificate issuance** — **built; one gap** | ordering, the challenge and issuance are proven against a real authority; **collecting the issued certificate is not** ([`04-ISSUES/020`](../../04-ISSUES/020-a-certificate-is-issued-and-never-collected/00-report.md)) |
**1.3 and 1.4 block later ones** and are listed now so they are not met as surprises. 1.3 is what
a mail system needs and nothing else so far does.
**Phase 1 is closed with 1.4 partly open**, deliberately. Two of its four items needed no code at
all; the network shape was built; and certificates are configured correctly, order correctly, and
are issued correctly — the client does not collect what the authority issued, against a server
that exists to be a test server. That is filed rather than chased, because the remainder may say
nothing about a real authority and the next thing to learn comes from moving a module rather than
from a fourth lab run.
**Checkpoint:** each is demonstrated in the lab before the module needing it is attempted.
### 1.2, and the same surprise twice
**A binding is per module per machine, and the two sessions are two modules** — the same mechanism
in different context roots, and a context root is what a module delivers. So `(node, module)`
already names them apart, and nothing needed adding.
[`14-model-access.md`](14-model-access.md) had called per-module-per-machine *a step toward it and
not it*, which is true of a **worker** — many run on one machine from one module — and not true of
a session, of which there is one per node and one for the mesh.
### 1.3, and the first one that needed building
**Ordering was already there** — the apply loop sorts nothing, so a module says *this before that*
by writing it first. Untested until now, and the kind of property a later change breaks silently.
Worth separating from readiness: a container started is not a container ready, and nothing waits.
What needs something *usable* retries, which is what both provisioners do and is the better answer
anyway, because a dependency can restart long after everything was applied.
**The network was a real gap, and the first thing in Phase 1 that needed a decision.** Adding a
shape widens what a compromised controller can express, so
[ADR 0029](../../02-DECISIONS/0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md)
records why this one is worth it: an `action` could create a network and **nothing could ever
remove it**, because an action leaves no footprint the host can undo. The vocabulary is nine.
**Three tasks in a row that were already possible.** Both were written from the design rather than
from the code, which is the review's finding arriving in the plan: *a claim here is counted, not
reasoned.* The remaining Phase 1 items should be checked against the code before being started,
not after.
### 1.1, and what it turned out to be
*Done 2026-08-31. Worth recording because the task was not the one written down.*
**The controller special-cases nothing.** `provides`, `requires`, `contributes` and `grants`
are name-agnostic — asking for a bucket needed no change to the mesh at all. What was missing was
a provider, and the last step where something on the machine turns a delivered secret into a key
that works. So "add an object-store provision" was never mesh work.
The provision is `s3-bucket`: a consumer's code is written against the S3 API and swapping one
store for another does not break it, so by
[ADR 0027](../../02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md) the name
says the protocol. A database is the other case, and names the engine.
**One assertion here that a database does not need.** One PostgreSQL server holds separate
databases and the product enforces the boundary; one object store holds every bucket behind one
endpoint, so *a consumer cannot reach another consumer's bucket* is a policy somebody wrote — and
a policy granting everything would pass every other test. **What is asserted is what the policy
does not say.**
## Data is the constraint, and it outranks the order below
*2026-08-31.* The modules being converted run live services — identity, mail — and **the data must
survive every step**. A data folder may move; it may never be lost.
**One thing was found by asking this and is fixed**
([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)): the host deleted
a directory and everything under it when the directory stopped being declared, which happens when
a module is unassigned or a manifest is edited to move a data folder — the exact operation this
plan needs. A directory holding anything the mesh did not put there is now kept and reported.
**That is not a backup and must not be read as one.** It stops the mesh destroying data. It does
nothing about a disk, a mistaken command, or a service corrupting its own store.
**So the rule for every step below:** the data is copied, the copy is verified by reading it back
through the service that owns it, and only then does anything point at the new location. Never
moved and then checked. **A backup nobody has restored is a belief, not a copy.**
## A module is adopted with the credentials it already has
*2026-08-31.* **Nothing is rotated during the conversion.** A service being adopted keeps the
password it is already using, because minting a new one is how a running service stops being able
to reach its own database in the middle of a migration.
The mesh has both paths and this needs the second:
| | |
|---|---|
| **generate** | a new secret, sealed to both ends. What a *new* module gets |
| **accept** | a value supplied from outside, sealed, plaintext discarded. **What an adopted module gets** |
**Rotation is a separate act, afterwards, once everything works.** The machinery for it is built
and proven — a credential moving at both ends with the old one ceasing to work — and it is exactly
the sort of thing to do deliberately on a quiet afternoon rather than as a side effect of moving a
service between systems.
**So there is a step before any of this: read the current environment out of the old system**, because
adoption means supplying those values and they live in its files today.
**And there is a failure worse than losing data, which is likelier.** A database image consumes its
password environment variable **only when its data directory is empty**. Everything here keeps its
data on a persistent directory, so the role holds whatever password it was created with, for ever.
Regenerate that variable and the application moves on while the database does not — permanently,
because nothing reconciles it. Eight modules are in that state today, working only because nobody
has regenerated their credential since their data directory was created.
*Where the detail lives:* this is operational and names machines, so it is in the mesh's own
knowledge base rather than here — `migration/where-service-data-lives`, which surveys where every
service's data actually sits and what each stop or removal would cost, and
`troubleshooting/db-password-frozen-at-first-init` for the lockout itself. **This document says the
rule; those say the specifics.**
*Corrected 2026-08-31 — an earlier version of this paragraph made that sound more dangerous than it
is.* A sealed secret is not unreadable; it is sealed **to the node**, which holds the private half
and writes the plaintext into the module's own file. The value is there, on the machine, as an
ordinary file. What does not exist is a way to ask *the mesh* what a secret is, and there is no
reveal command, because a mesh that can reveal a secret is a mesh that holds one.
## Where it starts, and what that costs
**On the node holding all the production data**, because that is where the services being
converted actually are.
Recorded plainly rather than argued with: this is the highest-risk order available. Everything
proven so far was proven on machines that could be destroyed and raised again, and the first real
exercise of the conversion will be on the one machine where a mistake is not recoverable. Nothing
about the lab work transfers automatically — a scenario proves the mechanism, not the state on
that machine.
**What makes it survivable is preparation rather than caution**: a restored backup before the
first step, one service at a time, and the previous arrangement left standing until the new one
has been read back. None of that is slower than the alternative, because the alternative includes
losing something.
## How the two systems hand over
*2026-08-31.* **The old system's brain is switched off; its services keep running.**
Not a migration and not a period of dual control. The old controller — provisioning, the
coordinator, the pipeline, the things that *decide* and *write* — is stopped. Every workload it
was managing goes on running exactly as it is, because nothing is managing it. Then the new mesh
takes ownership of them one at a time.
**Nothing is ever unassigned in the old system.** Unassigning is how it removes things, and
removing is how data is lost. The old system is never asked to take anything away; it is asked to
stop having opinions.
| | |
|---|---|
| **stopped, and disabled** | provisioning, the coordinator, environment and configuration sync, the pipeline — anything that decides or writes a file |
| **left alone entirely** | the units running the actual services: identity, mail, databases, the forge. They keep serving throughout |
| **never used** | unassign, remove, delete — any operation whose job is to take something away |
**Disabled, not merely stopped**, and this is the part that is easy to get wrong: those units are
enabled, so stopping them lasts until the machine reboots. A reboot mid-conversion would bring the
old controller back and it would resume regenerating managed files underneath the new one —
which is the one situation where two systems really would be fighting over the same machine.
**A service left running with nothing managing it is the safe state.** It has its data, its
configuration is already on disk, and nothing is going to change either. That is the whole trick:
the risk in a conversion is in the *managing*, not in the *running*.
**A brief interruption is acceptable. Losing data is not.** Where those two trade against each
other, the interruption wins every time — a service can be restarted, and there is no operation
that un-deletes a mail spool.
**The new host cannot remove what it did not put there.** Orphans are per-origin, so it only ever
removes resources it recorded itself. Services it has never been told about are not orphans to
it — they are simply not its business, which is what makes taking ownership one module at a time
safe.
## Phase 2 — the first real module
| # | task | done when |
|---|---|---|
| 2.1 | Port an **object store** module | it runs on the new mesh, serves a bucket to another module, and its credential rotates |
| 2.2 | Copy the data, and read it back through the service that owns it | the new location answers with what the old one holds |
| 2.3 | Point one dependent at it, old arrangement left standing | something real reads and writes through the new mesh's copy |
**Checkpoint, and it is a human one:** it runs for a week before anything else moves. The point of
going first is to find what Phase 1 missed, and a week is roughly how long that takes to show.
## Phase 3 — the modules that prove the shape
Each exercises something the first one does not.
| # | task | proves |
|---|---|---|
| 3.1 | An **identity provider** | a module that is itself a provider — the provides/requires chain, with consumers requiring it |
| 3.2 | A **forge** | a port claim against the machine's own daemon, and a module wanting both a database and an object store |
| 3.3 | A **mail system** | several containers as one module, a private network between them, and names that are not one-per-node |
**3.3 is the hardest thing in this document** and is deliberately last. If the declaration
language turns out to be insufficient, it says so here.
### Where Phase 3 actually stands — *2026-09-01*
All three have manifests. All three parse, resolve and plan. **None of them can start**, and the
two reasons are both filed rather than guessed at.
The **vocabulary held**. Nothing in 3.1–3.3 turned out to need a new shape: the identity provider,
the forge and the mail system are all expressible with what exists, including the mail system's
several containers on a private network — which was the one expected to break it. That is the
question this phase was designed to answer, and the answer is yes.
What did not hold was underneath the vocabulary:
- **[`022`](../../04-ISSUES/022-one-credential-per-node-per-provision-not-per-module/00-report.md)**
(fixed) — a credential belonged to a machine, so a node running several modules against one
database could not be planned. The refusal was loud on the provider and silent on the consumer.
- **[`023`](../../04-ISSUES/023-a-consumer-cannot-build-a-connection-string/00-report.md)** (open)
— a consumer gets its password and still cannot connect: the user name is invented by the
provisioner and recorded nowhere, and the bound values cannot reach a configuration file.
A third fault was in the manifests themselves rather than the design: each declared a secret at a
path named `.env` and read it as one, when a sealed file holds a password and nothing else. They
parsed and resolved and could never have worked, which is what a manifest checked only by a parser
buys. Two tests now refuse both halves of it.
### 3.1 needs a program, not a decision — *2026-09-01*
023 is fixed, and with it the two design faults are gone. What stands between 3.1 and a running
identity provider is now one concrete thing: **the realm provisioner does not exist.**
Its manifest named an image — `mesh-provision-keycloak` — that nothing builds and no program
backs. That has been removed rather than left standing, because a manifest describing a program
nobody wrote is the same mistake as the credential files that could never be read: it parses, it
resolves, and it could never work.
So Keycloak's manifest now says what is true today — a server the mesh runs, with its database
and its admin credential, both reaching it in a shape it can read. It no longer claims to provide
`oidc-client`, which means a consumer asking for one is **refused by name at plan time** rather
than resolving cleanly and waiting for a client nothing will create.
The provisioner is the same shape as the two that exist: it reads what the mesh granted and
reconciles a realm and a client per consumer. **It should be written against a real Keycloak in
the lab**, not from the API documentation — the object store's took three corrections that only a
running server produced.
The forge (3.2) and the mail system (3.3) need no provisioner and are not blocked on this.
### And they could not have run anyway — *2026-09-01*
Every one of the five named a container image that does not exist: sixty-four zeros where a digest
belongs, eighteen times over
([`025`](../../04-ISSUES/025-a-module-must-pin-a-digest-and-nothing-produces-one/00-report.md)).
They parsed, resolved and composed into a declaration a host accepts, and every one would have
stopped on the machine at the moment of fetching.
Nothing caught it because nothing could. A host checks the *shape* of a reference and no more —
verifying a digest exists means reaching a registry, which is the one thing a host must never have
to do. The refusal now sits where a declaration is composed instead, which is the last moment
before a machine sees one.
Twelve are pinned to real images. Two further faults surfaced only by pinning for real: the mail
system's seven images named repositories that **do not exist**, because it publishes to a
different registry than assumed, and one of the seven had been renamed upstream.
**The forge now runs**, on a database another module provides, with a password it did not choose
and a connection string it could not have written. That is the first of these descriptions to be
started rather than planned, and it exercises everything the credential work added.
What is still missing is the mechanism: nothing turns a tag into a digest as part of the mesh's
own work, so it was done by hand. Asking a registry takes about a second and pulls nothing, which
removes the main argument for leaving it undone.
## The conversion is done by hand, and that is a decision
*2026-08-31.* **Moving from the current system to this one is a person at a command line, working
through it.** Not a migration program, not a converter, not a period of dual-writing.
**What that removes from this plan is larger than what it adds.** Nothing below needs an importer,
a translation layer, a compatibility shim, or a mechanism for keeping two systems agreeing while
both are live — and every one of those is a thing somebody would otherwise reasonably build, use
once, and maintain for a year. The modules are the input; a person reads what one does today and
writes what it declares tomorrow.
**It also changes what "safe" means for the system being retired.** A fix to it has to be safe on
its own, because there is no careful rollout to sequence it into: the thing is being switched off
by hand, not managed into retirement. A change needing three steps in the right order is a change
that will be half-applied.
**And it is why the checkpoints below are weeks rather than gates.** Nothing enforces the order —
a person does — so the value of the sequence is entirely in what each step teaches before the next
one starts.
## Phase 4 — switch the old registry off
| # | task | done when |
|---|---|---|
| 4.1 | Move the remainder, by hand, a module at a time | nothing is assigned in the old system that is not assigned in the new one |
| 4.2 | The old one authoritative for nothing | a change to any module goes through the new mesh only |
| 4.3 | Switch it off | it is stopped, and nothing notices |
**4.3 is a day's work and the phases above it are not.** Naming it as a phase is what stops it
being mistaken for the goal.
## Sequencing
- **1 before 2.** Attempting a module without the vocabulary it needs produces a workaround, and a
workaround in a manifest is a design decision taken by whoever was in a hurry.
- **2 before 3, with the week.** Moving three modules before running one is how three modules
acquire the same defect.
- **3.3 last.** It is the only one that may send work back into the declaration language.
- **4 cannot start early, and there is no partial credit.** A registry still authoritative for one
module is still running.
## How this list is kept true
*This section exists because the document it replaces was wrong for weeks and nothing said so.*
**A claim here is counted, not reasoned.** The review of 2026-08-31 found a bundle described as
carrying two images that carries three, a bootstrap described as needing six shapes that uses
four, and ten documents calling themselves `designed` while naming lab-proven code. Each was
produced by describing the system from its design instead of reading it.
**A phase is done when the lab says so**, and the lab keeps a receipt of when it last ran and
against which commits. A phase marked done here whose assertions have not run is a claim about the
past.
**What is not proven gets said.** Phase 0 is done and its limitation is written into it. A list
that records only progress becomes a list nobody believes.
## Rules of engagement
These exist so the work can run largely unattended without accumulating the kind of
damage this refactor is meant to remove.
Unchanged from the previous version: they were about how work is done rather than what the work
is.
### Autonomous by default
An agent may, without asking:
- read anything, measure anything, query any database read-only
- create branches, write code and tests, open pull requests
- run the test suite and typechecks
- write and update `hq/` documents
Read anything, measure anything, query read-only. Create branches, write code and tests, run the
suites, and write or update documents here.
### Always stop and ask
- **destroying or overwriting data** — dropping a table, deleting a provision, rotating a
live credential, removing a module from a node
- **destroying or overwriting data** — dropping a table, deleting a provision, rotating a live
credential, removing a module from a node
- **merging anything** — every merge is a human checkpoint, without exception
- **a decision the ADRs do not already answer** — record the question in the relevant
research effort rather than picking and moving on
- **any change to `hq/00-META`** — it is stable by nature
- **anything touching a machine outside the lab**, including a configuration change that restarts
something people are using
- **a decision the records do not already answer** — record the question rather than picking and
moving on
- **any change to [`00-META`](../../00-META/)** — it is stable by nature
### Definition of done for every task
1. tests written **and failing first**, then passing
2. typecheck clean in every package the change touches
3. the local mesh (Phase 0) comes up, and the behaviour is demonstrated in it
4. `hq/` updated if the task changed or answered anything documented
5. deployed, and **delivery verified on every node** — not "the pipeline was green"
3. the behaviour demonstrated **in the lab, on real machines** — not asserted
4. documents here updated if the task changed or answered anything recorded
5. delivered, and the **effect** verified — not that a pipeline was green
### Non-negotiables carried from the current system
### Non-negotiables
- **Never edit mesh-managed files on disk.** Use the owning tool.
- **Never write to production databases directly.** Migrations for schema, application
code for data.
- **Every schema change ships twice** — consolidated schema *and* an incremental
migration.
- **Expand, then contract.** Add the new shape, migrate, verify, and only then remove the
old one — never in a single step.
- **A green pipeline proves transport, not effect.** Verify the effect.
---
## Phase 0 — A mesh that runs locally *(prerequisite)*
Nothing else starts until this exists. Every fault this refactor addresses was found in
production because there was nowhere else to find it.
| # | task | done when |
|---|---|---|
| 0.1 | Container image for a node runtime | a node process starts in a container and registers |
| 0.2 | Compose topology: broker, registry DB, object store, *n* nodes | `up` yields a mesh that elects a provider node and settles |
| 0.3 | Seed a minimal mesh: nodes, one module, one provision | a module deploys end-to-end with no external service |
| 0.4 | Run the pipeline inside it | a push-equivalent produces a cascade and a deployed artifact |
| 0.5 | Fixtures for the failure modes already known | credential rotation reaching a running session; a provider deploy rotating a shared credential; a migration that ships nothing — each reproducible on demand |
**Checkpoint:** a human confirms the local mesh reproduces at least one bug from
2026-08-22 before any decomposition begins.
---
## Phase 1 — Make the model expressible
The decomposition is impossible while a feature is a singleton per module.
| # | task | done when |
|---|---|---|
| 1.1 | Decision record — named features, per-node opt-in (next free number) | accepted |
| 1.2 | Manifest: declared `features:` with type + directory | a module declares two of one kind and both build |
| 1.3 | Selection: `always` / `optional` | a node installs a subset; artifacts stay selection-blind |
| 1.4 | `requires:` moves onto the feature | a schema feature's database is not provisioned where the feature is not installed |
| 1.5 | Assignment carries the opted-in feature set | opting a node in requires no rebuild |
**Checkpoint:** one existing module converted to declared features, deployed, verified —
before any others follow.
---
## Phase 2 — Draw the boundary the domain already has
Cheapest first, and each one proves the extraction pattern before the expensive ones.
| # | task | extracted from | risk |
|---|---|---|---|
| 2.1 | `hal/knowledge` — one store, review workflow ported | hippocampus + noxflow `knowledge_*` | low — additive |
| 2.2 | `hal/stream` — the record; notifications and messaging as views | axon, synapse, notifications, meetings, conversations | medium |
| 2.3 | `hal/agents` — identity, licence, runs, memory, thoughts | noxflow agents, `hal/thoughts` | **high** — touches credentials |
| 2.4 | `hal/work` — what remains of noxflow | noxflow tasks | medium |
| 2.5 | `hal/ai` — one module per provider, each providing *a model provider* | `hal/claude*` | medium |
Each extraction is expand-then-contract: new context alongside, dual-write, verify, cut
over, remove. **Never a move commit.**
**Checkpoint:** after 2.1, a human confirms the extraction pattern before 2.2 begins.
After 2.3, a human confirms credentials still reach every agent on every node.
---
## Phase 3 — Reclaim the kernel
Only possible once domains have modules to own their code.
| # | task | done when |
|---|---|---|
| 3.1 | Move work-domain code out of `hal/sdk` | `workflow-engine.ts`, `task-commands.ts` live in `hal/work` |
| 3.2 | Move provider code out | `claude-credentials.ts` lives in `hal/ai` |
| 3.3 | Move delivery code out | feature handlers, artifact manager, build executor live in `hal/delivery` |
| 3.4 | Decide the residue | ADR: what `hal/sdk` keeps (open question 4) |
**Measure:** `hal/sdk` line count, tracked per task. Today: **34,636** across **155**
files.
---
## Phase 4 — Separate what the mesh runs from the mesh
| # | task | done when |
|---|---|---|
| 4.1 | Decide the destination (open question 3) | ADR accepted |
| 4.2 | Cross-repository dependency resolution proven | a catalogue module builds against a published `@hal/*` |
| 4.3 | Move the 91 catalogue modules | this repository contains only mesh contexts |
**Checkpoint:** move one application first and run it for a week before the rest follow.
---
## Sequencing constraints
- **0 before everything.** Unverifiable refactors are how this list got long.
- **1 before 2.** Extracting into contexts without per-node features recreates the module
count inside the new names.
- **2 before 3.** A domain can only own its shared code once the domain has a module.
- **2.3 after 2.1 and 2.2.** Agents touch credentials; do it once the pattern is proven on
cheaper contexts.
- **4 last.** It is the only phase that is pure movement, so it is the only one safe to
defer indefinitely.
- **Never edit mesh-managed files on disk.** Use the thing that owns the file.
- **Never write to a production database directly.** Migrations for schema, application code for
data.
- **Every schema change ships twice** — consolidated schema *and* an incremental migration.
- **Expand, then contract.** Add the new shape, migrate, verify, and only then remove the old one.
- **A green pipeline proves transport, not effect.**
## What "done" looks like
Eight contexts. `hal/sdk` holding only what is genuinely cross-cutting. A mesh that stands
up on a laptop. A module count that grows only when the domain does.
The old registry is off. Every module runs on the new mesh, declared rather than scripted. A
machine that fails says what it could not do. And the number of modules grows when the work does,
not when the platform needs somewhere to put something.
+79 -7
View File
@@ -2,11 +2,12 @@
layer: to-be
status: in-progress
code: [mesh-lab]
updated: 2026-08-23
updated: 2026-09-21
decisions:
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0093-a-fixture-that-runs-a-modules-runtime-carries-its-name.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0019-how-this-repository-works.md
- 02-DECISIONS/0089-a-bed-reads-the-catalogue-it-proves.md
---
# End-to-end testing
@@ -41,13 +42,13 @@ not the first one built** ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
| | **Bootstrap scenario** | **Full scenario** |
|---|---|---|
| Contains | virtual machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules |
| Contains | virtual machines, the host binary, a pinned foundation bundle | a complete mesh: forge, coordinator, delivery, modules |
| Verdict from | what the host reports about the state it reconciled | a pipeline result ending in verify |
| Exercises | the node host and the substrate | the control plane and everything above it |
| Exercises | the node host and the foundation | the controller and everything above it |
| Exists to | **develop the mesh** | **test what runs on it** |
The bootstrap scenario is a **strict subset**: same virtualisation, same networking, same
lifecycle — it simply stops before a control plane exists. Everything from *"Where this sits in
lifecycle — it simply stops before a controller exists. Everything from *"Where this sits in
the way work happens"* onward describes the full scenario, and applies once there is a
coordinator to describe.
@@ -337,11 +338,14 @@ credential rotation reaches every consumer, that delivery to an absent node is r
pending rather than done, that a returning node catches up. These are fewer and change
rarely, but they are where the known production faults get encoded so they stay fixed.
The known faults become mesh tests that fail today. That is the Phase 0 checkpoint.
The known faults become mesh tests that fail today. Phase 0 of
[`00-work-breakdown.md`](00-work-breakdown.md) is now complete on this basis — twenty-two
assertions on real machines — and what it does **not** cover is recorded there: no module
from the existing system has run against any of it yet.
---
## The substrate
## The foundation
### A node is a system container
@@ -390,6 +394,74 @@ Everything a node itself does is real, because a node is a real machine.
---
## A suite too expensive to run on every push says when it last ran
*Written 2026-08-31, from resolving [04-ISSUES/005](../../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md).*
This suite needs a machine with a hypervisor. It therefore cannot run on every push, and a suite
that does not run on every push runs **when somebody remembers**. Remembering is not a mechanism,
and the harness this one replaces proves it: it had not built for two and a half months, nothing
said so, and the coverage was assumed rather than checked.
**The danger is not that the suite breaks. It is that nobody notices it stopped running** — and
that danger belongs to *this* design, not to the harness it retired.
So three rules, each held by a test:
**A run leaves a receipt** — when, what passed, what it ran, and the commit each repository was
at. Kept **outside version control**: the question is *has this machine run it*, and a receipt in
git would be a claim about everybody's machine made by whoever committed last.
**A receipt says why it does not count.** Old, failed, taken against commits the repositories have
moved past, or a run that never raised a machine. Something can be asked, and answers non-zero.
**A receipt that says nothing about something is not a receipt that clears it** — including a
receipt written before it recorded a given fact, which claims nothing rather than everything.
**The run rebuilds what it tests.** The suite consumes artifacts from other repositories, and an
artifact rebuilt from memory is one rebuilt sometimes. A stale binary reporting success against
rules that have since changed is the same fault wearing different clothes. This covers the module
runtimes a bed's scenario stocks as well as the host and the control plane
([issue 075](../../04-ISSUES/075-a-stocked-runtime-image-is-never-rebuilt-by-the-run/00-report.md)):
each is compared against the module's source and what it is built on, and rebuilt where older,
missing or uncommitted. *How it is checked:* unit tests on what a bed stocks and when it is stale;
a run with an image removed rebuilds it before the bed passes.
**The general rule, which outlives this suite:** *silence and success must never look alike.*
It is the same rule the host follows about a service that does not exist
([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)) — absence must be distinguishable
from a failure to answer — applied to coverage instead of to a machine.
## A bed reads the catalogue it proves
*Written 2026-09-21, from resolving [04-ISSUES/073](../../04-ISSUES/073-beds-carry-copies-of-catalogue-manifests/00-report.md);
decided in [ADR 0089](../../02-DECISIONS/0089-a-bed-reads-the-catalogue-it-proves.md).*
The rule above — the run rebuilds what it tests — was held for binaries and images and not for
manifests. Beds built the manifests they install inline, as literals copied from the catalogue
when each bed was written; the copies did not move when the catalogue did, and a catalogue change
was proven by no bed at all, while every bed stayed green against its copy.
**A bed that installs a catalogue module reads that module's manifest from the catalogue the run
was pointed at.** It rewrites what the lab must — a build artifact becomes the image the machine
holds, an image is pinned, a host port is remapped where one machine carries colliding modules,
an address may point at a stand-in the bed raises — and nothing else.
**A bed that needs less than the module declares is not testing that module.** No upstream
server, a secret in the environment, a requirement edge cut so no second provider is needed:
that is a mesh test. It does not get a name of its own if it runs the module's runtime — a
module's name is its tool namespace and its broker scope
([ADR 0093](../../02-DECISIONS/0093-a-fixture-that-runs-a-modules-runtime-carries-its-name.md)) —
so it reads the catalogue and installs what the module requires; a name of its own is for a
fixture that runs no real runtime.
**The receipt names the catalogue's commit** with the others', so a run taken before a manifest
changed says so — the same rule as for the binaries, for the same reason.
*How it is checked:* a unit test in the lab refuses an inline manifest literal that names a
catalogue module unless the bed is declared, with its reason, in the test's own list, and refuses
a declaration for a copy that is gone; the receipt test asserts the catalogue is claimed whenever
the run is pointed at one.
## Consequences
**Bringing a node into being is part of the framework.** A test creates its own nodes — one
@@ -2,7 +2,7 @@
layer: to-be
status: in-progress
code: [mesh-lab]
updated: 2026-08-25
updated: 2026-08-28
decisions:
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
@@ -205,7 +205,7 @@ machines:
place:
all: [host]
anchor: [substrate]
anchor: [foundation]
snapshot: raised
```
@@ -334,12 +334,12 @@ than a fork.
# bootstrap — tiers 0 and 1
place:
all: [host]
anchor: [substrate]
anchor: [foundation]
# full — adds a control plane, a forge, and a module under test
# full — adds a controller, a forge, and a module under test
place:
all: [host]
anchor: [substrate, control, forge]
anchor: [foundation, control, forge]
module: a-web-service
assert:
- the service answers on its published name
@@ -524,7 +524,7 @@ machines:
place:
all: [host]
anchor: [substrate]
anchor: [foundation]
snapshot: raised
```
+1 -1
View File
@@ -2,7 +2,7 @@
layer: to-be
status: in-progress
code: [mesh-lab]
updated: 2026-08-25
updated: 2026-08-28
decisions:
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md

Some files were not shown because too many files have changed in this diff Show More