204 Commits
Author SHA1 Message Date
mesh-admin e6402cf777 Merge pull request 'Issue 132: a module can be recorded without the directory it lives in' (#157) from issue/132-a-module-can-be-recorded-without-its-directory into main 2026-09-28 07:20:07 +00:00
jschoubben 4f0d144833 Issue 132: a module can be recorded without the directory it lives in
Nine modules could not be rebuilt: their record named the repository and no directory, so every
build looked for a manifest at a repository root that has never had one. Resolved by mesh-controller
— `module add` takes the directory and the forge, and the rule is checked rather than described.
2026-09-28 09:20:05 +02:00
mesh-admin 98d94ef71e Merge pull request 'Design 28: 5.5 done, the mesh has one bus; issue 131 resolved' (#156) from design/28-one-bus-issue-131-resolved into main 2026-09-28 01:59:41 +00:00
jschoubben 4e13280604 Design 28: 5.5 done, the mesh has one bus; issue 131 resolved
The AMQP transport is gone from the control plane and the hosts (mesh-controller #112,
mesh-host #39). On the way: no build had ever recorded its bases, so every bases-first order
walked an empty graph; the builder now reports what it was handed and the graph is read from
builds (mesh-controller #113/#114). Issue 131 is resolved by the forge module's merge event,
the control plane following it, and those edges.
2026-09-28 03:59:39 +02:00
mesh-admin 8783a13448 Merge pull request 'Design 28: the mesh runs on the new bus' (#155) from design/28-the-mesh-runs-on-nats into main 2026-09-28 00:40:29 +00:00
jschoubben a31cfcf461 Design 28: 5.3 is built and was used for the hand-over 2026-09-28 02:28:06 +02:00
jschoubben 75b3861911 Design 28: the mesh runs on the new bus
Tasks 4.3, 5.2 and 5.4 are done as of 2026-09-28 02:25: every machine reports on
the new bus, the seat is held by the module that provides it, the old broker is
unassigned and forgotten, and every credential was minted afresh at the end.

5.2 records what it took, in the order it was found and each fixed on the trunk
before the next step, and how the bootstrap loop was broken once, by hand.
2026-09-28 02:27:49 +02:00
jschoubben 694555214a Merge pull request 'Design 26: which assignment holds a seat is on record, and changes as one act' (#153) from design/26-a-seat-is-held-on-record into main 2026-09-27 21:22:56 +00:00
jschoubben a7249541df Design 26: which assignment holds a seat is on record, and changes as one act
Until now the holder was derived — assigned and claiming — and a second eligible
assignment was refused, so a seat could not pass from one holder to the next
without a moment where nobody held it. The controller finds its own bus through
one of these seats, and that moment took the control plane down on 2026-09-27.

The holder is now a row the controller keeps, written by `seat <name> --to
<node>/<module>` in the same write that removes the previous one. No row means the
old rule, so nothing changes for a mesh that never hands a seat over; with a row,
another eligible assignment is silent rather than refused, which is what lets the
next holder run beside the current one until the switch. A holding is the
assignment's and goes when it does. Each rule names the test that checks it.

Under ADR 0131; design 28 task 5.3 is the work.
2026-09-27 23:20:36 +02:00
jschoubben 95d8253f71 Merge pull request 'ADR 0131: everything on the mesh speaks to the broker seat, and AMQP is not a provision' (#152) from decision/0131-everything-speaks-to-the-broker-seat into main 2026-09-27 21:06:02 +00:00
jschoubben 784b487bf9 ADR 0131: everything on the mesh speaks to the broker seat, and AMQP is not a provision
Taken during the outage of 2026-09-27, when the protocol leaked into the seat's
contract: to hold mesh-broker a module had to provide amqp, so the module that
will carry the bus could not hold the seat that names the bus, while the module
being retired could. Supersedes 0127. Modules depend on the seat and reach the
bus through the sdk; no manifest provides or requires amqp; the old broker's
module and the two modules that required it leave the catalogue; the AMQP
transport is deleted once every node reports on the new bus.

Design 28 step 5 rewritten under it: the seat handover becomes its own task and
is built first, because the seat the control plane dereferences cannot be empty
in between — that emptiness was the outage. The cost note now carries what was
measured rather than what was assumed.

0128 and 0130 extended 0127; each now rests on 0131 with a dated note and
changes nothing it decided. Every other citation of 0127 names its replacement.
records.py still fails on 0120/0112, which predates this branch.
2026-09-27 23:03:05 +02:00
jschoubben 7ae711ba0b Merge pull request 'Building the bus: the decisions the work needed, and what it taught back' (#150) from feat/nats-genesis into main 2026-09-27 17:06:40 +00:00
jschoubben 51739c4302 The order the repositories land in is part of the rollout
Derived while merging and visible from no single repository, so it belongs
written down rather than re-derived later: the client library before the
catalogue, because it is where the subject is derived and a converted module
against the old one publishes the local name itself; the catalogue before the
controller, because the controller refuses an old-style name outright and would
make every unconverted module unregisterable.

Two tested properties are what make it safe and not merely ordered. An old-style
name passes through the derivation untouched, so an unconverted module keeps
working at every step. And a converted name derives to exactly the key the old bus
published, so nothing moves on the wire until 5.2 sets the variable.

The failure avoided is 127's own, which is why this is worth a table: a publisher
and a subscriber disagreeing about a subject log nothing anywhere.
2026-09-27 19:04:45 +02:00
jschoubben 5cac268457 Two checks were wrong about where a seat is judged
Correcting design 26 to match what the merged code does, and fixing issue 112's
status, which used a word the vocabulary does not have.

A claim on a seat the manifest does not itself declare is refused at registration
rather than by the parser. A module may hold a seat another module declared —
that is why ADR 0126 has a caller name the seat and not its provider — so whether
the name exists is a fact about the whole catalogue.

And a declared seat may promise nothing. That is a marker seat, and most
node-scoped seats are markers: which module is this machine's packet filter. ADR
0126's "a declared seat carries a protocol" says what a holder must satisfy, not
that every seat offers something.

`records.py` still fails on ADR 0120 resting on a proposed ADR 0112, which is not
this branch's and not mine to decide.
2026-09-27 18:57:49 +02:00
jschoubben ce6ae943b7 Merge main: renumber this branch's records around the trunk's
Both lines of work numbered from the same point, so four decision records and one design
document existed twice with different content. The trunk keeps its numbers and this branch
yields — the only rule that scales, because the trunk's are already cited by what merged
before them.

  0117 the bus is the only broker        -> 0125
  0118 a module declares its own seats   -> 0126
  0119 amqp is a provision, not the bus  -> 0127
  0120 the mesh bus is required          -> 0128
  0123 a seat carries its role's protocol -> 0129
  0124 the predecessor is ending          -> 0130
  design 29, what a module declares       -> design 32

Applied to the code repositories too, because a stale reference is worse when numbers
collide than when they dangle: the reader lands on a real record that decided something
else.

Two reconciliations the merge forced, both real:

**0110 was marked wholly superseded and was not.** Its successor says in as many words that
everything 0110 decided about what a seat *is* stands untouched — and two records that
landed on the trunk rest on exactly that part. So it is accepted again, extended rather than
replaced, with a note saying which of its claims moved and where.

**A seat's protocol becomes columns, not fields.** The trunk moved the seat set out of
compiled code into a table the controller owns. This branch had added what a role accepts,
emits and serves to the Go slice. The decision is unaffected and the mechanism is better for
it: giving a role a protocol is now a write rather than a rebuild, which is the trunk's own
argument applied to what this branch added.

One check still fails and it fails on main too: a record resting on ADR 0112 while that is
still 'proposed'. Left alone — it is not this merge's to answer.
2026-09-27 18:23:41 +02:00
jschoubben 9ef9830dcf ADR 0122: the predecessor is ending, and its broker goes with it
ADR 0119 rejected giving the old broker a retirement condition and said why: "its
clients are not only the predecessor's, so the retirement condition describes a day
that will not come". The operator has said that day is coming — the predecessor is
deprecated, some of it still running, none of it being migrated, left to stop rather
than moved.

Recorded because three documents reason from the premise it overturns. Design 25 §5's
"no day anything is waiting for", §9's "the predecessor's clients never notice", and
design 28's closing note that the predecessor's world does not need to move.

**And it needs no new machinery, which is 0119 being paid off rather than revised.**
Because that record made the broker an ordinary provider rather than a compatibility
module, ending it is unassigning a provider whose provision nothing requires — something
the module system has done since it existed. So step 5.3 finishes instead of trailing
off, and the transitional double announcement of a build outcome has a date.

The consequence worth planning around: the predecessor's own mesh talks over that
broker, so shutting it down ends the tooling that reaches this installation's machines
from a workstation. The rollout is driven from the node, or driven before the broker
stops. That is a sequencing constraint on 5.2, not an afterthought.

What survives is `amqp` as a provision: a module that genuinely needs an AMQP broker can
still be given one. What retires is this broker's role as the predecessor's.
2026-09-27 18:11:11 +02:00
jschoubben dfac01a6dd 5.2: the readiness half is in, and what a failed move actually costs
`rollout check` answers from records whether this mesh could move its bus, and names the
next step for each thing missing. The move itself waits on that check having been run
against a real mesh — writing the irreversible half before its question has ever been
asked of something real breaks the plan's own rule about beds by another route.

And the cost of being wrong is written down rather than assumed: the old broker stays for
its other clients, nothing in a served request's path goes over the mesh's own bus, and
what a failed move costs is the mesh's ability to change things rather than the services
its modules serve.
2026-09-27 17:59:26 +02:00
jschoubben 51f5ce5c0f 4.5 done: the catch-up needed nothing built, which was the answer
Three ways to do the replay were weighed and the right answer was that the bus being
moved to already does it. A queue on the old bus receives only what is published after
it is bound, so everything built before the catalogue existed was announced to nobody. A
stream is a log and a consumer is a position in it: a consumer created later starts at
the beginning and the builds are simply there. Checked against a running server, because
the decision rested on it.

So "who replays" has no answer because nothing replays. The mechanism was never about
builds — it was about a queue that could not remember, and carrying it across would have
carried a workaround for a limitation that no longer exists, with nothing looking wrong.

Retiring it belongs to step 5, with the rest of what only the old bus needs.
2026-09-27 17:23:49 +02:00
jschoubben fc64a2c1a4 4.4 done: a person's account and their client
The account existed as a permission model and as nothing a person could be given; there
is a record and three commands now. The client is two surfaces over one thing, a command
line and an MCP server, both using the client a module's runtime uses — so what a person
may do is answered by the same permission list that answers it for a module.

Design 25 §7 says nothing of this is built before its bed passes, and this was built
before. Noted in the task rather than quietly ignored.
2026-09-27 17:07:28 +02:00
jschoubben 9fc4e74b0d Merge pull request 'to-be 31: a module declares its fail2ban jail, mesh composes them per node' (#149) from design/a-module-declares-its-jail into main 2026-09-27 15:00:16 +00:00
jschoubben 225dfa9451 to-be 31: a module declares its fail2ban jail, mesh composes them per node
A node's intrusion filter should be composed from its assigned modules, like
its firewall (the Filtering mechanism): a service module (postgres, mssql,
mailu) declares its jail in its manifest (filter + stanza, no node/path per ADR
0112), and the mesh writes the jails of a node's modules into the fail2ban
holder's jail.d. The base (sshd, recidive, ignoreip=mesh-range) stays the
fail2ban module's. Records the model after novox's HAL per-module jails were
lost as dangling symlinks; the ignoreip is now safe on disk, the service jails
need this to be restored.
2026-09-27 16:59:56 +02:00
jschoubben bfefb1dbdb 4.3: the installer can raise a mesh on the new bus
A foundation template that stands the server up, writes its settings and the mesh's
first user list beside them, and starts a controller on the new bus. The first user list
is the installer's because at genesis there is no mesh to compose one — a bootstrap
credential, rotated like the store's.

The carried list is checked against what the controller derives, since a mesh cannot be
raised twice to discover they disagreed. That check immediately found the composer
granting a role's whole event branch as well as the one event it follows.

What is left of 4.3 is running it, which is 4.1's bed.
2026-09-27 16:39:48 +02:00
jschoubben 85c0a3e567 4.2 done: a build is work submitted to a role, on both buses
Both sides behind a seam, one implementation per bus, and the outcome is the role's
own event so one publish reaches the asker, the controller and the catalogue. Checked
against a running server, including the part the decision rests on: a third party
hears the same outcome the asker does.
2026-09-27 16:02:17 +02:00
jschoubben a39765c924 Merge pull request 'ADR 0122: a seat is data the controller owns; a rename is a database update' (#148) from design/a-seat-is-data into main 2026-09-27 13:41:10 +00:00
jschoubben a8921fe737 ADR 0122: a seat is data the controller owns; a rename is a database update
Reviews 0110/0121 after a session where renaming seats cost three freezes, a
builder deadlock, and hand-resolved manifests. The seat rules were right; the
set being a compiled Go slice referenced by name-string everywhere was the
mistake. Seats become a table keyed by a stable id; claims/held/production code
reference the id; a rename is one UPDATE, no rebuild, no re-registration, no
freeze. The build machine reads the set from the mesh instead of embedding it,
removing the controller/builder seat coupling. Closed set and scope naming
unchanged; only storage and reference change. Outstanding renames (registry
seats, private-network scope) wait for this — as data each is a write.
2026-09-27 15:40:55 +02:00
jschoubben c8f430290e ADR 0121: a seat carries the protocol of its role
The mesh's own seats said who does a job and nothing about what may be said to
them or by them, and that gap showed up three times in one day looking like three
different problems: a build machine with three audiences for one outcome and no way
to derive a grant for any of them; an event genuinely about a role with nowhere to
live but the namespace of whichever module holds that role today; and a catalogue
catching up on builds, where every option needed a grant the design refuses.

One cause — the mesh has roles it cannot describe. So the `mesh-*` seats take the
same three fields a module's seat has, and the machinery that already derives
authority, queues and consumers from a declared seat does it for these too.

Builds become work submitted to a role, and `mesh.build.request`,
`mesh.control.built` and the BUILDS stream retire. A work queue shared by several
build machines is exactly what a seat's `accepts` is, so a second mechanism for it
was two places a permission could be wrong. The outcome is the seat's own event,
which means one publish still reaches whoever asked, the controller that records it
and the catalogue that places it — the fan-out a shared exchange gave for free,
written as a subject the mesh derived rather than a topology somebody configured.

That also avoids the grant that ruled out the alternatives: no holder needs
permission to publish into an asker's inbox.

The blocking gap is now named rather than incidental: the shared library has no way
for a module to publish on a seat. The build machine is Go and reaches the bus
directly, so it is unaffected; the artifact-store event waits.
2026-09-27 15:24:17 +02:00
jschoubben d98d6fca11 Merge pull request 'to-be 30: the mesh updates itself on a push' (#147) from design/the-mesh-updates-itself into main 2026-09-27 12:49:32 +00:00
jschoubben b8cdfce16d to-be 30: the mesh updates itself on a push
Records the manual update process (module moved -> build -> reconcile; and the
breaking-change freeze/re-register recovery), and the two things that make
self-update more than a webhook: the build-on-push trigger is currently HAL's
(hal-gitea-tools on :9877), a retirement gap the mesh must replace with its own
forge-webhook trigger wired to every repo including mesh-controller; and the
builder validates manifests too, so a breaking change couples controller +
builder + manifests + hosts, and renaming the builder's own seat deadlocks its
rebuild. Names the transition discipline (accept old+new for one release) that
self-update needs so a push does not auto-freeze.
2026-09-27 14:49:02 +02:00
jschoubben c4a8455e2e Issue 127 resolved; design 29 says what wildcards are and how the rule is checked
Every module named its events the way the old bus spelled a routing key, so on the
new bus every cross-module subscription pointed at a namespace nobody publishes to.
Nothing failed — the services started and none of them reacted. Converted, and the
rule now has checks at both scales: at registration for one manifest, and as a test
across the whole catalogue where a consumed event's emitter is present.

It was larger than the report said, in two directions nobody had looked. Forty-three
files of module code pass the event name at runtime, so the code mattered as much as
the manifests. And both clients had to learn the mapping — without that, converting
the modules would have broken the mesh that is actually running, which is the
opposite of what fixing this was for.

Design 29 gained three things it did not say: what a wildcard is (`*` for one name,
`**` for the rest, spelled the mesh's way and derived to each bus's own), that an
event about a role belongs on the seat and why that is not yet possible, and how the
rule is checked — because "a subscription that matches nothing is silence" is exactly
why nobody noticed thirty-seven manifests being wrong the same way.

4.2 and 4.3 are unblocked. The catch-up half of 4.5 is not: it is a decision, and it
narrowed rather than closed. It cannot be a reply to a module's inbox, because that
needs the blanket grant design 25 §4 refuses.
2026-09-27 14:45:08 +02:00
jschoubben bf7297b7ed Merge pull request 'ADR 0121: keep distribution, retire only verdaccio; node-* seats migrated' (#146) from design/0121-keep-distribution-fix into main 2026-09-27 12:36:54 +00:00
jschoubben 1bb0ef5658 ADR 0121: keep distribution, retire only verdaccio; node-* seats migrated
Records the reversal: distribution stays as the mesh's OCI registry (it serves
every artifact-store:// image); only verdaccio, a redundant second npm registry,
is removed. The 'consolidate onto gitea / retire distribution' direction was
dropped. Also records that the node-* rename was executed as one controlled
migration with a brief compose freeze, and why the delivering registry seats
are deferred rather than folded in.
2026-09-27 14:36:33 +02:00
jschoubben 5a9917f802 Merge pull request 'ADR 0121: a system seat is named for its scope; a module may define its own' (#145) from design/system-seats-are-named-by-scope into main 2026-09-27 12:15:22 +00:00
jschoubben 8a6ee9177c ADR 0121: a system seat is named for its scope; a module may define its own
The control plane's seats grew a second naming style (the-*) beside mesh-*,
and the closed set was the only place any seat could be defined. This settles
both: system seats are mesh-* (one, mesh-wide) or node-* (one per node), named
for scope; a module may define its own seat outside the closed set. Folds in
the seat review: mesh-build-machine (scope fix), mesh-private-network (one
server + client modules, dropping per-node VPN choice), showcase becomes the
first module-defined seat, node-uplink, and the node-* renames — plus the
registry consolidation onto gitea, which reshapes the registry seats and gates
retiring distribution/verdaccio. Records why the renames are a coordinated
migration and why distribution cannot be removed until gitea serves images.
2026-09-27 14:15:05 +02:00
jschoubben dd577e9ebe 1.7 is done; 4.1 waits on nothing but the bed
Minting on both halves, the file delivered per push, and the bus's objects
asserted on every start — verified against a real server that asserting twice
changes nothing, that a machine joining an already-raised bus is accepted, that
each node's consumer is bound to its own declaration subject, and that CONTROL
does not dead-letter before the controller gives up.

Which bus the mesh is on is one fact, and being told about both is refused at
start rather than warned about: a mesh half on each is one where a declaration
goes out on one bus and the report comes back on the other while every component
logs success — ADR 0074's failure arriving through configuration rather than code.

So 4.1 no longer waits on code. Both links speak NATS, the composition happens,
and every claim behind them has a unit test or a check against a running server.
What none of those can stand in for is a mesh raising itself, which is what the bed
is — this is where the code stops and the lab starts.
2026-09-27 03:20:42 +02:00
jschoubben a9b8f41570 Design 25 §4: the server verifies no client certificate, and writes only accounts
Two corrections of fact, both found by building the module's image and connecting
to it as a host would.

The first composed configuration said `verify: true`, which makes the server
demand a *client* certificate — and nothing in the mesh presents one. A host pins
this server's exact certificate and authenticates with the password the mesh
minted, and so does a module's runtime. Every connection in the mesh would have
died at the TLS handshake before any password was looked at, with an error that
reads as a fault in the client. TLS is still required; verify only decides whether
client certificates are checked. Mutual TLS is a later question and would need
machinery the mesh does not have — a certificate per module per node.

And §4 read as though the controller wrote the whole file. It writes the user list
and nothing else: ports, TLS paths and a store directory belong to the container
the module raises. The two files share one directory of necessity, because an
absolute include path is resolved relative to the including file's own directory.

The decision stands in both cases — accounts are composed, not called for, and
passwords are minted and sealed. What changed is what the file says and who writes
which half.
2026-09-27 02:51:40 +02:00
jschoubben 92a5e8fc05 Merge pull request 'ADR 0120: a roster fact carries its format as a template; rewrite to-be 29' (#144) from design/roster-fact-is-a-template into main 2026-09-26 23:51:21 +00:00
jschoubben 8c91ba1cfa 1.7 in progress: the list is derived and the keys are kept
What is in: a bus user's hash is recorded and its plaintext returned once, and
the user list is derived from the machines, what each runs, every manifest and
which machines hold a live token. Permissions stay derived rather than stored,
because a stored copy could disagree with the records it came from while both
looked internally consistent.

What is out, with what each needs, so the next person does not rediscover it:
delivery, which has one open question about what a module declares in order to
receive the file — design 29's ground, not this document's; minting, which is
transport-coupled because an enrolment reply carries one password and a node on
the old bus must not be handed a credential for the new one; and calling the
assertions from a start path.
2026-09-27 01:50:04 +02:00
jschoubben 0f7f628730 ADR 0120: note the shared/region interaction with hq 128
A roster fact may be shared — written into a marked region of the machine's
file (into: block, hq 128) rather than as the whole file. The template
renders the content; shared decides how the host lays it down. Composes with
hq 128: the region mechanism is the host's, the format is the module's.
2026-09-27 01:42:00 +02:00
jschoubben 970da74136 1.7: the composer exists and the composition does not
Correcting a tick and a claim I made one commit ago. 4.1 does not wait on an
enrolment user per live token; it waits on the whole composition, of which that
user is one input.

Tasks 1.3 and 1.4 are honest about what they built — the composer, the
derivation, the permission model, the stream and consumer definitions, the
asserter, all pure and held by unit tests and a golden composition. Nobody wrote
the caller. Measured: outside the package that defines them there is not one use
of the composer, the permission derivation, the stream set, the stream asserter or
the principal type. Step 1's "done when" claims every account and permission
composed from the manifests, and a mesh raised today would stand up a server with
no user list at all.

It also needs state the mesh does not keep. Design 25 §4 says the file holds
bcrypt hashes and that passwords are minted and sealed exactly as today — but
today the mesh mints one, hands it to the broker through a management call, seals
the plaintext to the holder and keeps nothing. With no management call the hash
has to survive every later recomposition, because the first thing a new module or
a person's access change touches is a file that must still hold every other
user's password. No bcrypt hash is stored anywhere in the controller.

Named as its own task rather than folded into 1.3, so the gap between "the parts
of step 1 exist" and "the mesh does any of it" is visible.
2026-09-27 01:39:20 +02:00
jschoubben 4d4012cdf6 ADR 0120: a roster fact carries its format as a template; rewrite to-be 29 around it
The facts mechanism formatted the roster in Go in the control plane — one
formatter per fact, in the consumer's own configuration language. ADR 0120
makes a fact a path and a template: the mesh owns the data, the module owns
the format, and the control plane holds no format at all.

to-be 29 (operator accounts + what lives under a home) is rewritten to ride
it: the ssh files become roster templates, the whole ~/.ssh is owned with a
found/owned boundary that cannot lock the operator out, keys are mesh-owned
through an SSH CA (existing keys adopted not regenerated, the operator's
personal key signed not minted), and the ssh-agent is a user-scoped service.
2026-09-27 01:33:12 +02:00
jschoubben 0a61531c42 WBS 3.5 done; 4.1 waits on one thing, an enrolment user per token
All three halves of the host's link are through seams, and the reply address in
the payload is now proved from both ends rather than one — the test asserts the
transport's own field held the consumer's ack address by the time the request
arrived, so a server that stopped claiming it fails a test instead of letting the
reason become folklore.

Two things had to be built for the host to hear anything at all: a node's
declaration consumer, which only the controller may create, and the enrolment
user's inbox, which design 25 §6 names and the composer granted none of. Both
were silent gaps — a node with either missing looks correct and hears nothing.

What remains is a single piece: something that composes an enrolment user per
live token. On the old bus that account is made imperatively through the broker's
management API; here there is no management API, so issuing a token has to
recompose the server's configuration. It is the only thing between the two links
and a mesh raised on NATS from nothing, so 4.1 now says so.
2026-09-27 01:32:26 +02:00
jschoubben 4b0b659084 Merge pull request 'ADR 0119: a taken tunnel's predecessor is retired once the take is proven' (#143) from decision/0119-a-taken-tunnels-predecessor-is-retired into main 2026-09-26 22:59:15 +00:00
jochen 0042ca9258 0119: link ADR 0118 now that it is on main 2026-09-27 00:58:55 +02:00
jochen 63c19456b4 0119 review: rollback needs the private network unassigned first; a configuration written back is retired again with the first original kept; only the interface's own file, never a link 2026-09-27 00:58:54 +02:00
jochen 2f195d501e to-be 08: the found tunnel's configuration is retired once the take is proven (ADR 0119) 2026-09-27 00:58:54 +02:00
jochen 8a78ff4efe ADR 0119: a taken tunnel's predecessor is retired once the take is proven 2026-09-27 00:58:54 +02:00
jschoubben 762300a380 Merge pull request 'issue 130: undeclaring a service stops it, even one the mesh only reloads or keeps running' (#142) from issue/130-undeclaring-a-service-stops-it into main 2026-09-26 22:58:43 +00:00
jochen 248c99ca6c 0118/130: link ADR 0117 now that it is on main 2026-09-27 00:58:17 +02:00
jochen 13208f0f48 0118: give a unit back the state it was found in — never-stop broke the converge rollback; process removal found and fixed 2026-09-27 00:58:17 +02:00
jochen ebd19c4c6c ADR 0118: undeclaring removes what the mesh made, gives back what it changed, leaves the machine's units as they are — resolves issue 130 2026-09-27 00:58:17 +02:00
jochen dd4cbabffb 130: ADR 0117 named, not linked, until it is on main 2026-09-27 00:58:17 +02:00
jochen 5b76a09da6 issue 130: undeclaring a service stops it, even one the mesh only reloads or keeps running 2026-09-27 00:58:17 +02:00
jschoubben a363a605cb Merge pull request 'issue 129: nothing makes a machine trust the mesh's own certificate authority' (#141) from issue/129-nothing-makes-a-machine-trust-the-meshs-authority into main 2026-09-26 22:57:30 +00:00
jschoubben 24835ab710 Merge pull request 'issue 128: the machine's hosts file is written whole, and on a workstation it is shared' (#140) from issue/128-the-hosts-file-is-written-whole into main 2026-09-26 22:57:02 +00:00
jschoubben ba397d4cbe Merge pull request 'ADR 0117: a machine's uplink is a seat — the mesh configures the manager, never the link' (#139) from decision/0117-the-uplink-is-a-seat into main 2026-09-26 22:56:27 +00:00
jschoubben f555d523c7 WBS 3.4 is done both halves; issue 127 holds 4.2, 4.3 and catch-up
The controller's inbound is through a seam with both transports behind it, and
the store window is now the server's rather than the controller's memory. Seven
claims about that were asked of a running server rather than reasoned.

Wiring the controller's own subscription is what found issue 127: every event
name in the catalogue is still written the way a routing key is, so design 29's
derivation turns a consumer's declaration into a subject no emitter publishes.
Thirty-seven manifests, one that cannot be composed at all. It fails on the first
mesh raised on the new bus and not before, which is why nothing had caught it —
the conformance fixtures pin one emitter against one subject, and both halves of
that pair are correct.

The node-facing flows are unaffected: those subjects are the mesh's own and
derive from nothing a module declares.
2026-09-27 00:55:07 +02:00
jschoubben 4f93d304d7 WBS: a person's account is done; the client is not blocked by step 3 2026-09-27 00:18:05 +02:00
jschoubben d898bd87e8 Step 5.4 was wrong from 0119 onward; removed
It waited on a retirement condition 0119 abolished when it made the
deprecated broker an ordinary provider. A step waiting for a condition
nobody set would sit open forever.
2026-09-27 00:16:45 +02:00
jschoubben 7e4da874a9 Design 25 §2: the eaten reply address is verified, not assumed
A claim the whole enrolment handshake rests on, now measured against a
running server rather than reasoned from documentation — and held by a test
so it cannot become folklore if a server version changes.
2026-09-27 00:15:00 +02:00
jschoubben 53092020eb WBS: asking a tool is through the seam; a build is a different shape 2026-09-27 00:11:44 +02:00
jochen df4a3538c3 0117 review: a holder's service declares no state — the manager's lifecycle is the machine's; a start-only setting's gap on a fresh machine 2026-09-27 00:07:42 +02:00
jschoubben 00817fb9e3 Design 25: the store window, and what moving it into the server changes
The guarantee is the same and the mechanism is simpler — a nak with a
delay, no parked list, nothing lost when the controller restarts. It costs
one thing: a naked message comes back whatever happened meanwhile, so an
older report is redelivered after a newer was applied. A report already
carries the digest of the declaration it answers, so supersession becomes a
check rather than memory — ordering settled by what a message says, not by
when it arrived.
2026-09-27 00:06:15 +02:00
jochen 90b8eeff6b 129 review: the example name is a routed name, not a doubled suffix 2026-09-27 00:03:36 +02:00
jochen a865fc7d79 128 review: located; the fix as built (block, at, never held, order) 2026-09-27 00:03:30 +02:00
jochen c3313f6e17 0117 review: seat table row + decision cited, networkd/dhcpcd lines match the modules, block placement, dispatcher scope, references 2026-09-27 00:03:15 +02:00
jschoubben 9946e852e1 WBS: 3.5's outbound half is in 2026-09-27 00:01:56 +02:00
jochen 60a53f9b18 issue 129: nothing makes a machine trust the mesh's own certificate authority 2026-09-26 23:58:48 +02:00
jochen ee31f9f761 issue 128: the machine's hosts file is written whole, and on a workstation it is shared 2026-09-26 23:58:48 +02:00
jochen 0e066473b3 ADR 0117 accepted; a manager that cannot reload takes the setting at its next start (dhcpcd, measured) 2026-09-26 23:58:48 +02:00
jochen 504adef221 ADR 0117: a machine's uplink is a seat — the mesh configures the manager, never the link 2026-09-26 23:58:48 +02:00
jschoubben 9f6aa7ea9c The bus is the mesh's centre, not a transport that replaced one
Two things. A paragraph from the superseded 0117 survived beside the 0119
correction that reversed it, so §5 said both that the amqp interface
retires and that it does not. The stale one is gone.

And the framing. §1 opened with "the bus carries five kinds of traffic
today, and this design keeps the five", with a column mapping each to the
queue it used to be — which describes the mesh's nervous system as a port
of something that did a fraction of this. It now says what the bus is: a
role addressable without knowing its holder, the mesh's own state, work
that queues until somebody can do it, and permissions derived from what a
module declared. Conditions, observation and a person's client land there
too as they are built.

Glossary gains `bus` and `the deprecated broker`, with a note on why not to
say "compatibility broker" or name it after a protocol — the second invites
exactly the backwards framing this commit removes.
2026-09-26 23:54:09 +02:00
jschoubben 8deecf620a Merge pull request 'to-be 29: a node has operator accounts, and the mesh owns what lives under a home' (#138) from design/29-a-node-has-operator-accounts into main 2026-09-26 21:53:39 +00:00
jschoubben 64bbdc30c1 to-be 29: a node has operator accounts, and the mesh owns what lives under a home
The mesh models machines but not the people on them — a node record
holds no username, and no module places anything under a home. So who
you are on each node (jochens/ace/jochen) is unknown to the mesh, and
nothing owns ~/.ssh, dotfiles or ~/.config. HAL knew it; the nox mesh
dropped it. Proposes the account as a node fact and a home-scoped
resource class (the ~/ mirror of ADR 0112's /var/lib placement), with
the login key staying the operator's (ADR 0051). Not urgent — HAL's
generators still run — load-bearing at node-by-node retirement. Found
generating ~/.ssh/config from HAL's registry, which nox has no
equivalent for.
2026-09-26 23:53:13 +02:00
jschoubben 672c994afa WBS: 3.4's seam is in, outbound half through it 2026-09-26 23:47:40 +02:00
jschoubben 2f9bb73685 WBS: the first fixtures are in, and what byte-for-byte means 2026-09-26 23:41:15 +02:00
jschoubben 5c193b3f54 WBS: 3.8's check is written 2026-09-26 23:34:42 +02:00
jschoubben fbf9440e8e WBS: 3.7 and 3.8 done, with the one check 3.8 still owes 2026-09-26 23:34:04 +02:00
jschoubben 9510bf5311 WBS: 3.2 done 2026-09-26 23:33:10 +02:00
jschoubben d940e14ec8 Design 19: the protocol on NATS
Task 3.2. ADR 0074's model is untouched — floor plus capabilities, partial
implementations legitimate, identity from the credential, dedup on
x-event-id, conformance as executable fixtures. The transport beneath it is
rewritten: exchanges and queues become subjects and streams.

Statements marked *verified* were checked against a running server while
the runtime's client was written, not reasoned from documentation. Three
of them are things the specification would otherwise have got wrong:

- the payload is the body alone, with metadata in NATS headers; an
  implementation that nested the whole envelope would agree with nobody
- a durable name may not contain a dot, while the ack subject joins two
  names with one — conflating them looks right in a permission list and is
  refused as a consumer name
- a certificate must carry a name the bus is dialled by, because the NATS
  client has no hook to replace hostname verification the way pinning did
  on AMQP

And one limitation lifts: a module may now call another's tool. Issue 049
recorded that a scoped account could not declare the reply queue a caller
needs, and ADR 0095 routed every ask through the control plane because of
it. Per-account inbox prefixes plus allow_responses replace that. ADR 0095
is not reversed — the control plane is still how a person asks — but
module-to-module calling stops being a question about capability and
becomes one about policy, which `uses` already answers.
2026-09-26 23:32:48 +02:00
jschoubben 80456981be WBS: 3.6 done, and the certificate constraint it surfaced
The NATS client has no checkServerIdentity hook, so pinning no longer makes
the name check redundant — the bus's certificate must carry a SAN matching
the address nodes dial.
2026-09-26 23:29:56 +02:00
jschoubben d2ed3152d3 Seat renames done; 0118 was wrong that it was a migration
A holding is derived at resolution from manifests, never stored, so there
are no recorded old names to rewrite. The work is an edit plus a kept rename
table — kept because a module lives in its own repository and may be
registered long after the catalogue stopped using an old name.
2026-09-26 23:06:55 +02:00
jschoubben 24d99ddd24 WBS: 3.9 done, and 1.4's client with it 2026-09-26 22:28:59 +02:00
jschoubben b5730525fe Merge pull request 'issue 127: a declaration that shrinks to empty is skipped, not sent' (#137) from issue/127-shrink-to-empty into main 2026-09-26 20:25:19 +00:00
jschoubben bb334e138b issue 127: a declaration that shrinks to empty is skipped, so the node keeps what it should drop 2026-09-26 22:24:58 +02:00
jschoubben b759e36bfd Design 29: tools, not serves; WBS 3.9 partly done
The manifest already uses serves for a provision's facts, so a module's
tools take their own key. Declaring them is itself new — until now a
module's tools existed only in a runtime environment variable.
2026-09-26 22:17:17 +02:00
jschoubben 0b8e84334f Why module events share one stream, checked against the server
Storage is not a property of a subject — a stream is a separate object that
covers one — so the question is always how many streams, not which topics
are durable.

Three facts decide it, two of them verified rather than assumed: NATS
refuses overlapping streams instead of merging them, so a shared stream
plus a per-module one is not available at all; a filter cannot express an
exception; and a stream per module turns one cross-module consumer into one
per module. So one stream, with per-subject caps for the fairness that
matters. Per-module age is genuinely unavailable, and a module that needs
it declares a seat.
2026-09-26 21:49:55 +02:00
jschoubben 39c802cbd4 Step 2: adoption recreates the bus once, on purpose
2.3 was already true and is now proved — the seat refusal is generic, and
three tests pin what matters: a second bus is refused by name, a different
bus implementation is refused too (which is what makes the bus replaceable),
and the AMQP broker no longer contends so both run on one mesh.

2.1/2.2 turned out not to be a no-op. The host keeps a container only when
its spec matches exactly; genesis raises the upstream image and the module
declares the mesh-built one carrying the entrypoint, so assigning it
recreates the container. That is ADR 0067's pivot and it is safe only
because the bus carries nothing yet — which is why step 2 comes before
anything speaks NATS. After it, never again: the config is a directory
mount, so accounts change without touching the container's spec.
2026-09-26 21:40:02 +02:00
jschoubben 86a084b7ff The mesh bus is required, not ambient
Design 29 said no module requires the bus. The catalogue disagrees: 49 of
72 modules take a broker credential and 23 do not, so an ambient connection
mints an account for a third of the catalogue that never speaks — and the
49 each hand-write the path it lands at, which is provisioning done badly
by hand.

The bootstrap argument that made it ambient was narrower than it looked.
"A provisioner needs an account before it can run" is true of a provisioner
process and says nothing about a provision the controller answers, and the
controller is not waiting on a bus account to compose one.

So: the mesh-broker seat delivers mesh-bus; a module requires it and gets an
address, a sealed credential and the trust to verify the server; a module
that requires nothing has no account at all. The requirement delivers the
connection, the declarations shape the authority, and declaring a subject
without requiring the bus is refused as incoherent.

mesh-bus and nats are deliberately two names: a module may run its own NATS
as a backing service exactly as one provides amqp, and a manifest saying
"nats" would otherwise mean either the mesh's nervous system or a private
queue.

The seat's Delivers was wrong twice today — amqp, then empty — and the
comment says so rather than reading as though it were always right.
2026-09-26 21:17:21 +02:00
jschoubben 85f972749a AMQP is a provision, not the bus
0117 went a step further than it had grounds for. It was right that the bus
is the only bus, and wrong that the amqp interface must therefore retire —
because it conflated two reasons to want a broker. Using one to reach
another module is a second bus and stays refused. Needing an AMQP broker as
a backing service, the way something needs a database, is ordinary, and
forbidding it would make the mesh unable to run normal software while
calling that architecture.

So the broker becomes a plain provider module: no seat, not foundation,
never raised at genesis, no retirement condition. lavinmq now claims nothing
and provides amqp; nats claims mesh-broker and provides nothing.

The rule that survives is about direction, not software: inter-module
communication goes over the bus. A module may hold a broker for itself; it
may not use one as a channel to another module. That is a review judgement
where 0117 could have used a parser, which is the honest cost.

0106's progressive insight was itself wrong and is corrected by a second one
there — nothing moves off the old broker, so its "one purpose" sentence does
not become true, it is just not what that server is.

The insight check needed two fixes it found itself: a date may carry
trailing words, and a bold run with a link is discussing an insight rather
than marking one. All four bad shapes still fire.
2026-09-26 21:08:40 +02:00
jschoubben 7abb268de6 Step 1 done but for its bed
1.1 to 1.4 built and tested. 1.5 turned out to need no controller change:
it already resolves the broker by seat and names no broker module in its
source, which is what ADR 0079 was for. The genesis module set naming is
scenario and installer config, carried with the bed.

Recorded what must NOT change yet: the amqps:// credential shape and the
5671 default are correct until the rollout, because steps 1-4 leave every
node on AMQP.
2026-09-26 21:03:35 +02:00
jschoubben 78a2274baf Designs 25 and 29 disagreed about the subject space; implementing found it
29 put a module's events and tools in one namespace, 25 kept mesh.events.*
and mesh.tools.*. One namespace is right — a module's authority over its own
name becomes a single pattern the server enforces — but it needs a kind
token, because a stream is a subject filter and mesh.mod.*.> would persist
every tool call in the mesh. Tools stay on core NATS for the reason 25
already gives.

So: mesh.mod.<module>.event.<name>, .tool.<name>, and seats the same shape.
2026-09-26 21:02:18 +02:00
jschoubben 3f9b316015 Design 25: a scoped inbox needs allow_responses, or nothing can answer
Found composing the first real configuration. Scoping every inbox to its
owner is right and leaves a responder unable to reply, because the answer
goes to the caller's inbox. The fix is not a wider grant but the server's
own allow_responses: one reply to the subject of a message the user actually
received. Without it every tool call times out while the permission list
looks correct.
2026-09-26 20:58:34 +02:00
jschoubben e05825a881 Design 29: versioning, provisioning and secrets on the bus
Versioning: additive is free; a breaking change is refused while callers are
bound, and the refusal names them, because the mesh already holds the uses
graph; a real break versions the subject, not the seat name, so the role
does not fork; binding is a recorded pin, not a drift to whatever is newest.
Semantic change stays open — no fingerprint sees it, and saying so beats
implying the check is complete.

Provisioning: a provisioner's create/remove/holds IS a serves protocol, so a
provision interface is a seat that also delivers a credential — which is why
design 26 already allowed that. The per-consumer resource is what stops the
two collapsing into one.

Secrets: sealed, so the bus is never trusted with plaintext — but sealed is
not enough, because a stream persists and a durable ciphertext is an archive
the day a key leaks. So a secret never enters a stream: core request/reply
only, and a declaration names a secret rather than carrying one, which is
0098's fetch-don't-store applied where carrying is worst. The vault's own
credential and the bus's own accounts are the two bootstrap exceptions,
resolved the way 0067 resolves the control plane.

Also rewrote the addresses paragraph, which was too compressed to follow:
on-bus addresses disappear because nothing stores them, off-bus ones are
untouched and still 0098's problem, and the bus's own address is the one
that cannot be a subject.
2026-09-26 20:44:16 +02:00
jschoubben 7b4916e9ec Modules declare their own seats; the mesh reserves mesh-*
The architecture 0117 opened needs a module to offer a service as a role on
the bus — one holder, addressed by what it does. A closed table in the
controller cannot express that: a capability a module contributes would
require changing the mesh itself.

But 0110 closed the set for a good reason — nothing could say what seats a
mesh had, and the hand count came out at eleven of thirteen. That argues for
enumerable, not hardcoded, and 0110 weighed free-form against a fixed table
without considering a third option: closed at any moment and derived from
the catalogue. A derived list cannot drift, which is how the count broke.

So: the mesh's seats stay the mesh's, reserved by the mesh- prefix so the
prefix is the rule and there is no list to maintain; ten seats are renamed
to restore 0079's convention; everything 0110 decided about what a seat IS
survives untouched.

Design 29 carries the declaration model: three namespaces, subjects derived
from local names so a manifest survives the wire changing, queues never
declared, five relationships (the job and state shapes 0041 had no room
for), and the build-publish-deploy lifecycle with hard, soft and build-time
dependencies distinguished.

0041 gets a progressive insight: "no per-consumer setup, only a
subscription" was a fact about a topic exchange, and a JetStream durable
consumer is a real object someone creates.

WBS 1.3/1.4 were wrong and say so: streams come at registration and
consumers at assignment, so only the foundation set belongs at genesis.
2026-09-26 20:34:32 +02:00
jschoubben 3f14264c7b Merge pull request '112 is fixed: a carried peer is nameable, and the resolver answers for all four machines' (#136) from issue/112-fixed into main 2026-09-26 18:13:53 +00:00
jschoubben 25d599094f 112 is fixed: a carried peer is nameable, and the resolver answers for all four machines 2026-09-26 20:13:37 +02:00
jschoubben b5b68e8852 Design 25: a host directory bind, not a named volume
Issue 115 is resolved and converted four modules away from named volumes;
the bus's own data is not the place to bring one back. Also: NATS carries
TLS on the client port rather than beside a plaintext one, so there is no
5671/5672 pair to mirror.
2026-09-26 19:34:54 +02:00
jschoubben 814c9e563f The bus is the only broker; step 1 starts
A module does declare requirements the provisioner fulfils — but the broker
it gets that way is a private vhost, the analog of a database, not the
mesh's bus. Two modules of the new mesh depend on it, so the compatibility
broker was never single-purpose and its retirement would have stranded them.

NATS is the heart: one bus, a module's messaging is subjects on it scoped by
what it declares, and no module is handed a server of its own. The seat
delivers nothing; the interface retires with the broker. Also closes the
EVENTS question — one stream, on the bootstrap argument, not preference.

Designs 25 and 28 go in-progress: step 1 is starting.

The insight check caught a false positive on its own first real use — its
bold-run pattern crossed newlines and joined an unrelated `**` to the
marker. Constrained to one line, still catching all four bad shapes.
2026-09-26 19:26:07 +02:00
jschoubben db4ca9b043 Merge pull request 'The bus in five steps: a decomposition, a breakdown, and progressive insight' (#134) from feat/the-bus-in-five-steps into main 2026-09-26 17:05:43 +00:00
jschoubben 5a6d0e111d Merge remote-tracking branch 'origin/main' into feat/the-bus-in-five-steps
# Conflicts:
#	02-DECISIONS/README.md
2026-09-26 19:05:36 +02:00
jschoubben c94e2ece53 Merge pull request 'Give 0115 a home, and match its batch's status — main is failing both checks' (#135) from fix/0115-checks-on-main into main 2026-09-26 17:04:50 +00:00
jschoubben 0af6479b2c Give 0115 a home, and match its batch's status
Both repository checks fail on main. 0115 is cited by no design, and it is
marked accepted while resting on 0112, which is proposed.

Design 27 already states the rule the record decides — "a module is assigned
at most once to a node, and that pair is the assignment's identity" — so it
is the home, and now says so. And 0112, 0113, 0114 and design 27 are all
proposed: the batch is under review, so the record is too. Promoting 0112
instead would be marking a record accepted to satisfy a check, which
check_rests_on names as a failure this repository already made once.
2026-09-26 19:03:14 +02:00
jschoubben 77a1493df4 Renumber to 0116: another record took 0115 on main
PR #133 landed a different 0115 while this branch was open. The bus record
is now 0116, with every citation in designs 19, 25, 28 and the index
following it.

Note: cycle.py and records.py both fail on main as merged, on that record —
nothing cites it, and it rests on 0112, which is still proposed. Both
pre-date this branch and are left for their own change.
2026-09-26 18:56:34 +02:00
jschoubben 21b54ee52f Merge main: 0115 was taken by another record 2026-09-26 18:54:59 +02:00
jschoubben 3950c2b75d Merge pull request '0115: one assignment of a module per node — the requirement is dropped' (#133) from decision/0115-one-assignment-per-module-per-node into main 2026-09-26 16:53:12 +00:00
jschoubben 9516d31a62 0115: one assignment of a module per node — the requirement is dropped, the module's name is the identity 2026-09-26 18:52:59 +02:00
jschoubben 1c808898a5 Allow progressive insight, and apply two to the bus record
A record can assert a fact that goes stale while the decision it supports
stays right. Superseding for that buries a sound record under a second one
and makes every reader work out which is live. So a correction of fact is
now made in place, marked and dated, with the old wording quoted — bounded
by three conditions and checked by records.py, which fires on an unmarked,
undated or back-dated note. Judgements still supersede.

Applied to 0115: no conformance suite exists to recapture, and the full
genesis bed cannot run until the links exist. Designs 25 and 28 follow.
2026-09-26 18:39:38 +02:00
jschoubben 6ab113e6c6 Break the bus work down, measured, in dependency order
Counting the surface first changed the plan twice: the genesis bed cannot run
until the links exist, so it belongs to step 4, and there is no conformance
suite to recapture — step 3 builds one against the current bus before moving
it. Both corrections are recorded in the breakdown rather than edited into
ADR 0115. Also indexes design 25, which was never listed.
2026-09-26 18:24:14 +02:00
jschoubben fe0c1e9da2 Divide the bus work into five steps, each proved on its own
The NATS change was recorded as one undivided item, which hid three gaps:
a mesh already running had no adoption path, the protocol specification did
not know its transport was being replaced, and nothing was runnable until
everything was. Dividing it is what surfaced them.
2026-09-26 18:19:56 +02:00
jschoubben dc80116b09 Merge pull request 'to-be 27: the three gaps answered' (#132) from design/to-be-27-gaps-answered into main 2026-09-26 15:44:53 +00:00
jschoubben db142ffa26 to-be 27: the three gaps answered — root is a node setting, the mesh's writes need no module-visible reservation, resolution is the controller's at composition 2026-09-26 17:44:38 +02:00
jschoubben e38814962d Merge pull request 'Issues 125 and 126: two apply-layer gaps the novox session hit live' (#131) from issue/125-126-from-the-novox-session into main 2026-09-26 15:25:41 +00:00
jschoubben 7842457d4b Issues 125 and 126: two apply-layer gaps the novox session hit live
125: a hold is not a line in the apply report — sixteen resources held
for an untaken module while four surfaces reported success, and the
operator stopped the edge's predecessor on their word (the route-proxy
flip outage). 126: a changed volume path neither recreates a running
container nor warns, and a roll-out upgrade policy makes a build a
deployment — together they turned a data-path migration into a forge
outage (the /var/lib move). Filed as 119/121 in the novox session
before syncing; renumbered past the other session's 119-124.
2026-09-26 17:25:27 +02:00
jschoubben 782d5ace04 Merge pull request 'Issues 119 and 122: what a host path is, and what a node's layout would replace' (#130) from issue/119-122-what-a-host-path-is into main 2026-09-26 15:21:34 +00:00
jochen f709e8e0fb Issues 119 and 122: what a host path is, and what a node's layout would replace
Not every host path names this machine. A system file the mesh owns is at that path on every machine
of the kind — the path is the fact. The operator's shared data is already answered as an access. A
path inside a container is the software's contract. What is left, and what a node's default layout
would replace, is 514: a module's own data, and what the mesh writes for that module.

Documents the reservation model as the records already have it — a root per node, one directory per
assignment, a placement for adopted data — and the three things nothing states: where the root comes
from, what sits beneath it, and the order the 514 are retired in.

Stops there deliberately. Changing where a definition looks without moving the data does not fail: the
mesh creates the directory, the container starts, the service comes up empty. Retiring these is a data
migration with a verification step, module by module, and belongs with whoever can see the machine.
2026-09-26 17:21:14 +02:00
jschoubben c5e1a9a8d5 Merge pull request 'Issue 122: count host paths, not paths' (#129) from issue/122-host-paths-not-container-paths into main 2026-09-26 15:11:33 +00:00
jochen c4d9b515ea Issue 122: count host paths, not paths
A path inside a container is not a fact about the machine — /run/secrets and the directory a server
keeps its data in are the software's own contract, true in any mesh that runs it. Only the host side
of a mount names where it landed.

The first sweep matched path-shaped strings, so it counted both halves of every mount and every
in-container location a value mentioned: 798. Counted by role — directory and file resources, the
host side of mounts, accesses, and the targets of binds, grants, receives and secrets — it is 698
across 70 definitions.
2026-09-26 17:11:18 +02:00
jschoubben 1012fff607 Merge pull request 'Issue 122: count it properly — 30 of 71 definitions name this installation' (#128) from issue/122-the-census into main 2026-09-26 15:09:17 +00:00
jochen d5a4cb0d4f Issue 122: count it properly — 30 of 71 definitions name this installation
The five modules this report first named were what a first look found. A sweep of all 71: 49 public
domains across 26, the node's own name 58 times across 19, a routable IP 12 times in one, 798
absolute paths across 70 (issue 119's number, grown), and no email addresses at all.

The sharpest case is not a domain: a mail module states the node's public IPv4 as the address it
trusts a real-IP header from, so a node that moves or gains a second address stops attributing mail
correctly, silently. Two upstream resolvers are excluded deliberately — naming a public DNS service
is a policy default, true of any mesh, not a fact about this one.
2026-09-26 17:09:01 +02:00
jschoubben 5db79cd168 Merge pull request 'Issue 124: a consumer cannot be told a value its provider derived for it' (#127) from issue/124-a-consumer-cannot-be-told-what-its-provider-derived into main 2026-09-26 14:47:25 +00:00
jochen 146fd6b3a8 Issue 124: a consumer cannot be told a value its provider derived for it
The object store derives each consumer's bucket from the login the mesh minted, and never reads the
one a definition named. The consumer still has to tell its own software which bucket to use, and has
no way to be told: bound values come from the provider's serves, which is a literal block identical
for every consumer, and a provisioner returns nothing. So all three consumers wrote the answer down
by hand and one of them wrote the predecessor's bucket — a key scoped to one bucket and software
asking for another, which reads like a credential fault and is not one.

Records the general shape: any interface where the provider names the resource forces the consumer to
reproduce the provider's rule, kept in agreement by hand and checked by nothing.
2026-09-26 16:47:06 +02:00
jschoubben 007e3f4b03 Merge pull request 'Issue 123: the image registry is named after a role, and *artifact* is defined as one format' (#126) from issue/123-the-image-registry-is-named-after-a-role into main 2026-09-26 13:38:00 +00:00
jochen 3abb3a3c07 Issue 123: the image registry is named after a role, and artifact means one format
Three wordings disagree, and the confusion is the damage: the glossary defines artifact as an OCI
image while the build vocabulary already names four kinds in use, two of which are not images; the
image registry's seat is named after its job while ADR 0079 names foundation seats after their servers
and ADR 0109 names package seats after their ecosystem; and prose that says 'the module's image' reads
as though a module were an image.

Records the question the naming hides: ADR 0075 keeps two provisions because packages and images are
two protocols, and already allows the forge to provide the artifact store. The second implementation
rests on a bootstrap argument, and the forge has the same upstream-server shape the store and broker
have, which ADR 0078 raises as plumbing and adopts in place.
2026-09-26 15:37:41 +02:00
jschoubben 112963524f Merge pull request 'Issue 121: retract step 4 — the forge's service is not built' (#125) from issue/121-step-4-retracted into main 2026-09-26 13:36:36 +00:00
jochen a1d8b478ec Issue 121: retract step 4 — the forge's service is not built
Step 4 claimed the forge cannot exist as a container before the builder has built its image. The
forge's server is an upstream public image pinned by digest; only its runtime sidecar is built. The
store and the broker have the same shape, and ADR 0078 raises both at genesis as plumbing and adopts
them in place — so the forge can be raised the same way and serve git, packages and OCI before
anything is built.

What survives: a grant is minted by the provider's runtime sidecar, which is built, so the question
is whether raise-service, grant, build-sidecar simply works. Sequencing inside the mesh, not images.

The wrong version came from taking a record's bootstrap argument at face value instead of comparing it
to how the store and broker are raised — one command away in the manifests.
2026-09-26 15:36:19 +02:00
jschoubben 1489d17238 Merge pull request 'Issue 108: the second door was attached to the wrong thing' (#124) from issue/108-one-door-after-all into main 2026-09-26 13:22:55 +00:00
jochen c343fc68c5 Issue 108: the second door was attached to the wrong thing
Two things were treated as one. The mesh's own artifact store holds the store seat and is internal by
design — reached by name over the overlay, no accounts, ADR 0082. Serving a registry publicly is a
service the mesh can host: a module with its own name, accounts and storage, like anything else it
runs for somebody. The conversion this report was written beside gave the seat holder a second public
door over the same filesystem, which is neither.

Keeps the original issue whole — no garbage collection, and the settings a routine needs are not
enabled — drops the two-door complication, and sharpens one thing: deletion on the only door is
deletion on a door with no accounts, which the predecessor kept behind its authenticated one.
2026-09-26 15:22:09 +02:00
jschoubben 859ff49ff2 Merge pull request 'Issue 122: the pattern is already on main, in five modules' (#122) from issue/122-the-instances-already-on-main into main 2026-09-26 13:07:42 +00:00
jochen e399a2c148 Issue 122: the pattern is already on main, in five modules
Asked whether merging the three reviewed changes would set a precedent. It would not: five modules
already carry a name belonging to this one mesh — a workflow module stating its host, protocol and
absolute webhook URL, and two carrying a full clone URL for a repository on the mesh's own forge.

That changes what the issue is for. There is no version of this catalogue today that does not name
the mesh it was written in, so refusing three changes buys nothing and a mechanism is the only thing
that removes any of them. The three were merged on that reading, each PR saying so.
2026-09-26 15:07:22 +02:00
jschoubben bfcb660a8e Merge pull request 'Issue 122: a module cannot ask for its own public name' (#121) from issue/122-a-module-cannot-ask-for-its-own-public-name into main 2026-09-26 12:58:41 +00:00
jochen d7078061ea Issue 122: a module cannot ask for its own public name
Three open module changes independently wrote one mesh's names into the catalogue — two literal
public URLs, because the software generates absolute URLs behind a proxy, and one bucket renamed to
match what exists here. None was careless: the mesh composes <label>.<public-domain> for the proxy
and never hands it back to the module that asked for the route, and no interpolation yields a public
name, so writing the answer down is the only expressible option.

Files it rather than blocking the three, because the fix is a mechanism and the instances are live
needs. The cost is stated: a second mesh installing the identity provider gets the first mesh's
hostname, and nothing distinguishes a literal domain from a version number.
2026-09-26 14:58:18 +02:00
jschoubben bba9371832 Merge pull request 'The seats as they run, and to-be 26 implemented' (#120) from design/26-the-seats-implemented into main 2026-09-26 12:50:02 +00:00
jochen 3ff38d2ee1 The seats as they run, and to-be 26 implemented
Both halves are on their main branches, so the seats stop being an intention. Writes the as-is
document from the controller's code and the catalogue's manifests: the closed set of fourteen, the
three refusals a claim meets, the holder being an assignment and nothing else, and the one place a
seat changes resolution — which of several providers answers, never whether a requirement may go
unanswered.

Two things the as-is layer exists for are stated rather than smoothed over: a seat cannot answer
before it is held, which is the standing condition issue 121 records; and capacity is not
implemented at all, so the design's bench has no counterpart in the code.
2026-09-26 14:49:44 +02:00
jschoubben a89b0b5586 Merge pull request 'Issue 121 diagnosed: the seats work renamed the requirement, not the order' (#119) from issue/121-diagnosis into main 2026-09-26 12:48:12 +00:00
jochen 9766a3afce Issue 121 diagnosed: the seats work renamed the requirement, not the order
Asked first whether the seats change fixed this in passing, since it landed the same day and
touches both manifests the report names. It did not: the seat's holder is consulted only where
several nodes provide the thing, and with none providing it resolution refuses outright. Genesis
has no exemption — the unchecked first pass exists to learn what each node offers, and a
declaration is never built from it.

Records the part that did change: the three tests left failing on purpose were deleted by the
controller's seats PR and replaced with passing seat-based ones, so the gap is invisible again.
Adds the resolver to located-in, since that is where the refusal is.
2026-09-26 14:47:45 +02:00
jschoubben 8211dfd72c Merge pull request 'Accept ADR 0110 and ADR 0111; the seats design is in progress' (#118) from decide/0110-0111-accepted into main 2026-09-26 12:28:42 +00:00
jochen a7b2db9efc Accept ADR 0110 and ADR 0111; the seats design is in progress
The seats half of to-be 27's review is settled, so the two records it rests on are accepted and
the vocabulary catches up: the glossary's *seat* becomes a named role from a closed set, held by
an assignment and possibly delivering a provision, and 23 — Choosing a provider gains the seat
step in resolution, with ambiguity still refused rather than guessed. Both were held back when
0110 was proposed, because a document may not rest on a record that is not accepted.

26 — The seats moves to in-progress rather than designed: it names the files that implement it,
and naming a file claims implementation, which is only defensible once those files are on the
owning repositories' main branches. It becomes implemented when mesh-controller #63 and
mesh-catalog #69 land.

0112, 0113 and 0114 stay proposed; to-be 27 stays proposed with them.
2026-09-26 14:28:23 +02:00
jschoubben 99ffa6474d Merge pull request 'Issue 121: builder's real package-registry grant deadlocks a genesis bootstrap' (#117) from issue/117-builder-package-registry-deadlocks-genesis into main 2026-09-26 12:20:53 +00:00
jschoubben cb770f95c4 Merge pull request 'Issue 118: the analytics store answers the dial and times out the query' (#116) from issue/118-umami-store-query-timeout into main 2026-09-26 12:20:46 +00:00
jochen 09e502058b Issue 121: builder's real package-registry grant deadlocks a genesis bootstrap
Renumbered from 117, which is taken on main by 'a module's own code is a container in one record
and a process in another' — two reports claimed the same number and git would not have said so.

Scrubbed the node's name and a real registry path; this repository is public.
2026-09-26 14:20:05 +02:00
jschoubben 74b88d8efc Merge main 2026-09-26 14:19:58 +02:00
jochen b083790b21 Issue 118: the analytics store answers the dial and times out the query
Keeps 118: the other claimant to this number is on main as issue 119, where ADR 0112 points.

Scrubbed the service's public name — this repository is public — and completed the report's
frontmatter with the fixed-by and amended-design keys every other report carries.
2026-09-26 14:19:40 +02:00
jschoubben 0740de9d14 Merge main 2026-09-26 14:19:33 +02:00
jschoubben 79c692959e Merge pull request 'To-be 27 (proposed): a module requires, the mesh resolves — with ADRs 0109–0114, research 016 and issue 119' (#113) from design/27-a-module-requires-the-mesh-resolves into main 2026-09-26 12:17:56 +00:00
jochen d2044fb7b4 Merge main: the bus design, issues 113/114/117/120 and research 017 landed
# Conflicts:
#	04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md
2026-09-26 14:16:05 +02:00
jschoubben 01c6b89cc5 Merge pull request 'Design 25: the bus on NATS — proposed architecture for review; issue 103 resolved' (#93) from design/25-the-bus-on-nats into main 2026-09-26 12:15:19 +00:00
jschoubben 93502a05dc Merge pull request 'Issue 117: a module's own code is a container in one record and a process in another' (#110) from issue/117-a-modules-own-code-is-a-container-and-a-process into main 2026-09-26 12:14:38 +00:00
jschoubben 1272a97fd6 Merge pull request 'Research 017: a mesh that heals itself' (#115) from research/017-a-mesh-that-heals-itself into main 2026-09-26 12:14:24 +00:00
jschoubben de0c9b6cfc Merge pull request 'Issue 120: a provisioner remembers what it did, not what is there' (#114) from issue/120-a-provisioner-remembers-what-it-did-not-what-is into main 2026-09-26 12:14:17 +00:00
jschoubben 77934991c4 Merge pull request 'Issue 113 resolved by the repin, and what issue 064 did not cover' (#108) from issue/113-record-the-repin-and-fold-114 into main 2026-09-26 12:13:52 +00:00
jochen b5512bed32 Research 017: a mesh that heals itself
The operator's wish written as intended behaviour for the NATS bus:
every loop compares against what is, repairs by the ordinary path, never
destroys, and raises a condition for what it cannot fix. What is done
before NATS is limited to what survives the move.
2026-09-26 00:54:27 +02:00
jochen 0c5eac0218 0110 amends 0109: its seats are provisions, and moving npm takes a holdable seat 2026-09-26 00:44:35 +02:00
jochen b3bd50c588 Issue 120: a provisioner remembers what it did, not what is there
The harness compares against its own memory, so a backend that loses
what was provisioned (the cache's ACL users on a server restart) is
never provisioned again, silently.
2026-09-26 00:44:10 +02:00
jochen e387c4bd0e Apply review: two credentials, staged admin rotation, a ninth provider
The fact-check found mailu, whose user is its mailbox, so 0114 rotates
over two credentials rather than two logins, the adapter choosing what a
credential is. Also: minio keeps non-empty buckets; five backends take
their admin credential only at first init, so single-party rotation is
staged; postgres ownership moves to a non-login role; the harness keys by
consumer; rotation state lives with the vault. Consistency fixes across
0110-0113, 26 and 27; issue 103 resolved by mesh-host PR #22.
2026-09-26 00:38:06 +02:00
jochen 43f63ed41c ADR 0114: a two-party credential rotates over two logins
Graduates research 016. Retiring a login is separated from removing a
consumer, which closes a data-loss path in five providers; single-party
secrets rotate in place; the number of parties decides, not the provider.
2026-09-26 00:22:21 +02:00
jochen 942ebe350f Research 016: survey how each provider can rotate a credential
Overlap as drafted in 0113 would have deleted consumer data: seven of
eight providers name the resource after the login and five drop it on
remove. Rotation is now undecided in 0113 and to-be 27, pending the
survey. Also: a requirement naming a seat resolves to its holder, a
person chooses among remaining candidates at assignment, the controller's
secrets are requirements of its definition, genesis seals to the
control-node key, and moving the vault or broker is break-glass.
2026-09-26 00:14:23 +02:00
jochen 805df3f81e ADR 0113 and to-be 27: rotation overlaps old and new credentials
Decided with the author. A credential is never changed in place: each consumer has two logins, both
derived by the mesh, and uses one at a time. An applier adds the new login beside the old through the
adapter's existing create, and confirms both work; only then are readers released to the new one and
restarted by derivation; only when every reader has confirmed is the old login retired through the
existing remove.

It closes the three cases review found in applier-first rotation: an offline reader keeps working on
the old login until it returns; a bus account's owner keeps its bus until it has moved; a provisioner
restarted mid-rotation is still delivered both values. Nobody is ever without a credential that works,
which replaces to-be 13's all-or-nothing rule with a stronger one.

No consumer module changes. The alternation is the provider loop's. A provider's adapter gains one duty,
giving both logins the same rights over the consumer's data — in postgres, membership of one role that
owns it. The mesh derives two logins per consumer, both within ADR 0049's limit, which 0113 now names
among what it amends. Every rule has a check: overlap, offline reader, bus account, restarted
provisioner, equal rights, login length, and confirmation only once the old login is gone.
2026-09-25 23:50:20 +02:00
jochen 4a1b218706 Seats held by assignments, one assignment per module per node, and 0113's bottom of the stack
Decided with the author:
- A seat is held by one assignment, not claimed by a definition. A definition says which seats a module
  can hold; an assignment says which it does. The store module can run on every node and one assignment
  holds mesh-store; moving a role changes an assignment, never a definition. The foundation's seats name
  what the mesh itself uses and route no consumer — database and amqp consumers use co-location, the
  holder included. This replaces the wrong rationale that the foundation's store is "provider to nobody",
  which contradicted ADR 0078 and to-be 21. 0079's one-postgres rule becomes one mesh-store holder.
- A module is assigned at most once to a node. The instance identity in 0112 and 27 is withdrawn, and the
  login-length problem with it.

Review fixes to 0113:
- The bottom of the stack: the vault is installed as soon as the shared runtime base exists, and genesis
  generates everything needed until then — including the permanent controller's, the control-node
  agent's, the builder's and the broker provisioner's bus accounts, and the controller's store login.
  Genesis creates those accounts until the broker's provisioner runs and adopts them.
- Genesis's values are delivered recorded as the mesh's own, so 0092's never-replace rule for operator
  values does not make them unrotatable.
- Backend-issued secrets (a forge's once-only API token) enter through the vault. Non-module parties
  (the controller's logins, node agents' accounts) are answered the same way, the controller asking on
  their behalf; an enrolment token reaches the controller only as what verifies it.
- A secret with no provisioner to apply it is marked not rotatable by the mesh and refused, instead of
  a restart reported as done. Unused password generators in six provider clients are removed, and a
  catalogue scan checks no module mints.
- Rotation's lock-out cases (offline reader, bus account owner, restarted provisioner) are recorded as
  open, with overlap and re-confirm-with-safeguards as the two answers, to be chosen before acceptance.

0110, 0111 and 26 are marked proposed: they changed in meaning and are under review, and an accepted
record must not rest on proposed ones. To-be 23 and the glossary are restored to main; they change when
these records are accepted.
2026-09-25 23:47:26 +02:00
jochen 1b5f2c2c1a ADR 0113 and to-be 27: address the review of the vault rework
Two decisions taken with the author:
- Genesis delivers and the vault adopts. The vault cannot run first — it is built on the runtime base
  the installation makes after the store, broker and controller, and it learns its work over the bus.
  Genesis generates the foundation's first shared secrets, seals them to the operator key, and
  delivers them to the vault through the path an operator's value takes; from then on the vault holds
  and rotates them. This answers ADR 0085's own reason for rejecting vault-only minting, which 0113
  now names instead of stepping around.
- Rotation re-confirms on every pass. An applier repeats its confirmation until acknowledged, so a lost
  message costs one pass; an applier that stops after applying locks readers out until its supervised
  restart, and that window is stated and shown, not claimed away.

Fixes:
- Scope: a shared secret is made by the vault; a private key (node sealing keys, the operator's key,
  the certificate authority) is made where it is used. The inventory adds the makers the first version
  missed: node and builder broker passwords, and enrolment tokens.
- Broker accounts are created by the broker's provisioner, not the controller, so the controller never
  holds their plaintext; mesh-broker delivers amqp again — one broker per mesh — and only mesh-store
  delivers nothing.
- secret is a reserved provision: only the mesh-vault holder may provide it, and no pin routes around it.
- A secret's contract says whether a recipient applies it or reads it at start; appliers are never
  restarted for it, init-only secrets are applied, and confirmation is to-be 13's standard.
- Operator secrets are one rule everywhere: a secret requirement answered by the vault (0112 no longer
  says otherwise). A data provider's adapter may return fields; the data-return check names a lab consumer.
- 'Holder' now means a seat's holder only; a secret has recipients.
2026-09-25 23:27:58 +02:00
jochen fa2c09a2c5 ADR 0113 and to-be 27: the vault makes every secret — provisioning all the way down
A secret comes into being seven ways today: provider credentials, own secrets (54 modules), broker
accounts through a command that is easy to forget, a vault that only records what the controller
mints (6 modules), operator values, licences, and root secrets. The vault was built to end own secrets
and did not; the old path was never retired.

0113 is rewritten as a waterfall. The vault makes every secret and nothing else does. A provider that
needs a secret for a consumer requires it from the vault, declared once in its provision's contract
and expanded per consumer by resolution; the vault delivers it to both holders, each sealed to its own
node, so a provider's code is unchanged. Own secrets, broker passwords, operator values and licence
credentials take the same path. Genesis is not an exception: it raises the vault first and asks it,
so there is one way a secret is made from the first one on. The vault can sit at the bottom because it
requires nothing but a broker account.

One shared mint function in the SDK was considered and rejected: generation becomes uniform but custody
stays spread over every provider's machine, and each SDK language needs its own implementation.

Rotation is asked of the vault and is provider-first: the value goes to the holder that accepts it,
which confirms, before the holder that presents it gets it, so the lockout window shrinks to the
consumer's own restart, and an unconfirmed provider holds the rotation rather than half-doing it. The
host derives which processes to restart or recreate from the requirement a definition reads, so no
definition declares restart-on for a secret. A rotation shows unconfirmed until each consumer restarted
and passed its health check. Issue 103 becomes a prerequisite.

The file is renamed to match what it now decides. 0112 follows.
2026-09-25 23:10:52 +02:00
jochen 6e3373c879 Design pass: address the review
0113 — the plaintext claim was false under its own mechanism: handing a provider's answer to the
controller puts every secret on the broker and in the controller in the clear. The provider now seals
each secret field itself, to the consumer node's public key the mesh hands it, and the controller
carries sealed fields it cannot open. That is stricter than today, where the controller holds every
minted credential in the clear. Option 3 (plaintext to the controller) is recorded and rejected. The
foundation exception now covers root-secret rotation (0085) and forms like the broker admin's hash, so
no phase claims to remove the broker's bootstrap step. To-be 24 and 13 are named among what it amends.

27 — resolution is consistent with 0110: co-location and the only provider apply only where no seat
delivers the provision, so an unheld seat is refused even with one provider. The secret-field rule now
matches 0086 exactly (a declared env-file, never a container environment value). The seat placeholder
is the controller's, and the one module reading it moves to a host port. Contracts are held by the
controller and written down in phase 1, so they can be checked; every rule has a check. An operator's
secret is still the operator's, with the vault as custodian. Which seats a module holds is listed as
not settled.

0110 — the unheld-seat-with-one-provider case and the one-answer-for-everyone rule have checks; the
claim about moved manifests is corrected. 26 — the table governs and the code catches up, not the
reverse; scope and capacity agree with the glossary; moving a seat is described as it really is today.
0112 — aligned with 27, and lists 0049 and 26 among what it changes.

Issue 118 is renumbered 119: another branch took 118 first. 'Control-plane' is gone from 0110 and 0111.
2026-09-25 22:46:10 +02:00
jochen aad92ea8fe To-be 27 and ADR 0113 (proposed): a module requires, the mesh resolves
The design pass. Everything a module needs is a requirement: a name, a contract, and one of four kinds
of provider — a module, the node's host, the mesh, the operator. Installing a module resolves every
requirement or refuses, naming everything missing at once. It retires six mechanisms that grew
separately: provisions through bindings, settings, assigned ports, machine facts, minted secrets and
literals in the definition.

ADR 0113, proposed: a provider makes what it provides, and the mesh carries it back sealed to the
consumer's node. It is the return path ADR 0048 left "to a separate decision", now needed three ways:
data provisions with nothing to answer with, contracts needing a value the controller cannot make, and
a vault that generates nothing. Who a consumer is stays the mesh's (ADR 0049). Genesis is the one
exception. On acceptance it supersedes 0048 and amends 0085.

ADR 0112 is revised from three sources to that single concept.

ADR 0110 is amended for two points raised in review. The vault gets the mesh-vault seat (issue 106).
A seat's holder outranks co-location for a provision it delivers. Writing that down exposed an
inconsistency: mesh-store delivering postgres-database would have sent every database consumer to the
control-node, against to-be 23's node-local stores. So a seat delivers a provision only where the mesh
has one answer for everyone — artifact store, npm registry, git, vault — and mesh-store and mesh-broker
deliver nothing. 23 and 26 follow.

'Control plane' becomes 'controller' in the records written today.
2026-09-25 22:34:46 +02:00
jochen 7668190154 ADR 0112 and issue 118: address the review
- Secrets follow ADR 0085 as amended: a module's own secret is a provision the controller mints and
  the vault records. The previous commit had that backwards. Whether the vault should generate
  instead is recorded as an open question, not decided.
- A directory's contract is owner and mode only. The persistence flag was the keep flag ADR 0030
  refused; a directory is kept while it holds anything, and disposable data is a named volume (0107).
- An operator's shared data stays an access (ADR 0051), which rejected an operator-owned directory.
  Only where its path is written moves to the assignment.
- The records it changes on acceptance are named: 0051, 0091, 0046 (settings keyed by instance),
  0084 (a provider is a node and an instance), and the glossary, which gains its new words only
  when the record is accepted.
- How it is checked covers every stated rule. Container-side paths are no longer flagged by the
  host-path rule, and code fallbacks are covered.
- Provisions are what other modules provide. A seat's occupant is not listed as one, and the vault
  is not described as selectable per assignment.
- 'Control plane' becomes 'controller'. The provider count is ten of eleven, not eleven of twelve.
2026-09-25 22:29:38 +02:00
jochen 90fb7ae4ad ADR 0112: a module's own secrets are a provision from the vault, not something the mesh generates
The first draft listed minted secrets under what the mesh generates. ADR 0085 made a module's own
secret — a password, an internal token, an external key it was handed — a secret provision answered
by the vault, like a database by the store. What the mesh still mints is the delivery credential for
each provision a module takes (ADR 0048), the vault's own included.
2026-09-25 22:29:38 +02:00
jochen ca235e775f Issue 118 and ADR 0112 (proposed): a module definition names no node, no mesh and no path
Issue 118 records what a review of where module code reads its files found: 789 host-path strings
in 70 of the catalogue's 71 definitions, every one a decision the definition makes about a machine.
Mounts are checked (ADR 0091); the same paths retyped as values are not. It records what that has
already allowed — a DNS provider that would provision nobody silently, a contributions file that
names credentials by host path and so forces every provider to mount at the identical path, an SDK
loop that treats an unwritten contributions file as empty without a word, defaults in code that
disagree with their own manifests — and that no module can be assigned to one node twice, because
every identity is keyed by the module's name.

ADR 0112, proposed for review, answers it the way ADR 0038 answered ports: a definition names
variables, and installing it resolves every one or refuses, from three sources — the assignment's
own configuration, provisions the mesh resolves against a contract, and what the mesh generates or
knows. A directory becomes a provision: the module requires one by name with its owner, mode and
persistence, and where it lands is the assignment's. The mesh's own files stop carrying host paths.
An assignment gets an identity of its own, so a module may run twice on one node.

Checking copies for agreement was rejected as checking something that should not exist; rewriting
paths per assignment was rejected as inferring which strings are paths by their shape. Syntax, a
node's default layout, and when a second instance becomes possible are left to the design.
2026-09-25 22:29:38 +02:00
jschoubben 74ae0609cb issue 118: umami's store answers the dial and times out the query 2026-09-25 22:01:28 +02:00
jochen d94fe8f638 To-be 26: name the files that implement the seats and the build source 2026-09-25 20:47:48 +02:00
jochen c4cac767f8 ADR 0110: admit the-private-network, claimed by a manifest the control plane composes in code
The enumeration behind the first set read manifests in two repositories and missed a claim made
in the control plane's own code: the private-network module it ships claims the-private-network at
node scope. A closed set without it would refuse the control plane's own module. Thirteen claims in
use, naming twelve seats.
2026-09-25 20:35:31 +02:00
jochen dbe100ca96 ADR 0110 and 0111: a seat is a module assignment from a closed set, and a build source may live on the git seat
Seats have been doing two jobs and neither is written down. The mechanism ADR 0009 introduced is
enforced — a second holder is refused — but any well-formed name becomes a seat by being claimed,
and nothing can say which seats a mesh has or who holds them: holdings are assembled while planning
and discarded. The enumeration done while preparing this missed the control plane's own manifest,
because core modules' manifests live in its repository rather than the catalogue.

0110 closes the set. Each seat has a name, a scope, what occupying it delivers, and the record that
made it one; a claim outside the set is refused. A seat is held by a module assignment, and what the
mesh knows about the holder is what it knows about that assignment — nothing is stored beside it. A
seat may deliver a provision, and then its holder answers for it among several providers: pin, then
the holder, then the only provider, then refused. That keeps 0009's "refused, never guessed": the
seat is the choice made once, mesh-wide, instead of a pin per consumer node. The first set is the
eleven seats already claimed plus 0109's npm-package-registry, so nothing in use is refused.

Two concepts — seats for exclusion, a new word for consumable singulars — was rejected: both mean
"this mesh's one X", and the overview a person wants is one list.

0111 gives the mesh a git seat and makes a build source one of two explicit forms: a repository on
the seat's holder, recorded by its path and cloned from wherever the holder runs at build time; or
an external URL, recorded and cloned exactly as given. Recognising self-hosted sources by matching
URLs against the forge's address was rejected — it fails in the one case it exists for, after the
forge moves. Credentials for private repositories are left undecided and said so.

Design: new to-be 26 (the seats); 23 gains the seat step in resolution; 18's source entry names
the two forms; the glossary's seat and provision entries say where they meet. 0109 is carried from
its own branch so every link here resolves.
2026-09-25 20:33:14 +02:00
jschoubben 59c93dcfe4 109: a package registry seat is one per ecosystem, not one for all of them
Extends ADR 0075. Surfaced fixing builder's hand-faked package-registry
binding tonight: gitea's manifest declares the provision once with a single
npm-path, conflating what should be independently assignable per ecosystem
(npm/cargo/docker/...) the same way artifact-store and package-registry
were themselves split. Cited in 22-the-work-ahead.md's Phase 2, where the
target state this decision points at was already described a week ago.

Numbered 109, not 108: route-proxy's policy feature (mesh-controller PR
still-unwritten decision record — reserved but never committed. Renumbered
around it rather than colliding.
2026-09-25 20:29:22 +02:00
jschoubben 2ca63ae54e 117: builder's real package-registry grant deadlocks a genesis bootstrap
Fixing builder's hand-faked package-registry binding tonight (requires:
package-registry, a real mesh grant instead of a hardcoded JSON fragment)
broke three tests describing a deliberate carried-binding fallback for
exactly this: gitea's own image is built by builder, so builder cannot
yet hold a real grant from gitea the first time either has to exist.
Invisible on novox (already bootstrapped, gitea already live) — real on
any genesis from scratch. Fix left in place, tests left failing rather
than reverted or hacked, so the gap stays visible.
2026-09-25 16:59:04 +02:00
jochen 82a6badc7c Issue 114: land the controller's container-or-process question, renumbered
Filed 2026-09-24 on a branch of its own and never merged, numbered 113, which is taken. 114 is
free because a sibling branch folded it, so it takes that number and keeps its commit.

Kept separate from issue 117 rather than folded into it. 117 asks the same question of every
module and locates the missing decision; this asks it of the controller, where `network: host`
means container network isolation — the property that resource type usually buys — is not in use.
That observation is this report's own and is nowhere in 117, and folding would lose it.

Its first open question is answered by 117's diagnosis and now says so: the host's `process` shape
is built, applied and tested, restart and run-to-completion semantics included, so deciding this
does not wait on host-side work.
2026-09-25 16:16:21 +02:00
jschoubben 10a2b706c6 Issue 113: should the controller be a container or a process the host supervises
Filed after a session where every mesh-controller interaction went through
docker exec — its manifest runs it as a container with network: host, using
none of the isolation that resource type usually buys, while ADR 0006 makes
it the mesh's single point of coordination. Open question, not a claimed
defect: does type: container get the controller anything type: process
(supervised the way the host supervises its own unit, per ADR 0005) would not.
2026-09-25 16:15:27 +02:00
jochen 35db2aaa41 Issue 117: a module's own code is a container in one record and a process in another
Asked what the "sidecar" is and whether a supervised process would do instead. The repository
answers both ways. ADR 0047 (accepted, unsuperseded) says a module with tools or events runs a
container carrying its compiled code. To-be 18 and 20 (both proposed) define a `process` resource
type — the module's own code, a unit the machine's supervisor keeps up — and the worked guide says
plainly "it is why these are `process` rather than four containers." Neither design doc names 0047,
and no decision record mentions a `process` shape at all.

Diagnosed rather than left open, because the ground truth settles what the report could not.
The shape is real: mesh-host defines TypeProcess, applies it, and tests it, and the host's
vocabulary is twelve shapes rather than the nine ADR 0029 counted. So the alternative the report
offered — that two proposed documents describe a type that does not exist — is disproven.

ADR 0029's mechanism is intact and was not enough. The vocabulary-count test names the decision
behind each addition: network 0029, access 0051, opening 0100. The eleventh names a *proposed
design document*, and TypeProcess is the only shape in the vocabulary whose doc comment cites no
ADR. Requiring every addition to name something does not require it to name a decision.

The argument this issue asked for already exists — as a Go test comment. "It is a full-host shape
rather than a portable one: it needs a process supervisor to install into. It does NOT need a
container runtime, which is the point — only software that genuinely needs isolation asks for a
container." That is a decision's context and consequences, in another repository.

What the catalogue does is a third thing: 115 container declarations against 3 process, all three
in showcase — the module to-be 20 documents. There the tools resource is a container running
`sleep infinity` on a bare upstream base with the broker credential mounted, and the tools and
provisioner entrypoints are run by nothing. That is the condition 0047 was written to end, back
in a new shape.

Where the isolation argument leaks is narrower than expected and worth having precisely: the
serving key and the credential shape both conform. But serveTools serves every registered module
over one broker connection, the runtime takes its modules from a comma-separated list, and
x-source is stamped from the single credential — so two modules in one runtime means the second's
events are attributed to the first. Nothing refuses it and no test asserts against it.

Located on hq rather than on a code repository: the implementation and the design layer agree,
and the missing thing is the record. Which shape is right is left open, deliberately — this
establishes that the question was answered in practice and never written down, not which answer
is correct.

One correction kept in the trail: the first search here was for len(Vocabulary()), found nothing,
and was two steps from being written up as "the mechanism ADR 0029 relied on is gone." The test
binds the slice to a local first. A negative search result read as a fact about the world is the
same error issue 113 recorded.
2026-09-25 16:15:00 +02:00
jschoubben 56669ee23b Merge pull request 'Issue 116: route-proxy has no authentication or IP-restriction mechanism' (#109) from issue/116-route-proxy-has-no-auth-or-ip-restriction into main 2026-09-25 12:28:50 +00:00
jochen 367df38e6d Issue 116: resolved by mesh-controller PR #58
The gap is closed in the proxy: policy applies, the four capabilities exist, the table is keyed
by host and path with a total ordering, and the two failure modes that rot quietly are held by
tests — a declaration carrying a credential refused rather than served, an unreadable secret
failing closed.

Resolved rather than left open because the issue reports a gap in the proxy and that gap is
gone. But the record says plainly what it does not yet allow: an operator still cannot move the
affected routes, because that needs the mesh side — a manifest able to declare these values and
the controller minting the secret auth names. Until both exist the capability is reachable only
by writing the routes file by hand. That is the ordinary build-out of a contract this issue's
decision created, and it belongs to to-be 08 rather than here.

The open questions are marked answered and kept rather than deleted, pointing at ADR 0108 —
what was rejected and why is the half worth having, and a section still saying "the fix should
not be written before these are answered" after the fix was written reads as though nobody
looked.

One finding kept in the record: priority was read with the reader for ports, which caps at
65535, and the one real rule this reproduces is declared at 100000. It parsed to zero, so
refusal and path scoping would have shipped looking complete and doing nothing on the only case
that motivated them. A validator borrowed from a neighbouring field is a silent default.
2026-09-25 14:25:24 +02:00
jochen a11da86591 ADR 0108: a route carries the policy applied to a request
Issue 116 found the mesh's proxy applies nothing to a request — host lookup, forward. Against
what the replaced ingress actually relies on, four capabilities are missing: authentication
(three dependents, each gating an admin surface with no login of its own), refusal scoped to a
path (one, a live incident mitigation), path-scoped routing with priority, and redirect.

Policy goes on the route rather than beside it. A proxy-side settings layer keyed by route name
would keep the grant literally clean, but then "what protects this route" is answered from two
files nothing keeps in step — and a route's protection is part of what a route is.

The set is closed at those four, so a fifth is an amendment and each addition is earned by a
dependent that exists. An open middleware surface was rejected: it recreates what is being
replaced, and narrowing one later is far harder than widening a closed one.

Where policy needs a credential the declaration names a secret and never carries the value,
which keeps the existing secret machinery the only thing holding credentials. Inlining a hash
was rejected as the first credential in a declaration — a precedent easier to set than withdraw.

This re-keys the routing table by host and path with priority, which follows from the decision
rather than being a separate one: two of the four need one host routed more than one way. Equal
priorities must resolve identically every time or the proxy stops being reproducible.

The record says how it is checked, including the negative case that rots quietly — a
declaration carrying a credential value rather than a reference must be refused, so the
rejected option cannot return by accident.

08-connectivity §3 names the record and gains the subsection; issue 116 gains amended-design.
2026-09-25 13:48:18 +02:00
jochen c839d9ac26 Issue 116: scrub the disclosure, and correct the count and the shape of the gap
Two things the report got wrong, and one it could not have found the way it looked.

Disclosure first: it carried a real hostname and an absolute node path, in a public
repository. Both are gone; the ingress, the modules and the routes are named by role, as the
rest of 04-ISSUES does.

The count was low. Basic authentication has three dependents in the catalogue, not one — the
key-value store's browser UI, a database web UI, and the ingress's own dashboard. All three
are credential-less admin surfaces whose only gate is a middleware the mesh's proxy lacks.
The earlier version read only the node's dynamic configuration directory, which cannot see
what modules declare as container labels; counting needs both sources, and the report now
says so.

Two gaps were missing entirely. Redirect rules: two live routes canonicalise a www name onto
its apex, they exist only on the node and not in the catalogue, and they fail silently rather
than erroring. And path-scoped routing with priority, which is the one that reorders the
issue: the table maps host to exactly one target, so a host cannot be routed two ways, and
the refusal rule matches a path on a host already routed elsewhere. Authentication and a
source filter would not make it expressible. Path scoping is a prerequisite, not a sibling.

Also corrected: the refusal rule was described as an address-scoped deny. It is an allow-list
holding a single documentation-range address — deny-everyone — so reading it as address-scoped
points at the wrong fix. And its severity was understated: its own header records it as
incident response closing an abused write primitive, which is not "a real exposure" but a live
mitigation.

The open questions now say plainly that they are design questions and the fix should not be
written before they are answered, and one is added: whether a declaration may carry a
credential at all.
2026-09-25 13:24:34 +02:00
jschoubben 226d556743 Issue 116: route-proxy has no authentication or IP-restriction mechanism
Comparing route-proxy against what HAL's actual Traefik config does today,
not Traefik's general feature set, per the standing rule that the nox mesh
must do at minimum what the HAL mesh it replaces already does. Everything
else checked out even or better; these two are real, confirmed gaps —
RedisInsight has no login of its own and depends entirely on Traefik's
basicauth middleware, and the gitea-internal route depends on an IP-scoped
deny rule. Neither has any equivalent in route-proxy's single-lookup
request path.
2026-09-25 11:51:56 +02:00
jochen ab7d216fc6 Issue 113 resolved by the repin, and what issue 064 did not cover
The object-store module was repinned to a maintained fork of the withdrawn server image, its
runtime sidecar built rather than pulled, and its data moved off the predecessor's live
directory. The instance is closed; the three general points the report makes are not, and What
was done says so rather than letting a resolved status imply otherwise.

Folds in the one thing a duplicate report of this symptom had that this one did not: issue 064
asked whether the build environment can reach a declared vendor image and assumed that, once
declared, it stays fetchable. Withdrawal is the case that assumption does not cover. The
duplicate is not merged — it carried the reading this report's diagnosis retracts.
2026-09-25 00:55:13 +02:00
jschoubben cdcd4da27e Merge pull request 'Research 015: reopen the comparison — the premise for narrowing to one candidate was false' (#105) from storage/015-reopen-the-candidate-comparison into main 2026-09-24 16:47:18 +00:00
jochen 1d524a1fa8 Research 015: rewrite the comparison — wrong axis, and a missing candidate
The previous version ranked candidates on whether they preserved single sign-on to the
object store's console. That is not a requirement: a "user" of the store is normally an
application, so the requirement is per-application keys scoped to buckets — which the mesh
already mints. And the console login it ranked on never worked; the module's own hook comment
records "policy claim missing", a failing login written up as progress.

It also omitted the incumbent's own maintained fork, which changes the question from "which
product replaces it" into two decisions: repoint, or migrate — and if migrating, to which.
Repointing costs an image reference; migrating costs a data copy, two handler rewrites and a
maintenance window. Repointing does not foreclose migrating, which is the argument for taking
it first.

On the corrected requirement Garage ranks first — its per-key-per-bucket model is the
requirement verbatim, its admin API matches how the mesh provisions, and the highest-risk
consumer is first-party documented against it. Its remaining gap (no versioning, no
server-side encryption, partial lifecycle) is unmeasured against the buckets and is the one
thing that could still disqualify it.

Measured and folded in: 230 GiB logical, 82,496 objects, 468 GiB raw at 2.03x, eight drive
directories on one filesystem on one machine. That last fact decides more than any feature —
the erasure coding is not buying independent-drive redundancy, so the redundancy model is
close to irrelevant and only storage overhead remains, which at this volume is a rounding
error against the headroom.

Both errors are recorded at the end of 01 rather than quietly fixed. A configured feature is
not an observed one; and when a dependency dies, "who took it over" precedes "what replaces
it" — searching for alternatives by construction returns things that are not the incumbent.
2026-09-24 18:46:47 +02:00
jochen 999636e2a8 Issue 113: retract the diagnosis table — the original report was right
The diagnosis carried a table headed "claims that could not be substantiated", denying a
module.json, a digest pin, and an all-zeros runtime digest. All three exist. The table is
withdrawn in full and replaced with what is actually true, plus the two claims that remain
genuinely unverified rather than disproven.

The cause: one repository was searched and absence in it was written up as absence. The
catalogue of the mesh being built is a separate repository, not checked out where the search
ran, and all four claims were about that repository. Compounding it, the predecessor's
object-store module and the one being cut over to were treated as one thing — they are
different files in different repositories, one pinning a tag with no sidecar, the other a
digest with two container resources.

Also corrected in the report: located-in named the wrong repository; the "pins a tag" passage
described the predecessor; the open question about pinning by digest is struck, because this
module already does and it made no difference — a deleted digest resolves to nothing either
way. The section on why nothing broke is now scoped explicitly to the predecessor's
machinery.

The lesson kept in the record: "zero occurrences anywhere in the tree" is only as strong as
the tree searched, and a diagnosis must say which tree. A confident rebuttal of a correct
report is worse than no diagnosis — it sends the next person to the wrong place with a
written record behind them.
2026-09-24 18:46:47 +02:00
jochen 48abc36b5b Merge remote-tracking branch 'origin/main' into work/object-store-records 2026-09-24 18:41:53 +02:00
jschoubben 2c5f805467 Merge pull request 'Add hq-defer: park a thought without moving the work off course' (#107) from meta/hq-defer-skill into main 2026-09-24 16:39:22 +00:00
jochen db2c950ba6 hq-defer: make the MEMORY.md pointer an explicit placeholder
Review caught it reading as a real relative link, so a link checker flags
.claude/skills/hq-defer/file.md forever. Angle brackets say placeholder.
2026-09-24 18:38:43 +02:00
jochen 971d0839f5 Add hq-defer: park a thought without moving the work off course
A thought raised mid-task needed remembering but not working on, and there was no
mechanism for that — so it was recorded by hand. This is that, made repeatable.

Records to Claude's persistent memory rather than the repository, deliberately. A parked
thought has no number, owner or status: giving it one asserts triage that deferring says
has not happened. A shared "deferred" document would be a central status file, which
AGENTS.md forbids. And a repository write means a branch, a commit and an MR — the drift
the skill exists to prevent.

Wraps no playbook, because deferring precedes the development cycle rather than being part
of it. It does say which playbook a thought would need if it graduates, and that recording
"undetermined" is the honest answer when the evidence does not say.

The stop condition is the substance: at most two lines, then return to what was in
progress. No plan, no triage question, nothing opened.
2026-09-24 16:40:24 +02:00
jschoubben 5677e97508 Merge pull request 'ADR 0107: persistent data is a directory bind, never a named volume' (#106) from decide/0107-persistent-data-is-a-directory-bind into main 2026-09-24 14:24:17 +00:00
jschoubben a93743708c ADR 0107: persistent data is a directory bind, never a named volume
Records the rule the operator gave directly, mid-session, after checking
that HAL's own postgres and lavinmq both used a directory bind and the
mesh's adoption of them three weeks ago switched to a named volume without
a reason recorded anywhere.

Already built and rolled out on novox (mesh-catalog PR #54) before this
record -- urgent enough to fix first and write down after. Includes the
incident: the new host directories needed the container's own UID, which a
named volume gets for free and a directory bind does not; mesh-store
crash-looped on Permission denied until ownership was matched to what the
original volume already had.

Closes issue 115. Checks pass.
2026-09-24 16:21:51 +02:00
jochen c8cbbcfb8a Research 015: reopen the comparison — the premise for narrowing to one candidate was false
SeaweedFS was scoped as primary because it looked like the only candidate preserving
OIDC console login. Measured: its admin UI is Apache-2.0 but its identity-provider
integration is not — console SSO sits behind the per-TB commercial licence, alongside
point-in-time recovery and automatic EC repair. The free build gives OIDC on the S3 API
via STS and a console authenticated by local username and password.

So the answer to the gating question is that no candidate preserves the current feature
set for free, which this effort had written down as a possible outcome. Reopened across
three candidates with the requirement-by-requirement evidence in 01.

Two corrections to what the overview recorded. RustFS is not a binary-level drop-in
retaining existing data: API and on-disk compatibility are separate paths and the on-disk
one is preview-scoped. And it carries an open defect in the credential path the bucket
provision depends on, which gates it specifically.

Nothing graduates before two measurements named in 01: whether an authenticating proxy
is an acceptable answer to console SSO, and which S3 endpoints consumers actually call —
the latter because Garage does not implement the full span and cannot be ranked until
that is counted.
2026-09-24 16:10:33 +02:00
jschoubben a458f751c6 Merge pull request 'Accept ADR 0038; close issue 091' (#104) from issue/091-ports-are-mesh-assigned-not-manifest-fixed into main 2026-09-24 13:53:36 +00:00
jschoubben 402b798ab6 Accept ADR 0038; close issue 091
The decision (the mesh assigns a container's machine-side port; a module
says only what it needs) was proposed 2026-09-01, and the machinery
already implements it in full -- internal/inventory/ports.go's PortFor,
declaration.go's publishedOn. What was missing was the catalogue actually
complying: 14 of 46 modules baked a machine-side number into their own
manifest anyway. mesh-catalog PR fixes 11 of them (the two defensible
kinds -- foundation, protocol-fixed -- are left alone, per the issue's own
categories). Accepting the decision now that it's actually enforced, and
closing the issue it was blocking.

Checks pass.
2026-09-24 15:52:45 +02:00
jschoubben f10dce4f9e Merge pull request 'Issue 113 and research 015: the object store's images are gone upstream, not access-restricted' (#103) from storage/113-the-object-store-lost-its-upstream into main 2026-09-24 13:45:34 +00:00
jochen 4beb6629db Issue 113: ground the rebuildability point in what the design actually says
Review of my own text found an unattributed claim — "the mesh's claim that a node
can be rebuilt from its declarations" — which is not a stated principle anywhere.
Replaced with the design position that genuinely covers it: to-be 07 chooses
references over payload because "reproducibility comes from pinning the identity of
a thing rather than carrying its bytes". This incident is that choice's failure mode
when the identity stops resolving, which is a sharper point than the one I made.

Scope stated honestly: the passage is about the foundation bundle and this module is
not in it, but pin-identity-fetch-bytes is how every module gets third-party images.

Also names the tension the mirroring question actually carries — mirroring is a move
away from references-over-payload, so it is a decision, not a fix.
2026-09-24 15:44:44 +02:00
jochen d497b37e43 Issue 113 and research 015: the object store's images are gone upstream, not access-restricted
The symptom arrived diagnosed as "the registry disabled anonymous pulls for the
whole vendor namespace". It did not hold: sibling repositories in that namespace
pull normally, the "$disabled" token field appears on every repository including
working ones and describes signing rather than access, and "actions": [] with a
401 is byte-identical to what an invented repository name returns. Both registries'
own APIs establish deletion instead.

Recorded because the correction is the expensive part to rediscover, and because
the instance was harmless while the standing condition is not: no node that does
not already hold the images can ever provision the module again, and nothing
detects that until one tries.

Research 015 scopes the replacement. It is not a redesign — the foundation design
already commits to S3 the protocol rather than the product, and the object store
is an ordinary module, so this instantiates an existing principle. The live OIDC
wiring is the requirement that gates the choice, and it is checked first.
2026-09-24 15:27:03 +02:00
jschoubben 183b22997c Merge pull request 'Issue 112: diagnose — the predecessor's own DNS config already names the carried peers' (#101) from issue/112-diagnosis into main 2026-09-24 12:49:50 +00:00
jschoubben 8d67cf63c5 Issue 112: status located, not diagnosing
Playbook 03 step 2: move status to diagnosing, then located once the
owner is known. located-in is filled with four confirmed packages —
the owner is known.
2026-09-24 14:27:25 +02:00
jschoubben 75c104c355 Issue 112 diagnosis: correct located-in attribution
The carried-peer record (CarriedPeer/TunnelPeer) lives in mesh-controller
internal/inventory, not internal/catalogue. internal/catalogue is the
right package for the zone-generation side of the fix (facts.go's
nodeZones), but a different concern from where the name field itself
would go. Split the two so a decision record doesn't get pointed at the
wrong package.
2026-09-24 13:54:20 +02:00
jschoubben 6abfec7433 Design 25: address first review's four findings before any code
Fixes, each named where it was wrong:

- reload-on is a service field; a container only has restart-on, which
  recreates. Cited precedent (registry-trust-reload) is a service resource,
  not a container. Fix: nats-server's own SIGHUP reload, triggered by an
  in-image entrypoint watching a directory-mounted config file (issue 103's
  recreate-on-change applies to a directly-mounted file, not a directory's
  contents) -- asks nothing new of the host.
- A JetStream-delivered message's Reply field is already claimed by the
  consumer's own ack address, so a responder using it answers nobody. Fix:
  every CONTROL message needing a reply carries its reply subject in its own
  payload; the controller publishes there explicitly, never via Respond().
  Enrolment is the case this design actually depends on, so it's fixed there
  too, not just noted.
- The listed permissions never granted publish on a durable consumer's own
  ack-reply subject -- a module could receive but never ack, so every
  message redelivers forever. Fixed with a scoped grant per module's own
  consumer.
- One account (a deliberate choice, kept) means inbox privacy is the
  permission list or nothing. The design granted 'its reply inbox' without
  scoping it, which read as any user reaching any inbox. Fixed: each user's
  inbox prefix is derived from its own identity and its permissions name
  only that prefix.

New open question from this revision, not closed: whether the in-image
watch-and-SIGHUP shape belongs in mesh-sdk if a second module ever needs it.

Checks pass (records.py, cycle.py, index.py).
2026-09-24 13:53:53 +02:00
jschoubben 6a56738d7b Issue 112: diagnose — the predecessor's own DNS config already names the carried peers
Checked why ADR 0104's forward-to-predecessor shape doesn't transfer to the
resolver the way it did the proxy: DNS is one process on one port, and
assigning the mesh's dnsmasq module replaces it in place, so there is no
predecessor process left standing to forward to.

But /etc/dnsmasq.d/hal-dns.conf's static address= lines for ace/shanks/g14
match the mesh's own carried-peer addresses from overlay show exactly. The
name a carried peer needs isn't a guess the operator has to make under
pressure — it's a transcription of a record the predecessor already has and
has been correctly serving for six days. Located in mesh-controller (no way
to attach a name to a carried peer today) and the dnsmasq module (doesn't
emit a wildcard for a named-but-uncarried peer). Not implemented.
2026-09-24 13:25:44 +02:00
jschoubben 91a5c63d65 Merge pull request 'Name the migration repository in the map, so nobody has to be told it exists' (#99) from meta/name-the-migration-repository into main 2026-09-24 00:02:07 +00:00
jschoubben 545d038198 Name the migration repository in the map, so nobody has to be told it exists
hq cannot hold the migration's operational record — it names machines, addresses and
paths, and this repository is public — but it can say where that record is, which is what
this map is for. Asked for by the operator, who had to be told.
2026-09-24 02:01:41 +02:00
jschoubben f3ad60b98c Issue 103 resolved by mesh-host #22 2026-09-23 23:44:31 +02:00
jschoubben 0c433d51ad Design 25: the bus on NATS — subjects, streams, accounts as configuration, enrolment, a person's client, the cutover, the beds
The architecture ADR 0106 asks for, proposed for review before any code.
2026-09-23 23:42:42 +02:00
88 changed files with 9497 additions and 66 deletions
+78
View File
@@ -0,0 +1,78 @@
---
name: hq-defer
description: Use when a thought is raised that should be remembered but NOT worked on now — an aside during other work, a "we should look at X someday", a known gap nobody is assigning yet. Triggers on "defer this", "park this", "register this thought", "note this for later", "don't work on it, just remember it". Records and returns to whatever was already in progress.
---
# hq-defer
Parks a thought so it is not lost, **without moving the work off course.** The defer is the
point: the thought is recorded and the previous task resumes.
This skill wraps no playbook, because deferring is not part of the development cycle — it is
what happens *before* something enters it. A parked thought has no number, no owner and no
status, and that is correct.
## Where it goes, and why not the repository
Record it as **one memory file** in Claude's persistent memory directory for this project (the
path is given in the session's memory instructions), with `metadata.type: project`, plus a
one-line pointer in `MEMORY.md`.
**Not** in `04-ISSUES`, `01-RESEARCH` or anywhere else in the repository:
- A parked thought is not an issue or a research effort. Giving it a number asserts it has been
triaged, which is exactly what deferring says has not happened.
- A shared "deferred" or "someday" document is a **central status file**, which
[`AGENTS.md`](../../../AGENTS.md) forbids. Status lives in frontmatter on real records, and a
parked thought has no real record yet.
- A repository write means a branch, a commit and a pull request — drift, which is the one thing
this skill exists to avoid.
## Steps
1. Write the memory file. Slug is kebab-case and descriptive of the thought, not of the act of
deferring.
```markdown
---
name: <kebab-slug>
description: Deferred note — <one line>
metadata:
type: project
---
Raised and deliberately deferred on YYYY-MM-DD: **<the thought, in the user's own terms>**
**Why:** what was being worked on when it came up, and that deferring was intentional so
that work was not pulled off course.
**How to apply:** treat as an open thread, not an assignment. Do not start on it
unprompted. If it graduates it needs an HQ home first — an issue under playbook
[03](../../../00-META/process/03-issues.md) if a stated behaviour does not happen, or
research under playbook [01](../../../00-META/process/01-research.md) if it is still an
idea. Say which is undetermined, if it is.
```
2. Append one line to `MEMORY.md`: `- [<Title>](<kebab-slug>.md) — deferred YYYY-MM-DD; parked, no HQ
record, do not start unprompted`.
3. Convert relative dates to absolute before writing. "Last week" is worthless in six months.
4. Check for an existing memory covering the same thought and update it instead of adding a
duplicate.
## Then stop
Reply in **at most two lines** — what was recorded, and that it is parked — and **return to
whatever was in progress before.** Do not summarise the parked thought back at length, do not
propose a plan for it, do not ask which playbook it belongs to, and do not open anything.
If nothing was in progress, say only that it is recorded.
## Do not
- Do not create an issue, a research effort, a decision record or a design document.
- Do not create a branch, commit or pull request.
- Do not start investigating the thought, however cheap the first check looks.
- Do not name nodes, domains, addresses, absolute paths or usernames in the memory file — the
thought may later be quoted into this repository, which is public.
- Do not decide whether it is an issue or research when the evidence does not say. Recording
"undetermined" is the honest outcome and costs nothing later.
+71
View File
@@ -284,6 +284,76 @@ def check_numbering(failures, records):
)
def check_progressive_insights(failures, records):
"""A correction made inside a record is marked and dated, or it is a silent rewrite.
A record may be corrected in place when a *fact* in it went stale and the decision still
stands (`02-DECISIONS/README.md`, "Progressive insight"). The whole safety of that allowance
is that the correction is legible in the record rather than only in a diff nobody reads, so
the form is what is checked here: every mention of an insight is the marker, the marker
carries an ISO date, and that date is not earlier than the decision's own — an insight
predating the decision it corrects is a copied marker, not a correction.
What this cannot check is an edit made with no marker at all. Nothing mechanical can; that
one is the reviewer's, reading the diff. The check keeps the *marked* path honest so that an
unmarked change stands out as the anomaly it is.
"""
phrase = re.compile(r"progressive insight", re.I)
# Both patterns stay on one line: a bold run does not span paragraphs, and `[^*]*` across
# newlines will happily join an unrelated `**` far above to the marker below, reporting the
# whole span between them. It did exactly that the first time this ran.
# Trailing words after the date are allowed — "— 2026-09-26, correcting the one above." — so
# an insight can say what it relates to. Only the date's presence and position are fixed.
marker = re.compile(r"\*\*Progressive insights?[ \t]*[\u2014\u2013-][ \t]*(\d{4}-\d{2}-\d{2})[^*\n]*\*\*")
loose = re.compile(r"\*\*[^*\n]*[Pp]rogressive insights?[^*\n]*\*\*")
iso = re.compile(r"^\d{4}-\d{2}-\d{2}$")
for number, record in sorted(records.items()):
text = record["text"]
if not phrase.search(text):
continue
decided = str(record["front"].get("date", ""))
good = [(m.start(), m.end(), m.group(1)) for m in marker.finditer(text)]
for m in loose.finditer(text):
if any(s <= m.start() and m.end() <= e for s, e, _ in good):
continue
# A bold run carrying a link is discussing an insight — usually another record's —
# rather than marking one. A marker never needs to cite anything.
if "](" in m.group(0):
continue
failures.add("insights", rel(record["path"]),
"a progressive insight is not in the dated marked form "
"'**Progressive insight \u2014 YYYY-MM-DD.**': %s" % m.group(0))
for _, _, stamp in good:
if decided and iso.match(decided) and stamp < decided:
failures.add("insights", rel(record["path"]),
"a progressive insight dated %s predates the decision (%s)"
% (stamp, decided))
covered = [(s, e) for s, e, _ in good]
for m in phrase.finditer(text):
if any(s <= m.start() and m.end() <= e for s, e in covered):
continue
line = text.rfind("\n", 0, m.start()) + 1
end = text.find("\n", m.end())
whole = text[line:end if end != -1 else len(text)]
if whole.lstrip().startswith("#"):
continue
# A line that also carries a link is discussing the rule, not marking a correction:
# a marker never needs to cite anything, and a record that reasons about the policy
# must be able to name it. Bare prose with no citation is the informal marking this
# is here to catch.
if "](" in whole:
continue
if loose.search(text, line, text.find("\n", m.end()) + 1 or len(text)):
continue
failures.add("insights", rel(record["path"]),
"'progressive insight' appears unmarked; a correction is marked and "
"dated, or it is a silent rewrite")
def check_status_against_code(failures):
"""A design document naming specific code may not still call itself `designed`.
@@ -325,6 +395,7 @@ def main():
check_numbering(failures, records)
check_topics(failures, records)
check_status_against_code(failures)
check_progressive_insights(failures, records)
print(f"records: {len(records)} decision records checked")
return failures.report()
+27 -6
View File
@@ -31,8 +31,22 @@ another — and a mesh you cannot name precisely is a mesh two people describe d
- **store** — the one postgres server. It holds the controller's own context databases
(`inventory`, `identity`, `licences` — a context owns its store, [ADR 0008](../02-DECISIONS/0008-a-context-owns-its-store.md))
and every module's own database. One server, many databases — never one shared "mesh database".
- **broker** — the one lavinmq message bus. It carries the mesh bus on the `/` vhost and a vhost per
consumer that requires `amqp`.
- **bus** — the mesh's own nervous system: NATS, one per mesh, carrying every link the mesh has —
control, declarations, builds, events, tool calls
([ADR 0106](../02-DECISIONS/0106-the-bus-is-nats.md)). A module reaches it by requiring
`mesh-bus` ([ADR 0128](../02-DECISIONS/0128-the-mesh-bus-is-required-not-ambient.md)); one that
does not require it has no account on it. Held by the `mesh-broker` seat, which is named after
the *role* rather than the server, so the server can change without the seat doing so.
- **the deprecated broker** — the lavinmq module. It was the mesh's bus and is not any more. It
keeps running as an **ordinary provider** of the `amqp` provision, for modules that need a
message broker of their own the way something needs a database
([ADR 0127](../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md))) — no seat, not foundation,
never raised at genesis, and a mesh that never installs it is complete.
Say *the deprecated broker*, not "the compatibility broker" (it serves the mesh's own modules,
not only the predecessor's) and not "the AMQP broker" (naming it after a protocol invites
describing the bus by contrast with it, which is backwards: the bus is the mesh's nervous
system and this is a module).
## What the mesh stores and serves
@@ -44,15 +58,22 @@ another — and a mesh you cannot name precisely is a mesh two people describe d
## How modules relate to the mesh
- **seat** — a named position at a scope (node / site / mesh) with a **capacity**. A capacity-1 seat
is exclusive (one holder); a higher-capacity seat is a **bench** (several holders coexist).
- **seat** — a named role at a scope (node / site / mesh), held by a module assignment, from a
**closed set** the mesh defines: a claim naming a seat outside the set is refused. A seat may
**deliver a provision**, and its holder is then the mesh's answer for it when several modules
provide it ([ADR 0126](../02-DECISIONS/0126-a-module-declares-its-own-seats.md) (superseding [ADR 0110](../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md))).
The set, with who holds each seat, is the overview of what a mesh has
([26 — The seats](../03-DESIGN/01-to-be/26-the-seats.md)). A seat has a **capacity**: a
capacity-1 seat is exclusive (one holder); a higher-capacity seat is a **bench** (several holders
coexist).
- **claim** — a module taking a spot on a seat. `claims: [{name, scope}]` in a manifest. A
mesh-scoped exclusive claim is how the mesh says "there is one of me". A foundation seat is
named after the server it guards: the `mesh-controller`, `postgres` and `lavinmq` modules claim
the `mesh-controller`, `mesh-store` and `mesh-broker` seats ([ADR 0079](../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md)).
- **provision** — a service one module `provides` and others `require`; the mesh resolves a provider
and wires the two with an endpoint and a credential. This is separate from seats: a provision is
a service you offer, a seat is a slot you occupy.
and wires the two with an endpoint and a credential. A provision is a service you offer, a seat
is a role you occupy, and the two meet where a seat delivers a provision: occupying the seat is
what makes a module *the* provider of it.
## How this page is kept
+6 -2
View File
@@ -33,7 +33,10 @@
A design changes only through a decision.
1. Write the decision record. If it reverses an earlier one, the earlier record's `status:`
becomes `superseded-by: 02-DECISIONS/NNNN-....md` — **its text is never edited**.
becomes `superseded-by: 02-DECISIONS/NNNN-....md` — **its reasoning is never rewritten**. If the
earlier record is sound and only a *fact* in it went stale, that is a **progressive insight**,
corrected in place and marked in the record rather than superseded
([`02-DECISIONS/README.md`](../../02-DECISIONS/README.md)).
2. Edit the to-be design document and set `updated:` to today.
3. If the amendment came from an issue, set that issue's `amended-design:` to the document
path.
@@ -54,4 +57,5 @@ Implementation state is a third axis, independent of both design and decision.
- Do not move a to-be document into `00-as-is/`. Write the as-is document; both stand.
- Do not edit an as-is document to describe an intention. That is what the to-be layer is for.
- Do not change a decision record's meaning. Supersede it.
- Do not change a decision record's meaning. Supersede it. Correcting a fact it got wrong, while
the decision stands, is a progressive insight — marked and dated in the record, never silent.
+2 -1
View File
@@ -1,6 +1,6 @@
---
status: canonical
updated: 2026-08-23
updated: 2026-09-24
---
# The Novox repositories
@@ -17,6 +17,7 @@ and a forge address is an operational detail (see [`README`](../README.md)).
| `hal` | The monorepo — the node runtime, the module catalogue, the delivery machinery, and the bootstrap scripts. Every core module lives here. |
| `hq` | This repository, under the company organisation — mission, research, design, decisions, issue diagnosis. Company-scoped ([ADR 0019](../02-DECISIONS/0019-how-this-repository-works.md)); the mesh is its first product. The source of truth for *why*. Carries no implementation. |
| *(one per application)* | Every standalone application, site or side-project gets its own repository, with `module.yml` at the root. Registered with the mesh as a build source; built and deployed by the same pipeline as anything in the monorepo. |
| `migration` | **Private.** The record of one installation replacing the predecessor mesh with this one: the runbook, a dated log of every step and what it cost, the per-service data procedures, the readiness checks, and the scripts. Private because it is the opposite of this repository in every way that matters — it names machines, addresses, ports and paths, because a procedure that cannot be followed is not one. Where hq asks *what did we decide and why*, that repository answers *what happened on the machines, in what order, and what to do next*. Its `HANDOFF.md` is where somebody picking the work up starts. |
## What the mesh becomes
@@ -0,0 +1,142 @@
---
status: active
initiated: 2026-09-24
touches:
- 02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md
- 02-DECISIONS/0033-the-substrate-is-a-store-and-a-broker.md
- 02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md
- 02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md
- 02-DECISIONS/0078-the-store-and-broker-are-modules.md
- 02-DECISIONS/0084-which-provider-serves-a-consumer.md
- 03-DESIGN/00-as-is/03-provisioning.md
- 03-DESIGN/01-to-be/07-the-foundation.md
- 04-ISSUES/113-the-object-stores-images-were-withdrawn-upstream/00-report.md
---
# 015 — The object store after MinIO: which S3 implementation, and how the data moves
**The question.** The mesh's object store is MinIO. Its community edition is archived upstream,
its server and client images have been deleted from every public registry, and the pinned release
is four and a half years old and will never be patched
([issue 113](../../04-ISSUES/113-the-object-stores-images-were-withdrawn-upstream/00-report.md)).
Which S3-compatible implementation replaces it, and what is the migration track for the data and
the provisioning model that sit on top of it?
**Why now, and why not sooner.** Nothing is on fire: nodes that already hold the images keep
running, and issue 113 establishes that the deploy path tolerates an unfetchable-but-present
image by design. The forcing function is not an outage but a one-way door — **no node that does
not already hold the images can ever provision the module again**, so the mesh's ability to stand
a node up from its declarations is already broken for this module, and silently.
**The direction is not a departure from the design; it is the design.** The foundation document
already states the commitment:
> The dependency is on the **protocol**, not the product: AMQP for the bus, S3 for the object
> store, the OCI protocol for the registry. That is what keeps the naming safe rather than a
> commitment that cannot be revisited.
The object store is also **not** a foundation service — ADR 0028 removed it, and it is an
ordinary module required through the module graph by whatever wants one. (The "exception that is
not a swap" in that passage is the relational store, whose provisioning model borrows PostgreSQL's
own meaning of databases, roles and schemas. The object store carries no such coupling: a bucket
is a bucket.) So this effort is an instantiation of an existing principle, not a redesign — which
is the cheapest kind of decision to make and the strongest kind to cite.
## What the replacement has to carry, measured
Taken from the module's manifest, its composition, its tool surface, and a search for its
consumers across the catalogue — not from assumption.
| Requirement | Evidence in the module today |
|---|---|
| S3 API | The protocol every consumer speaks; already the design's stated dependency. |
| ~~OIDC login against the mesh's identity provider~~ | **Struck 2026-09-24. Not a requirement, and it never worked.** Six variables are wired and an entrypoint blocks on the provider, which reads as a live feature. The module's own hook comment records the end state as *"policy claim missing"* — a failing login. See [01](01-candidate-comparison.md). |
| **Per-application access keys, each scoped to a bucket** | The real requirement. A "user" of the store is normally an application; the mesh already mints a credential per provisioned bucket. |
| **One live consumer using it as opaque primary storage** | A file-sync application, since early 2023: objects named by internal id, metadata in its own database. Highest-risk consumer — a live copy drifts, and its bucket name must be preserved. |
| Erasure-coded multi-node topology | Four server nodes with two data directories each, behind a load balancer. |
| A single-node form | Declared as a flavour, for development and small nodes. |
| Buckets as a typed provision | The module declares a provision type of `bucket` on a named network; the mesh mints the credential and the provider creates it (ADRs 0048, 0084). |
| A tool surface | Bucket create/list/delete, object list/info/delete, presigned URL, and provisioning. |
| A console | Published on its own subdomain through the reverse proxy, with an unlimited request-body middleware for uploads. |
**Consumers, counted:** one application module, one capture module that takes a private bucket per
node, one workflow module's tools, and the delivery/rescue internals of the shared library. The
surface is small — the cost is concentrated in the provisioning handler, the tool handlers and the
OIDC story, not spread across the catalogue.
## Candidates
**Four candidates, not three.** The comparison was briefly narrowed to SeaweedFS on the strength of
console single sign-on; that axis turned out not to be a requirement, and the incumbent's own
maintained fork had been omitted altogether. Both errors, and why they happened, are recorded in
[01 — the candidates measured](01-candidate-comparison.md), which carries the evidence and the
requirement-by-requirement detail.
In short, and only in short:
- **The maintained fork of the incumbent** — the community edition was archived and its images
deleted, but a fork publishes, tracks CVEs, and preserves the on-disk format, S3 API and
environment surface. Costs **an image reference** where every other option costs a data
migration, two rewrites and a maintenance window. Does not end the dependence on an abandoned
codebase; buys time to choose deliberately.
- **Garage** — its permission model *is* the requirement (per access key, per bucket), its admin
API is the closest match to how the mesh provisions, and the highest-risk consumer is
first-party documented against it. Remaining cost: no object versioning, no server-side
encryption or object locking, partial lifecycle — **unmeasured against the ten buckets, and the
one thing that could still disqualify it**.
- **SeaweedFS** — longest field record and erasure coding. Its console sign-on is a paid feature,
which is now beside the point. What weighs against it is narrower: its S3 surface is a gateway
translating onto its own file-system API, with no first-party support for the opaque consumer.
- **RustFS** — closest in shape to the incumbent, so the least porting. But it reached general
availability eight days before this was written, and carries an open defect in the credential
path. Two earlier claims about it are corrected in 01: it is **not** a drop-in that retains
existing data.
- **Ceph RGW** — remains rejected as disproportionate where the object store is an ordinary
module rather than a platform.
**This is now two decisions, not one:** whether to repoint to the fork or migrate, and — if
migrating — to which. Repointing does not foreclose migrating, which is the argument for taking it
first. On the corrected requirement the migration ranking is Garage, then SeaweedFS, and not yet
RustFS. Two measurements gate any graduation: **which S3 endpoints the consumers actually call**
(Garage cannot be ranked fairly until counted), and **whether the fork can read the incumbent's
on-disk format in place** — tested on a copy, because the migration between them is one-way. Both
are in [01](01-candidate-comparison.md#what-is-still-unmeasured).
## The migration track, in outline
Data movement is the easy half, and deliberately reversible.
1. **Stand the replacement up beside the incumbent**, on its own ports, its own provision type and
**its own data directory**. Nothing removed. The data directory matters: reusing one the
incumbent already holds would put a fresh single-drive store on top of a live erasure set.
2. **Copy bucket by bucket with a neutral tool.** `rclone` rather than the incumbent's own client
— the client has been withdrawn upstream too, so building the migration on it would inherit
the same dependency this effort exists to remove.
3. **Verify per bucket** — object counts and checksums, not a transfer exit code.
4. **Repoint consumers through the connection the module already publishes.** Consumers read an
API URL from the module's declared connections rather than addressing the store directly, so
the cutover surface is that value plus the provisioning and tool handlers.
5. **Freeze writes, final incremental sync, flip**, and keep the incumbent read-only as the
rollback until confidence is earned. For the opaque consumer this is **not optional and not
instant**: it stores objects by internal id with metadata in its own database, so a copy taken
while it runs will drift. It needs a maintenance window for the final sync, and the window is
proportional to 82,496 objects rather than to 230 GiB.
6. **Retire**, and only then remove the module.
The genuinely new work is not the copy. It is the **provisioning handler** and the **tool
handlers**, both written against the incumbent's admin API. *The OIDC wiring was previously listed
here and is struck: it is not a requirement and it never worked.*
## Open questions
- ~~How much of the OIDC requirement survives, and in which build?~~ **Answered, and it was the
wrong question.** The console requirement does not exist, and the login it referred to never
worked. What replaced it: which S3 endpoints consumers actually call, and whether the fork reads
the incumbent's format in place.
- Does the mesh's bucket provision translate to the candidate's identity model without weakening
what ADR 0049 says about a consumer's identity fitting the tightest backend?
- Should this effort also answer issue 113's general question — mirroring third-party images into
the mesh's own registry — or is that a separate decision? Replacing one withdrawn product with
another unmirrored upstream leaves the same one-way door in place, just further from the hinge.
- Is the four-node erasure-coded topology still warranted, or was it inherited? Worth re-asking
while the product is being chosen, rather than reproducing a shape by default.
@@ -0,0 +1,176 @@
# 015 / 01 — The candidates measured
*Rewritten 2026-09-24. An earlier version of this document ranked the candidates on whether they
preserved single-sign-on to the object store's **console**. That was the wrong axis — it is not a
requirement — and a fourth candidate was missing entirely. Both errors are recorded at the end,
because how a comparison came to be ranked on the wrong thing is worth more than the ranking was.*
## The requirement, corrected
Taken from the operator and from the running system, not from the module's shape.
**A "user" of the object store is normally an application.** The requirement is therefore
**per-application access keys, each scoped to its own bucket** — not per-human single sign-on. The
mesh already works this way: it mints a credential for every provisioned bucket, and the consumer
reads an endpoint from the module's declared connection rather than addressing the store directly.
**The console is not a requirement.** It was the axis the previous version ranked on, and it should
not have been.
**The identity-provider login never worked.** The predecessor's module wires six OIDC variables and
blocks startup until the provider answers, which reads like a working feature. It is not: the
module's own hook comment records the end state as *"policy claim missing"* — a **failing** login,
written up as progress because it proved the provider had registered. The identity provider emits no
such claim, nothing in the module creates the mapper, and the configured scope alone would not carry
a custom one. Two days of logs show no genuine login attempts, only internet scanners failing on an
STS API version. **Nothing should be carried forward on the assumption this works**, and no
candidate should be credited or penalised for matching it.
**One consumer is live, opaque, and holds real user files.** A file-sync application has used the
store as its **primary storage** since early 2023: objects named by an internal id, with all
metadata in its own database. Three consequences — a copy taken while it runs will drift, its bucket
name must be preserved or its database references break, and it is the highest-risk consumer of the
lot.
## What is actually stored, measured
| | |
|---|---|
| Logical | **230 GiB, 82,496 objects, 10 buckets** |
| Raw on disk | **468 GiB** — eight drive directories at 59 GiB each |
| Implied scheme | 468 ÷ 230 = **2.03×**, confirming erasure coding at half parity |
| Headroom | ~1.3 TiB free on the filesystem holding it |
**All eight "drives" are directories on one filesystem on one machine.** The erasure coding is
therefore not buying independent-drive redundancy; the real failure domain is the array underneath,
which has its own. This single fact decides more of the comparison than any product feature: a
scheme's redundancy model is close to irrelevant here, and what remains is its storage overhead.
At 230 GiB with 1.3 TiB free, **storage overhead is not a deciding cost either.** Replication at
three copies would run ~690 GiB against the present 468 GiB — about **+222 GiB**, comfortably
absorbed. Erasure coding at a wider stripe would *save* roughly 146 GiB. Both are rounding errors
against the headroom, and neither should decide this.
## The candidates
Four, not three. The previous version omitted the first.
### The maintained fork of the incumbent
The community edition was archived upstream and its images deleted
([issue 113](../../04-ISSUES/113-the-object-stores-images-were-withdrawn-upstream/00-report.md)),
but **a fork is maintained and publishing** — `pgsty/minio`, from the Pigsty project. It restores
the console stripped from the community
build, rebuilt image and package distribution, tracks CVEs, and states that it preserves the on-disk
format, the S3 API and the environment-variable surface. Verified by pulling it: it reports a
current release, permissive-to-copyleft licensing unchanged from upstream, and identifies itself as
a community fork. Adoption is real — the server image has been pulled three quarters of a million
times.
**Why it reorders the comparison.** Every other candidate costs a data migration, a provisioning
handler rewritten against a different admin API, a tool surface ported, and a maintenance window for
the opaque consumer. The fork costs **an image reference**. It also closes the issue's one-way door:
a node holding nothing can provision the module again, and patches resume.
**What it does not do** is end the dependence on a codebase its original authors abandoned. It is
maintenance mode, largely one project's effort, with no new features intended. It buys time to
choose deliberately rather than under pressure — which is worth a great deal, and is not the same as
a decision.
### Garage
**The best fit for how the mesh provisions.** Its permission model is *per access key, per bucket,
read/write/owner* — which is the requirement above stated verbatim rather than approximated. Its
admin API is a first-class REST surface with tokens scopeable to exactly the two operations a bucket
provision performs. The opaque consumer is **first-party documented** against it, for primary
storage, including client-side encryption support.
Its previously-recorded penalties mostly dissolve under the corrected requirement: it has no console
and no identity-provider integration, neither of which is wanted; and it replaces AWS-style ACLs and
bucket policies with its own per-key-per-bucket model, which is the thing being asked for.
**What genuinely remains.** It replicates rather than erasure-codes — immaterial at this volume and
on a single array, as above. It does **not implement the full span of S3 endpoints**: object
versioning is absent, object locking and server-side encryption endpoints are absent, and lifecycle
is partial. **Whether any of the ten buckets depends on those is unmeasured, and it is the one thing
that could still disqualify it.**
### SeaweedFS
Longest field record of the group, permissive licence, erasure coding, and identity-provider
integration on the S3 API through token exchange. **Its console sign-on is a paid feature** — the
admin UI itself is open, its identity integration is not. That finding is what falsified the
previous version's narrowing, and it is now largely beside the point, since the console is not a
requirement.
What weighs against it here is narrower and more specific: its S3 surface is a **gateway
translating onto its own file-system API**, with acknowledged divergence from AWS behaviour at the
edges, and there is no first-party documentation for the opaque consumer. For a store already
holding real user files in an opaque layout, first-party support is worth more than a feature list.
It also carries more moving parts than a single-machine deployment needs.
### RustFS
Closest in shape to the incumbent — a similar admin API and client compatibility, so the existing
handlers would port with least effort — under a permissive licence, with erasure coding and a
console that does integrate an identity provider.
Two things were recorded about it earlier that were **wrong, and are corrected here**: it is *not* a
binary-level drop-in that retains existing data (API compatibility and on-disk compatibility are
separate paths, and the on-disk one is preview-scoped with documented encryption limits), and it
therefore offers no shortcut around the migration. It also carries **an open defect in the exact
area the mesh depends on** — an access key created by an identity-provider user reported denied on
all S3 operations.
Decisively for now: **it reached general availability eight days before this was written.** For a
component holding 230 GiB of real user files, field record is a feature, and it does not have one.
### Ceph RGW
Remains rejected, for the reason already recorded: disproportionate where the object store is an
ordinary module rather than a platform.
## Where this leaves it
**The decision is no longer "which product replaces the incumbent".** It is two decisions, and they
can be taken in either order but should not be confused:
1. **Repoint to the maintained fork, or migrate now?** Repointing is an image reference and it
closes the issue. Migrating now costs a data copy, two rewrites and a maintenance window, and
buys independence from an abandoned codebase sooner.
2. **If migrating, which?** On the corrected requirement the ranking is **Garage first** — its
permission model *is* the requirement, its provisioning API is the closest match, and the
highest-risk consumer is first-party supported. SeaweedFS second, on field record, with a
translation-layer caveat that matters more here than its feature list. RustFS not yet, on age.
Taking (1) does not foreclose (2), and that asymmetry is the argument for taking (1) first.
## What is still unmeasured
1. **Whether any of the ten buckets needs object versioning, server-side encryption or lifecycle.**
This gates Garage specifically and nothing else here answers it.
2. **Whether the fork's release can actually read the incumbent's on-disk format in place.** The
format is claimed compatible across a multi-year gap; the migration between them is one-way, so
this is tested on a copy or not at all.
3. **Whether the opaque consumer's maintenance window is acceptable**, and how long it actually is
at 82,496 objects.
4. **Whether the eight-drive erasure-coded shape is warranted at all.** The evidence above says it
is not buying what it appears to: eight directories, one array, one machine. It looks inherited.
**Nothing graduates to a decision before 1 and 2.**
## Two errors in the previous version of this document
Recorded because the shape of both survives anonymisation and neither is unique to this effort.
**It ranked on a requirement that did not exist.** Console single sign-on was treated as the axis
because the module's configuration showed it wired up, and a wired-up configuration was read as a
used feature. It was neither used nor working. *A configured feature is not an observed one*, and
the evidence needed was the operator's answer and the logs — both cheap, neither consulted before
the ranking was written.
**It omitted the incumbent's own fork.** The whole effort began because an upstream withdrew its
images; whether anyone had continued that upstream was the first question to ask and it was not
asked. The candidate list was assembled from a search for *alternatives*, which by construction
returns things that are not the incumbent. *When a dependency dies, "who took it over" precedes
"what replaces it".*
@@ -0,0 +1,57 @@
---
status: graduated
became:
- 02-DECISIONS/0114-a-shared-credential-rotates-over-two-credentials.md
- 03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md
initiated: 2026-09-26
touches:
- 02-DECISIONS/0113-the-vault-makes-every-secret.md
- 02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md
- 02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md
- 03-DESIGN/01-to-be/13-credentials-and-their-rotation.md
- 03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md
- 04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md
---
# 016 — How a credential can be rotated
**What.** Which rotation mechanisms the mesh's providers can actually support, measured against
their code rather than assumed. Every provider in the catalogue was read, found by listing every definition that provides something: how it names what it
makes for a consumer, what its remove destroys, whether it re-applies a password, whether its
backend can hold two secrets for one login or two logins on one resource, and how its own
administrative credential is set. The consumer side was read too: when a module reads a secret, and
what makes it read a new one.
**Why.** [ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md), as first drafted,
chose *overlap*: add a second login beside the first, move every reader, then remove the old one,
"through the adapter's existing create and remove", with "no consumer changes". A review showed that
claim false. In most providers the consumer's data is named after its login, and remove drops the data
with the login. Overlap as written would have deleted every consumer's database on its first
rotation. The mechanism has to be chosen on what the providers do.
**What it touches.** Rotation in 0113 and [to-be 27](../../03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md),
which [ADR 0114](../../02-DECISIONS/0114-a-shared-credential-rotates-over-two-credentials.md) decided on
these findings. The identity budget in
[ADR 0049](../../02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md), if a consumer
gets two logins. The rotation already implemented, which [to-be 13](../../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md)
describes.
**Documents.**
- [01 — The providers](01-the-providers.md): the survey, one row per provider, and what it shows.
- [02 — The readers](02-the-readers.md): how a secret reaches a running process, and what already
recreates it.
- [03 — The options](03-the-options.md): each rotation mechanism against those facts, and a
recommendation.
**Finding, in one paragraph.** All nine credential providers already re-apply a consumer's password
in place on every create, and the controller's `rotate` command relies on that. It is a working
rotation with a stated window. Eight of the nine name the consumer's resource after its login, and five
destroy the consumer's data when they remove the login. The harness, keyed by login, would do the same
on any change of login. Only one backend holds two passwords on one login, and two more hold several
tokens. Eight backends can grant two logins the same rights over one resource; the ninth can give one
login a second token. So every provider can hold **two credentials** over one resource, but only after
each adapter separates *the consumer's resource* from *the credential that reaches it*. In postgres
that also means the resource belongs to a role no login owns. Administrative credentials are a
different case. They have one party and a fixed name, and five backends take them only at first
initialisation, so changing one needs the old and the new value at once.
@@ -0,0 +1,103 @@
# 01 — The providers
Read from the catalogue's main branch: each provider's provisioner adapter (`create`, `remove`),
the client functions they call, and each definition's own credentials. The providers were found by
listing every definition that provides something and has a provisioner, not from memory. A first pass
of this survey worked from memory and missed one, mailu.
**The provisioner harness** in `mesh-sdk` calls `create` for a consumer when its contribution appears
or changes (its login, password or values), and after the provisioner restarts. It calls `remove` for
a login it applied earlier in the same process that is no longer contributed. Its record of what was
applied is kept in memory and keyed by login. Two things follow:
- a consumer whose derived login changes is removed under the old login and created under the new one,
in one pass;
- a contribution that disappears while the provisioner is down is never removed, and is left behind.
## The credential providers
`login` is the consumer's derived identity, which the adapter receives as `as`
([ADR 0049](../../02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md)).
| provider | the consumer's resource is named | remove destroys | create re-applies the password | two secrets on one login | two logins on one resource |
|---|---|---|---|---|---|
| postgres | a database named `login`, owned by the role `login` | the database and the role | yes, `ALTER ROLE … PASSWORD` when the role exists | no: a role has one password | yes, but only through a role that cannot log in owning the database, with each login working as it. Otherwise whatever one login creates is its own, and dropping that login means handing its objects over first. Not done today |
| mssql | a database named `login`, with the login mapped into it | the database and the login | yes, `ALTER LOGIN … WITH PASSWORD` | no: a login has one password | yes, two logins mapped to users in `db_owner`. A user owning a schema cannot be dropped, and a login with an open session cannot. Not done today |
| mongodb | a database named `login`, with a user holding `dbOwner` | the database and the user | yes, `updateUser` with the new password | no: a user has one credential | yes, two users with `dbOwner` on one database. Not done today |
| redis | the key prefix `login:` on an ACL user named `login` | the user, **not** its keys | yes: `ACL SETUSER … reset … >password` replaces all of them | **yes**: an ACL user holds several passwords, added with `>` and removed with `<`. Today's `reset` discards all but the new one | yes, two users on one key prefix, once the prefix is not the login |
| minio | a bucket derived from `login`, and a service account whose access key is `login` | the access key; the bucket **only if empty**. A bucket holding objects is left, and the failure logged | yes, by removing the access key and adding it again, which leaves a moment with no key | no, but an access key *is* the login: a second key is a second login | yes, two service accounts with one bucket policy. The access key is capped at 20 characters |
| lavinmq | a virtual host named `login`, and a user named `login` with permissions on it | the virtual host, with any queued messages, and the user | yes, the user is written again with the password | no: a user has one password | yes, permissions for two users on one virtual host |
| mosquitto | a client named `login`, with a role named for it on the topic prefix `login/#` | the client and its role | yes, the password is set when the client exists | no: a client has one password | yes, two clients holding one role, once the prefix is not the login. The MQTT client identifier is chosen by the consumer, not tied to the login; a duplicate one takes the older session over |
| mailu | a mailbox `login@domain`, unless the consumer contributes its own account name | the mailbox with its mail, for a login-named one; a contributed name is left for an operator | yes, the password is set when the user exists | no for the password; a user can hold several authentication tokens, per the backend's documentation | **no**: a mail user *is* its mailbox |
| gitea (npm) | a user named `login` on a team of an organisation that owns every package | the user; **packages survive**, because the organisation owns them | yes, the user's password is set on every run | no for the password; a user can hold several access tokens | yes, trivially: a second member of the same team |
## The other providers
| provider | answers with | credential |
|---|---|---|
| umami | a website, found by its public name | none. The site id it makes has no way back to the consumer today |
| cloudflare-dns | a public name derived from `login` | none handed to the consumer; its own API token is an operator value |
| showcase | a route | none |
| mesh-vault | custody: it records and withdraws sealed values in a ledger | it holds secrets; it makes none today |
verdaccio provides the npm registry too, and has no provisioner.
## Each provider's own administrative credential
| provider | identity | how the backend takes it |
|---|---|---|
| postgres | a fixed superuser | from a file **only at first initialisation** |
| mssql | `sa` | from the environment at first setup. The image documents no file form, and the definition records that as a declared exception |
| mongodb | a fixed `root` | from a file **only at first initialisation**, when the data directory is empty |
| mosquitto | a fixed admin client | seeded into the broker's dynamic-security file **once**; the seeding step skips when the file exists |
| lavinmq | a fixed admin name | per its own bootstrap code, **only on a first boot** with an empty data directory. No resource in the definition runs that bootstrap; what sets it on a running mesh is outside the catalogue |
| redis | the default user | from `requirepass` in a configuration the mesh renders, read when the server starts |
| minio | a fixed root user | from a file, read when the server starts |
**In five of seven, a new administrative value takes effect only through a command run with the old
one.** The credential file is mounted directly into both the server and the provisioner. So replacing
it recreates the provisioner, which then holds only the new value while the backend still expects the
old one, and the provisioner is locked out. That is worse than changing nothing.
Every provider module also has its own bus account, an own secret, read at start.
## What the tables show
1. **Every credential provider already rotates in place.** All nine re-apply the password on the
same login each time `create` runs. The controller's `rotate` command relies on that: it replaces
the credential in the inventory and sends both ends in one push. Its own comments state the window,
between the provider applying and the consumer restarting, in which the consumer cannot
authenticate.
2. **Eight of nine name the consumer's resource after its login.** Only gitea separates them,
because an organisation owns the packages. A second login therefore has no resource of its own to
reach, and cannot share the first one's without the adapter granting it.
3. **Five of nine destroy the consumer's data when they remove the login**: postgres, mssql and
mongodb drop the database, lavinmq drops the virtual host with its queued messages, and mailu
deletes the mailbox with its mail. minio drops only an empty bucket, and redis leaves the keys. In
those five, *retire a login* and *delete the consumer's data* are one call. With the harness keyed
by login, a changed login triggers it too.
4. **One backend holds two passwords on one login** (redis). Two hold several tokens beside one
password (gitea and mailu). A rotation built on two secrets per login would work for three
providers out of nine.
5. **Eight of nine can give two logins the same rights over one resource.** Group roles in postgres,
database roles in mssql and mongodb, permissions in lavinmq, a shared role in mosquitto, a shared
policy in minio, a shared key prefix in redis, a shared team in gitea. mailu cannot, because its
user is its mailbox, but it can give one user a second token. So every provider can hold **two
credentials** over one resource, though not every one as two logins. No adapter does either today.
6. **Ownership is a trap in two backends.** In postgres whatever a login creates is that login's, so a
second login cannot alter the first one's tables, and the first cannot be dropped while it owns
them. The one-step way out deletes them. In mssql, a login cannot be dropped with a session open,
nor its user while it owns a schema.
7. **The administrative credentials have one party and a fixed name**, and five backends take them
only at first initialisation. The provisioner needs the old and the new value at once to change
them. Today nothing can give it both.
8. **A consumer's identity is already the resource's name.** The login is derived from the
assignment, which is a module on a node, so the current login and "the consumer" are the same
string today. A second login would need a new name. The resource can keep the one it has.
## Seen on the way
The redis configuration names no ACL file, so a consumer's ACL user exists only in memory. A restart
of the redis server erases every consumer's user. The provisioner does not create them again until it
restarts itself, because its in-memory record says they are done. That is not a rotation finding, but
it is a live fault, and it is recorded here so it is not lost.
@@ -0,0 +1,46 @@
# 02 — The readers
How a secret reaches a running process, and what makes the process take a new one.
## No module watches a secret
A search of every module's code in the catalogue found no file watching of any kind, and no
re-reading of a secret while running. **Every reader reads a secret when it starts.** There is no
consumer that takes a new value live, so every rotation that changes what a consumer presents ends in
the consumer restarting.
## The host already recreates what read a changed file
The node host records, for every long-running container, the digest of each file it read when it
was created: its env-files, and every file bind-mounted into it directly. When a digest changes, the
host recreates the container, even though its spec is otherwise unchanged. This is the fix for
[issue 103](../../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md).
It is on the host's main branch, while the issue is still recorded as located, not fixed.
Two cases are deliberately left out and need `restart-on` in the definition:
- a file read out of a **mounted directory**, because the host cannot know whether the service reads
it once or watches it (a route proxy re-reads its routes live; a provisioner polls what it receives);
- a **process** rather than a container.
## What the catalogue does with it
22 definitions declare a secret they receive. In 16 of them it reaches the service through a
rendered file, usually an env-file. That case the host already covers. 14 declare `restart-on` for
something. Whether each of the 22 is fully covered depends on how its secret travels: through an
env-file or a direct mount, which the host covers, or through a directory or into a process, which
needs `restart-on`. **That was not classified module by module.** It is the check to run before a
rotation mechanism relies on it.
The count covers only the `secrets` field. The 49 modules with their own bus account, and 54 with any
own secret, are readers too, and their bus accounts are rotated like any credential two parties hold.
Their files are mounted directly, which the host covers, but the classification has to name them.
## What this means for rotation
- The *read at start* half of 0113's recipient model is already true, and mostly already handled by
the host. The restart is derived from the files a container reads, not declared per secret.
- Any mechanism, in place or overlapping, ends with the reader being recreated. What differs is
whether the credential it held until then still works.
- For a single-party secret, a module's own, the reader is also the only holder. There is nobody to
overlap with, and delivering the new file recreates the reader.
@@ -0,0 +1,79 @@
# 03 — The options
Three mechanisms, weighed against [01](01-the-providers.md) and [02](02-the-readers.md).
## A. In place, as today
The vault makes a new value. Every applier re-applies it on the same login, which all nine
providers already do. Every reader is recreated by the host.
- **Works with:** every provider, unchanged. It is what `rotate` does now.
- **Costs:** a window per consumer, from the provider applying to the consumer being recreated. They
are on different machines, and nothing orders them. A reader whose machine is unreachable from the
mesh but still reaches its provider stays locked out until the mesh reaches it again.
- **Admin credentials:** the natural form. The provider module is the only party, and it has to
apply the new value with the old one anyway (finding 6).
## B. Two secrets on one login
The applier adds the new password beside the old one, readers move, and the old one is removed.
- **Works with:** redis natively, and gitea and mailu through tokens. **Not** with the other six, whose
backends hold one password per login (finding 4).
- **Verdict:** not a mechanism, a special case. Using it where it exists and something else
elsewhere is the "this way or that way" the design is trying to remove.
## C. Two credentials per consumer, over one resource
The consumer has two credentials and uses one at a time. The applier ensures the other with the new
value and gives it the same rights over the consumer's resource. Readers move to it, and then the old
credential is retired, which removes the credential only, never the resource. **What a credential is,
is the adapter's**: a second login for eight providers (finding 5), a second token on the same login
for mailu. The mesh sees one mechanism.
- **Works with:** every provider, **after** each adapter changes:
- the resource is named after the consumer, not the login. Today the two are the same string (finding
8), so existing resources keep their names, and the current login stays one of the two;
- the resource is owned by the resource, not by a login. In postgres that is a role no one logs in
as, which each login works as, and ownership of an existing database moves to it once (finding 6);
- both credentials get the same rights, over data and structure;
- *retire a credential* and *remove the consumer* become two operations. Today they are one call, and
in five providers that call destroys data (finding 3). The harness must key by consumer, so that a
changed login is not a removal. This is the whole of the danger, and it has to be split, whatever
else is chosen.
- **Costs:**
- every credential adapter changes;
- the harness learns the alternation and a confirmation per step, and rotation state has to live
somewhere that survives a restart, which the harness's memory does not;
- the second login's name must fit the tightest backend. That is 20 characters for a minio access
key ([ADR 0049](../../02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md)), and
a suffix spends part of it;
- retiring an mssql login has to end its sessions first.
- **Gains:** no window. A reader that cannot be reached keeps a working credential until it can.
- **Does not apply to** single-party secrets: admin credentials and a module's own secrets. There is
no second party to overlap with.
## Independent of the choice
- **Split remove, and key the harness by consumer.** Retiring a credential, or a login changing, must
never be able to destroy a consumer's data. That holds under A too, because A's remove is the same
call.
- **Admin credentials are applied by their own provider**, using the old value, with the new one staged
beside it. Five backends take the value only at first initialisation. Replacing the file first locks
the provisioner out (finding 7).
- **Classify the readers** ([02](02-the-readers.md)), bus-account readers included, before relying on
derived restarts.
## Recommendation
- **Two-party credentials, consumer credentials and bus accounts: C**, because it is the only
mechanism every provider supports, and it closes the window instead of shortening it. Its prerequisite, separating the resource from the login and
retiring a login from removing a consumer, is worth doing on its own, because it removes a
data-loss path that exists today.
- **Single-party secrets (admin credentials, a module's own): A, staged.** In place, applied by the
provider that holds them, with the new value beside the old until it has taken.
- **Until the adapters are changed, A stays** as `rotate` implements it, with its window stated. It is
not replaced by a mechanism the providers cannot yet carry.
This is two mechanisms, split by a property of the secret rather than by provider: whether it has one
party or two. Every provider is treated the same way for the same kind of secret.
@@ -0,0 +1,47 @@
---
status: active
initiated: 2026-09-26
touches:
- 00-META/mission.md
- 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0010-delivery.md
- 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
- 03-DESIGN/01-to-be/06-the-controller.md
- 03-DESIGN/01-to-be/09-the-node-lifecycle.md
- 03-DESIGN/00-as-is/09-interfaces-and-observability.md
---
# 017 — A mesh that heals itself
**What.** The behaviour the operator wants: a mesh that runs itself. It notices what is wrong,
repairs what it can, and hands what it cannot repair to someone who can, with the reason. This effort
writes that wish down as intended behaviour, designed for the bus the mesh is moving to
([ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md): NATS). It also records what can be done
pragmatically before that move.
**Why.** The mission is *a mesh that controls itself* ([mission](../../00-META/mission.md)). The
mesh can tell whether it is up. It cannot tell whether it is right. The as-is page on observability says so
([as-is 09](../../03-DESIGN/00-as-is/09-interfaces-and-observability.md)). To-be 06 names an
`observability` context in the controller and leaves its store undecided. Nothing routes a condition
the mesh cannot fix to anyone. The cost is measurable: **46 of the 116 issue reports in this
repository describe a failure that was silent.** A mesh that heals itself is, first, a mesh that stops
failing silently.
**What it touches.** The controller's observability context, the node lifecycle's liveness, delivery
([ADR 0010](../../02-DECISIONS/0010-delivery.md), [ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)),
the provisioner harness, and rotation, which is proposed alongside to-be 27 as ADR 0114.
**Documents.**
- [01 — The intended behaviour](01-the-intended-behaviour.md): the wish, as principles and as how
the mesh behaves once the bus is NATS.
- [02 — Now, pragmatically](02-now-pragmatically.md): what is done before NATS, why it does not
build anything the move would throw away, and what has been done already.
**Next.** Two measurements this effort owes before it can graduate:
1. **Every loop in the mesh**: what it converges, and whether it compares against observed state or
against its own memory. Issue 120 found the provisioner harness trusting memory. The same pattern is
expected elsewhere.
2. **The 46 silent failures, classified**: a missing observation, a loop trusting memory, or a missing
escalation. That shows which mechanism removes the most of them.
@@ -0,0 +1,85 @@
# 01 — The intended behaviour
The operator's wish, written as behaviour: what a person or an agent sees the mesh doing. This is a
target to design toward, not a design. Every part of it is to be decided through a record before it
is built.
## Principles
**1. Every loop compares what should be with what is, never with what it did.** Desired state is the
mesh's: assignments, requirements, seats. Observed state is read from the thing itself: the container,
the backend, the node. A loop that compares against its own memory of what it applied is blind to
anything that changed behind its back. That is issue 120, and it is the pattern this whole effort is
written against.
**2. Healing is the ordinary path run again, never a second path.** Repairing a lost login is
provisioning it. Repairing a dead container is converging the node. Repairing a stale declaration is
delivering it. A repair that needs its own code is a second way of doing something, which is exactly
what the mesh is removing everywhere else.
**3. A repair never destroys.** Healing may recreate, re-provision, re-deliver and restart. It may
never delete a consumer's data, retire a credential someone still uses, or pick a winner between two
contradictory states. Where the only repair is destructive, it is escalated.
**4. Nothing fails silently.** Every condition the mesh cannot repair within its budget becomes
visible. It is named, it says since when, why, and who can resolve it. It is visible until it is
resolved, and resolved by observation, not by someone clicking it away.
**5. What the mesh cannot fix goes to an agent.** Per the mission, an agent may be human or not. A
condition that needs judgement is handed to one, as work, with what the mesh knows. It is not handed
over as a notification that someone may or may not read.
**6. Correctness, not only liveness.** A running process that authenticates with a dead credential,
serves an old version, or routes nowhere is not healthy. What a provision's contract promises is what
is checked: the credential authenticates, the route answers, the version is the declared one.
## The loop, everywhere
Every part of the mesh that owns something runs the same loop:
1. **know** what should be true: from assignments, requirements and seats;
2. **observe** what is true: from the thing itself, on its own cadence;
3. **repair** the difference by running the ordinary path again, within a budget of attempts and
time;
4. **raise** a *condition* when the budget is spent or the only repair is destructive;
5. **clear** the condition when observation shows it resolved.
A **condition** is a durable fact about something the mesh owns, such as a node, an assignment, a
provision, a seat or a rotation: what is wrong, since when, the evidence, what was tried, and who can
resolve it. Conditions are the one thing a person or an agent looks at to know whether the mesh is
right. `status` is the list of open conditions. When it is empty, the mesh is right, not just up.
## On NATS
[ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md) moves the bus to NATS, and NATS makes most of
this cheaper, because observation becomes something every component publishes rather than something
a central process polls.
| the wish needs | on NATS |
|---|---|
| every component says it is alive | a heartbeat on a subject per node and assignment; silence past its interval is a condition, and nobody polls |
| every component says what it observed | observations published on subjects (`mesh.observed.<node>.<assignment>`, for instance), consumed by whoever owns the comparison |
| the last known state survives restarts | a JetStream key-value bucket of observed state per owner; the provisioner's "what I applied" and a rotation's step live there, not in memory |
| conditions are durable and watchable | conditions as entries in a key-value bucket, watched by anyone who cares: a surface, an agent, the controller |
| the bus itself is observed | the server's advisories (a consumer exceeding its deliveries, a slow consumer, a client disconnecting) and its monitoring endpoint become observations like any other |
| a repair is retried, not lost | JetStream redelivery with delay, which is the same mechanism [ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)'s guarantee moves to |
| work handed to an agent | a condition that needs judgement published as a task on a subject an agent's queue group consumes |
**Who compares.** Each owner compares its own: the host for its node's containers and files, a
provisioner for its backend, the vault for rotations, the controller for delivery and seats. The
controller's observability context does not repair anything. It holds conditions, their history,
and the view across the mesh. It notices what no owner can see about itself: an owner gone silent.
## What stays human
Some repairs need the operator's key, and the mesh says so rather than pretending otherwise:
re-raising the vault or the broker, and recovering a node's identity. These are conditions too, with
the procedure named, and they are the only ones that can never clear themselves.
## Open
- The budgets: how many attempts, over how long, per kind of repair.
- How a condition that needs judgement reaches an agent, and how the agent's action is recorded.
- Where the observability context stores history (to-be 06 left it open; volume argues against the
relational store).
- Which correctness probe each provision's contract offers, and how often it runs.
@@ -0,0 +1,46 @@
# 02 — Now, pragmatically
The intended behaviour lands on NATS. The bus moves after the migration's core
([ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md)). Until then, work toward it is chosen by one
test:
**Does it survive the move?** A change to what a loop compares against, or to what an adapter can
tell about its backend, survives, because it is independent of the bus. A new AMQP queue for health
reports, a poller written against the broker's management API, or a condition store built on the
current broker does not survive, and is not built.
## Done
**The provisioner asks the backend, not memory** (issue 120). The SDK's harness gained an optional
`holds` on the adapter, asked for every applied consumer every minute. A consumer the backend no
longer holds is provisioned again. Being unable to ask is not treated as loss. The cache module
implements it first, because its server keeps its users in memory and forgets them all on a restart.
That was verified against a real server: a restart erases every consumer's user, and `holds` answers
correctly for absent, present, wrong-password, disabled and deleted.
Changes: mesh-sdk PR #7 (0.1.1) and mesh-catalog PR #84.
This is principle 1 applied to one loop. It survives the move unchanged. On NATS, the harness's
record of what it applied moves from memory into a key-value bucket, and `holds` stays as it is.
## Next, in order of silent failures removed
1. **`holds` for the other credential providers.** Each backend can answer whether a login exists
with the mesh's password without changing anything. Where a backend cannot check a password without
logging in, logging in is the check.
2. **The harness's other blind spot.** A consumer that goes away while its provisioner is down is never
removed. The fix is the same principle in reverse: list what the backend holds, and compare it with
what the mesh asks for. Removal stays subject to ADR 0114's rule that it never follows from a login
changing.
3. **`status` reports what owners already know.** The host knows which containers it recreated and
why. A rotation knows who it waits on. Delivery knows what is outstanding. Surfacing those as
conditions in the existing `status` needs no new transport. It is the shape the NATS condition store
will hold.
4. **The loop inventory and the classification** in [00](00-overview.md). They decide what comes after
these three.
## Not now
- Heartbeats, observation subjects, key-value state, advisories: all NATS, all after the move.
- Handing conditions to agents: designed with NATS, where a task on a subject is native.
- Choosing the observability store: decided when there is something to store, which is after the
move.
@@ -1,6 +1,6 @@
---
topic: what runs on it
status: proposed
status: accepted
date: 2026-09-01
deciders: jochen
reconstructed: false
@@ -76,6 +76,22 @@ three relationships, one broker, one runtime, all declared on the manifest.
- The runtime must dispatch a module's event handlers as well as its tools; that generalisation is
small (both arrive by importing the module's entrypoint) but it is real work.
## Progressive insight
> **Progressive insight — 2026-09-26.** *"No provisioner and no per-consumer setup" was a fact
> about the transport, and the transport changed.* This record's table says an event's machinery is
> "nothing but the broker's topic routing", and the text that an event needs "no per-consumer setup
> — only a subscription". That was true of a topic exchange, where a binding cost nothing and the
> broker fanned out. On NATS
> ([ADR 0106](0106-the-bus-is-nats.md)) a subscription is a **durable consumer**: a real object
> with a name, an ack policy, a delivery limit and its own ack subject, created when a module is
> assigned and removed when it is not. Per-consumer setup exists, and the controller does it.
>
> The decision is untouched — events are declared on both sides, 1:many, credential-free, and
> still provisioning's lighter sibling; the lightness is now relative rather than absolute.
> [ADR 0126](0126-a-module-declares-its-own-seats.md) adds the relationship this record's two
> columns had no room for: work addressed to a role, where exactly one holder must act.
## References
- [ADR 0002](0002-nodes-communicate-over-a-broker.md) — the broker events ride.
+24
View File
@@ -75,6 +75,30 @@ after its deliveries are exhausted; a module's account cannot publish outside it
subscribe outside its `consumes`. Then the cutover bed: a mesh on AMQP with the predecessor's
compatibility broker beside it moves its bus in one rollout with every node reporting afterwards.
## Progressive insight
> **Progressive insight — 2026-09-26.** *The compatibility broker was not single-purpose when this
> was written.* This record says the adopted AMQP broker is "kept as a module with one purpose —
> the predecessor's clients". Two modules of the new mesh also depended on it, through a `requires:
> ["amqp"]` grant its provisioner answered with a private vhost — `amqp-ping` and
> `amqp-email-forwarder`. On the retirement condition below, both would have been left requiring
> something no provider answers.
> [ADR 0125](0125-the-bus-is-the-only-broker.md) resolves it by moving them onto the bus and
> retiring the interface, which makes this record's sentence true rather than merely intended. The
> decision — the bus is NATS, the AMQP broker becomes the predecessor's compatibility broker and
> retires with the last of them — is unchanged.
> **Progressive insight — 2026-09-26, correcting the one above.** *The broker is not a
> compatibility module at all, and the sentence does not become true.* The insight above said
> [ADR 0125](0125-the-bus-is-the-only-broker.md) would make "one purpose — the predecessor's
> clients" true by moving the mesh's own modules off it.
> [ADR 0127](0127-amqp-is-a-provision-not-the-bus.md) supersedes that: nothing moves off, because
> a module may legitimately need an AMQP broker as a backing service the way it needs a database.
> The broker becomes **an ordinary provider module** — no seat, not foundation, not raised at
> genesis, and with no retirement condition, because the day its last client disappears is not a
> day anything is waiting for. What this record decided — the mesh's bus is NATS — is untouched
> by both; what was wrong was the sentence describing what happens to the old server, twice.
## References
- [research 014](../01-RESEARCH/014-the-bus-on-nats/00-overview.md)
@@ -0,0 +1,71 @@
---
topic: building it
status: accepted
date: 2026-09-24
deciders: jochen
reconstructed: false
---
# 107. Persistent data is a directory bind, never a named volume
## Context
Measured 2026-09-24, mid-migration: four modules mount a named Docker volume for real state —
`mesh-store` (every database the mesh holds), `mesh-broker` (its data and TLS material),
`mesh-registry` (every image), and `searxng`'s cache. Every other module in the catalogue —
more than forty — mounts a host directory, `/var/lib/<module>/...`.
The predecessor did not make this choice. HAL's own `postgres` bound `./db-data`, and its
`lavinmq` bound `./data` — directories, both. The mesh's adoption of them
([`a63ef3d`](https://git.novox.be/novox/mesh-catalog/commit/a63ef3d), "the postgres module
adopts mesh-store instead of raising its own", 2026-09-16) introduced the named volume; three
other modules followed the same shape since. Full account of what was found and fixed:
[issue 115](../04-ISSUES/115-a-named-docker-volume-is-invisible-and-one-flag-from-gone/00-report.md).
## The two guarantees are not the same guarantee
A named volume and a host directory both survive **ordinary** container recreation — a rebuild,
a `take`, the `docker rm -f` and push this migration already uses routinely. Neither loses data
to that. That was never the question.
What they do not both survive:
- **`docker rm -fv`, `docker volume rm`, `docker system prune --volumes`** all target a named
volume specifically. The first is one character from the command this migration's own rules
already call for after every address change. A host directory has no equivalent command that
destroys it by accident — removing it is always a deliberate `rm -rf` on a path someone typed.
- **Visibility.** Every tool this migration has used all night to find and verify data —
`ls`, `find`, `grep`, a backup job — reaches a host directory for free. A named volume requires
knowing to ask Docker (`docker volume inspect`) before its bytes, at
`/var/lib/docker/volumes/<name>/_data`, are reachable at all.
## Decision
**A container mount holding data that must survive is a host directory bind. A named volume is
permitted only for data that is disposable if lost** — a cache, a scratch space, something the
module rebuilds on next start without consequence. `searxng`'s `valkey` cache is close to this
line and was converted anyway, for consistency and because it costs nothing to.
Ownership is the one thing a host directory does not get for free that a named volume does:
Docker initialises a fresh named volume's ownership to what the container's first process needs;
a host directory is whatever created it. A directory made for this purpose must be given the
image's expected UID before the container using it starts — read from the running instance being
replaced when one exists, rather than guessed.
## Consequences
- The four modules were converted: `distribution`, `lavinmq`, `postgres`, `searxng`. Data copied
and verified byte-for-byte before each manifest changed; `mesh-store` stopped cleanly first, so
its copy is crash-consistent rather than a live read of a running postgres.
- **The ownership gap above was not theoretical — it is what happened.** The new directories
were created by the operator's tooling running as root; `mesh-store` crash-looped on
`mkdir: ... Permission denied` until its directory's ownership was set to match what the
original volume already had. Worth a check at `module add` time — nothing catches this today
beyond the container failing to start.
- Old named volumes were not deleted. They remain the rollback path until confidence in the new
mounts is established over time, not one clean start.
## References
- [Issue 115](../04-ISSUES/115-a-named-docker-volume-is-invisible-and-one-flag-from-gone/00-report.md)
- `mesh-catalog` PR #54
@@ -0,0 +1,102 @@
---
topic: the tiers
status: accepted
date: 2026-09-25
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0007-connectivity.md
---
# 108. A route carries the policy applied to a request, and names a secret rather than holding one
## Context
[ADR 0007](0007-connectivity.md) settled that **a route is a grant**: a module contributes the name
it wants and the port it listens on, and the proxy hands back the public name. That governs the
**grant**. It says nothing about what a request arriving at the name is permitted to do, and the
mesh's proxy currently permits everything: its request path is a host lookup and a forward, with no
authentication, no source check, no redirect handling and no middleware anywhere in it.
The standing requirement is that the mesh does **at minimum** what the system it replaces already
does. Measured against the predecessor's live configuration and its module catalogue
([issue 116](../04-ISSUES/116-route-proxy-has-no-auth-or-ip-restriction/00-report.md)), four
capabilities are relied on and absent:
| Capability | Dependents, counted |
|---|---|
| Authentication | **three** modules, each gating a credential-less admin surface — a key-value browser UI, a **database web UI**, and the replaced ingress's own dashboard |
| Refusal scoped to a path | **one**, and it is a live **incident mitigation** closing a write primitive that was abused |
| Path-scoped routing with priority | **two** — the refusal above, and a certificate-challenge path on a host that otherwise routes elsewhere |
| Redirect | **two** live routes canonicalising a `www` name onto its apex |
Counting needed two sources and neither alone is complete: the catalogue cannot see what was
hand-written on a node, and a node's configuration directory cannot see what modules declare as
container labels. An earlier count read one source and undercounted authentication by two.
**The third row is a prerequisite, not a sibling.** The proxy's table maps a host to exactly one
target, so a host cannot be routed two ways. Adding authentication and a source filter would not
make the refusal rule expressible.
## Considered Options
1. **Leave policy out of the mesh; keep the affected routes on the adopted ingress.** *Rejected.*
The mesh would run two reverse proxies indefinitely with no principled division between them, and
one of the routes held back is a live mitigation — leaving it on a component being decommissioned
means its removal date is whenever somebody forgets.
2. **A separate, proxy-side settings layer keyed by route name.** *Rejected.* It keeps the grant
literally clean, but answering *"what protects this route"* then requires reading two files that
nothing keeps in step. A route's protection is part of what a route is.
3. **The grant carries a reference; the detail lives in a second layer.** *Rejected.* Both costs of
option 2, plus a naming indirection to maintain.
4. **Carry the password hash in the declaration.** *Rejected.* It would be the first credential
value in a declaration, and a precedent is easier to set than to withdraw. A hash is not a
plaintext password, but it is sufficient to pass the gate it protects.
5. **An open middleware surface the proxy applies.** *Rejected.* It recreates the thing being
replaced, makes the proxy's behaviour unbounded, and an open surface is far harder to narrow later
than a closed one is to widen.
## Decision
**A route contribution carries the policy applied to requests arriving at its name**, alongside the
name, the port and the location it already carries.
**The set is closed, and it is these four:** authentication; refusal scoped to a path; path-scoped
routing with priority; redirect. A fifth is an amendment to this record, deliberately — each
addition should be earned by a dependent that exists.
**Where policy needs a credential, the declaration names a secret the mesh mints and holds. It never
carries the value.** This keeps the existing secret machinery as the only thing that holds
credentials, and keeps hashes out of anything regenerated, synced or committed.
**Consequently the routing table is keyed by host and path, with priority** — not by host alone.
This follows from the decision rather than being a separate one: two of the four capabilities need a
single host routed more than one way.
## Consequences
**What this makes possible.** The affected routes can leave the adopted ingress, and "at minimum"
becomes a satisfiable claim rather than a standing exception. The incident mitigation gets a durable
home in the mesh, which its own note already asked for.
**What got harder.** The contribution shape grows, and every provider of `route` must understand
more of it. The table is no longer a flat map, and priority introduces ordering that has to be
deterministic rather than incidental — equal priorities must resolve the same way every time or the
proxy becomes non-reproducible. A closed set means a new need is a decision, not a patch.
**What does not change.** The proxy remains a reference implementation: the contract is the file the
mesh writes, not the program that reads it. Another proxy may implement the same file.
**How this is checked.** A rule states how it is verified, so this one does. A lab bed must show,
against a mesh that declared them: a route with authentication refusing an unauthenticated request
and admitting an authenticated one; a path-scoped refusal shadowing an ordinary route on the same
host while that ordinary route still serves every other path; a redirect answering with the
redirect; and — the negative case, which is the one that rots quietly — **a declaration carrying a
credential value rather than a reference being refused**, so option 4 cannot return by accident.
## References
- [Issue 116](../04-ISSUES/116-route-proxy-has-no-auth-or-ip-restriction/00-report.md) — the count,
the evidence, and the structural finding about the table.
- [ADR 0007](0007-connectivity.md) — a route is a grant. This record extends it to the request.
- [ADR 0009](0009-modules-and-the-graph.md) — the provision vocabulary a contribution belongs to.
- [To-be 08 §3](../03-DESIGN/01-to-be/08-connectivity.md) — where exposure is specified.
@@ -0,0 +1,106 @@
---
topic: the tiers
status: accepted
date: 2026-09-25
deciders: jochen
reconstructed: false
extends: 0075-two-stores-and-which-provides-what.md
---
# 109. A package registry seat is one per ecosystem, not one for all of them
## Context
Fixing `builder`'s consumption of `package-registry` tonight surfaced the shape ADR 0075 actually
left implicit. 0075 split `artifact-store` from `package-registry` and said the second is "an
ecosystem's own registry — npm, cargo, PyPI, Go" — but it defined one provision for all four,
not one each.
What that produces, read from the manifests as they stand:
- `gitea`'s `module.json` declares `provides: package-registry` once, and its `serves` block
carries exactly one path: `npm-path`. Nothing names a cargo or PyPI endpoint, though gitea's own
package API serves both.
- `builder` had, until tonight, a hand-written JSON fragment standing in for a real grant —
`{"provision": "package-registry", "from": "gitea", "at": "127.0.0.1", ...}` — because nothing
in the interface gave it a real one to ask for. The fragment named `npm-path` specifically; there
was nowhere to put a second ecosystem's endpoint even if one had been wired.
- The fix applied tonight declares `requires: ["artifact-store", "package-registry"]` and lets the
mesh mint the grant properly — correct for what exists today, but it is one seat standing in for
what should be several, the same conflation 0075 itself named and did not resolve for this
provision specifically.
**The version-skew problem 0075 wrote down for the artifact-store/package-registry split repeats
one level down, inside "package registry" itself:** npm resolves by name and range from one
namespace, cargo from another, and a single grant conflates them exactly the way one store for
both digests and ranges would have.
## Considered Options
**1. One `package-registry` provision, gitea answers every ecosystem it can.** What exists today.
Simplest to grant — one credential, one binding, done once per consumer. Rejected: a consumer
that only ever needs npm still receives a grant shaped to cover cargo and PyPI, and there is no way
to hand off *only* npm to a different provider (verdaccio, say) without renegotiating the whole
provision for every consumer of any ecosystem.
**2. One provision, parameterised by ecosystem.** `requires: package-registry` plus a declared
`ecosystem: npm` alongside it, still one interface. Rejected: the `provides`/`requires` refusal
mechanism this mesh already uses (two providers of one provision is a naming conflict until
resolved) would need to become conditional on a parameter it does not otherwise carry anywhere in
the mesh's resolution — a special case for exactly one provision, rather than the mesh's existing
mechanism applied again.
**3. One provision per ecosystem — `npm-package-registry`, `cargo-package-registry`,
`docker-package-registry`, and so on, each independently `provides`/`requires`.** Chosen.
## Decision
**A package registry seat is one per ecosystem.** `npm-package-registry`, `cargo-package-registry`,
`docker-package-registry` — each its own provision, resolved, granted, and refused exactly the way
`artifact-store` and today's single `package-registry` already are. Adding an ecosystem is adding a
provision, not widening one.
**Gitea may hold several seats at once.** Nothing here says gitea answers only one; ADR 0075
already established that a provider may answer more than one named thing on one machine ("a mesh
running gitea for git and packages alongside a registry serving artifacts is an ordinary
arrangement"). Gitea fulfilling `npm-package-registry` and `cargo-package-registry` both is the
expected shape, not an exception.
**Each seat's grant is independent.** A consumer that only needs npm holds only the
`npm-package-registry` grant. Moving that one ecosystem to a different provider — verdaccio,
named directly as the motivating case — means assigning `npm-package-registry` to verdaccio and
leaving every other seat exactly where it was. No consumer of `cargo-package-registry` observes
the change; no manifest naming `package-registry` broadly needs to be found and re-read.
**`builder`'s fix tonight is the interim shape, not the target.** It correctly consumes the one
seat that exists today (`package-registry`, npm in practice). Splitting it becomes, later,
replacing that one line with the ecosystems `builder` actually uses — a manifest change, not a
redesign of how `builder` asks for anything.
## Consequences
- `gitea`'s `module.json` gains a `provides` entry per ecosystem it actually serves, each with its
own `serves` block (`npm-path`, a cargo path, a PyPI path) in place of the one `package-registry`
entry with a single `npm-path` inside it.
- `gitea`'s provisioner (`modules/gitea/provisioner/index.ts`) currently runs one `runProvisioner`
registration for `package-registry`; each seat needs its own registration, or one provisioner
keyed by which seat's `create`/`remove` fired — the mesh's `Provision` type does not yet carry
which named provision a call is for when a module answers more than one, and that is worth
checking before assuming the harness already supports it.
- Every consumer's `requires` moves from the one name to however many ecosystems it actually uses.
`builder` is the only known consumer today; widening later is one manifest line per module, not
a migration.
- `verdaccio`'s role sharpens: not "package-registry, an alternative for npm alone" (0075's phrasing)
but a named `npm-package-registry` *provider*, a straight swap against gitea's answer to the same
seat.
- Not solved here: whether `cargo-package-registry` and `pypi-package-registry` are needed at all
before something actually consumes them. This record names the shape; building unused seats is
its own decision.
## References
- [ADR 0075](0075-two-stores-and-which-provides-what.md) — the record this extends; defined
`package-registry` as the second provision without splitting it per ecosystem.
- `mesh-catalog modules/builder/module.json` — tonight's fix, the interim single-seat shape.
- `mesh-catalog modules/gitea/module.json`, `modules/gitea/provisioner/index.ts` — today's
single-provision, npm-only implementation.
@@ -0,0 +1,230 @@
---
topic: what runs on it
status: accepted
date: 2026-09-25
deciders: jochen
reconstructed: false
extends: 0009-modules-and-the-graph.md
---
# 110. A seat is held by one assignment, from a closed set, and it may deliver a provision
> **Narrowed, not replaced — 2026-09-27, on merging two lines of work.** This was marked superseded by
> [ADR 0126](0126-a-module-declares-its-own-seats.md), and that overstated it: 0126 says in as many
> words that *"everything 0110 decided about what a seat is stands untouched"*. What moved is where the
> set lives and who may add to it —
> [0126](0126-a-module-declares-its-own-seats.md) lets a module declare one and makes the set derived,
> [0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md) names the mesh's own
> for their scope, and [0122](0122-a-seat-is-data-a-rename-is-a-database-update.md) moves them out of
> code into a table. **What a seat *is* — one holder at its scope, a definition saying what a module
> can hold against an assignment saying what it does, a role made singular rather than a module — is
> this record and still current**, which is why those three rest on it.
## Context
[ADR 0009](0009-modules-and-the-graph.md) introduced claims: a module declares something
singular, at a scope, and a second holder is refused. [ADR 0079](0079-the-foundation-seats-are-named-after-their-servers.md)
named the foundation's three after their servers. That mechanism is enforced and works. What it
means has drifted, and four things are now true of it that no record says.
**Any well-formed name becomes a seat by being claimed.** The controller's manifest check
refuses a claim only for a malformed name or an unknown scope. Nothing says which seats a mesh has.
The names in use were each invented by the module that claims them: `the-showcase`,
`the-build-machine`, `the-intrusion-prevention`.
**Nothing can say what a mesh has, or who fills it.** There is no seat table and no command that
lists seats. Holdings are assembled while planning, one node at a time, and discarded afterwards.
The only way to answer "which seats does this mesh have, and which module holds each" is to read
every manifest in two repositories, because the controller's own manifest lives in its own
repository ([ADR 0069](0069-a-module-is-a-repository-and-a-path.md)), and then the controller's code,
because one module it ships has its manifest composed there. While this record was being prepared,
that enumeration was done by hand, and it missed both of the last two sources: eleven claims were
reported where there are thirteen.
**A claim in a definition makes a module singular, not a role.** The store module's definition
claims `mesh-store`, so every assignment of it claims the seat, and a second store module on any other
node is refused. What is singular is *the store the mesh itself uses*, not postgres. Any module can
run on any node whose capabilities match, which is a core principle of the module system, and a claim
written into the definition breaks it for every module that claims anything.
**Some seats are the mesh's one of something everyone consumes, and nothing uses that fact.** A
requirement for a mesh-scoped provision with more than one provider is refused until a person pins,
**per consumer node**, which provider to use. [ADR 0109](0109-a-package-registry-seat-is-one-per-ecosystem.md)
anticipates exactly that case, gitea and verdaccio both answering npm, and under today's resolution it
would mean a pin on every machine that builds anything.
## Considered Options
**1. Leave seats as free-form exclusion, claimed in definitions.** Rejected. The overview stays
unanswerable, a module that claims a seat can run on only one node, and a second provider of anything
costs a pin per consumer node.
**2. Two concepts: seats for exclusion, and a new word for the mesh's one consumable thing.**
Rejected. Both mean "this mesh's one X". Every existing claim would first have to be classified into
one or the other, and the overview a person wants is one list, not two.
**3. A seat is held by one assignment, from a closed set, and holding it may deliver a provision.**
Chosen.
## Decision
**The mesh defines a closed set of seats.** Each entry has a name, a scope, what holding it delivers
(if anything), and the decision that made it a seat. A seat outside the set is refused wherever it
is named. Adding a seat is a decision, for the same reason adding a shape to the host's vocabulary is
one: the set is what a person reads to learn what a mesh can have, and a name added without an
argument is a name nobody can explain later.
**A definition says which seats a module *can* hold. An assignment says which it *does* hold.** The
store module can hold `mesh-store`, and it may be assigned to every node. Exactly one of those
assignments holds the seat, because that assignment said so. A second assignment saying so, at the
seat's scope, is refused. So a seat makes a *role* singular, never a module, and moving the role is
changing which assignment holds it, with no definition changed and nothing unassigned.
**What the mesh knows about a seat's holder is what it knows about that assignment**: its node, the
node's settings for it, and what it serves. Holdings are not stored separately. The seat points at an
assignment, and a second record of the same fact would be a second thing to disagree with the first.
**Holding a seat may deliver a provision.** A seat that delivers a provision may only be held by an
assignment of a module that provides it, at the seat's scope.
**A requirement may name a seat, and then the seat's holder answers it.** Naming the seat asks for
*the mesh's* one, not for whichever provider is nearest, so the holder answers **even when another
provider runs on the consumer's own node**, and nothing is asked of anyone. Unheld, the requirement is
refused, naming the seat. A builder asks for the mesh's npm registry this way, and is served by the
holder of `npm-package-registry` wherever it runs, with no pin on any machine.
**A requirement that names no seat resolves as [ADR 0084](0084-which-provider-serves-a-consumer.md)
has it**: a pin, then the provider on the consumer's own node, then the only provider. Where several
remain and none is local, **a person chooses when the module is assigned**. Assignment lists the
candidates, with the holder of a seat that delivers the provision suggested first, and records the
answer on the assignment as its pin. Without an answer the module is not assigned. This keeps
[ADR 0009](0009-modules-and-the-graph.md)'s rule that a requirement with several answers is never
guessed. The choice is made either by the requirement naming the seat, or by a person at assignment,
and never silently by what happens to run nearby. That is the failure
[issue 106](../04-ISSUES/106-the-vault-claims-no-seat/00-report.md) names for the vault.
**A seat delivers a provision only where the mesh has one answer for everyone.** The artifact store,
the npm registry, git and the vault are each one per mesh by their own records, so their seats
deliver them.
**The foundation's seats deliver nothing.** `mesh-controller`, `mesh-store` and `mesh-broker` name
which assignment the mesh *itself* uses: the controller, the store holding its records, the broker
carrying its bus. The store and broker modules may run on other nodes too, and a database or `amqp`
consumer that names no seat is served by co-location from whichever runs on its own node, the seat's
holder included. Were `mesh-store` to deliver, a consumer could name it and be sent to the store the
mesh keeps its own records in. That is not a store for consumers.
**A seat may reserve its provision.** Where a second provider would break the reason the provision
exists, only an assignment holding the seat may provide it at all: the parser refuses a definition
that provides it without being able to hold the seat, resolution refuses an assignment providing it
without holding the seat, and a pin cannot choose anyone else: a requirement for it always names the seat. `secret` is the one
reserved provision.
The vault is one per mesh because a second *"would be a second place to lose"*
([ADR 0085](0085-a-secret-is-a-provision.md), as amended), and a second `secret` provider is exactly
that, whether a pin chose it or not.
**Seats are also informational.** The controller lists every seat in the set, what it delivers, and
which assignment holds it, including seats nobody holds. An unheld seat is an answer, "this mesh has
no X", not an error.
**The first set is the twelve seats already claimed, plus two.** Thirteen claims are in use, and they
name twelve seats because two alternative modules claim `the-resolver-configuration`. This record
admits every seat the catalogue and the controller claim today, so no definition is refused by it:
| seat | scope | delivers | can be held by | made a seat by |
|---|---|---|---|---|
| `mesh-controller` | mesh | — | `mesh-controller` | [0079](0079-the-foundation-seats-are-named-after-their-servers.md) |
| `mesh-store` | mesh | — | `postgres` | [0079](0079-the-foundation-seats-are-named-after-their-servers.md) |
| `mesh-broker` | mesh | — | `lavinmq` | [0079](0079-the-foundation-seats-are-named-after-their-servers.md) |
| `mesh-vault` | mesh | `secret`, reserved | `mesh-vault` | this record, for [issue 106](../04-ISSUES/106-the-vault-claims-no-seat/00-report.md) |
| `the-artifact-store` | mesh | `artifact-store` | `distribution` | [0075](0075-two-stores-and-which-provides-what.md) |
| `the-catalogue` | mesh | — | `mesh-catalog` | this record |
| `npm-package-registry` | mesh | `npm-package-registry` | `gitea` | [0109](0109-a-package-registry-seat-is-one-per-ecosystem.md) |
| `the-build-machine` | node | — | `builder` | this record |
| `the-dns-port` | node | — | `dnsmasq` | this record |
| `the-intrusion-prevention` | node | — | `fail2ban` | this record |
| `the-packet-filter` | node | — | `nftables` | this record |
| `the-private-network` | node | — | the controller's private-network module | this record |
| `the-resolver-configuration` | node | — | `resolv-conf` or `resolved-split-dns` | this record |
| `the-showcase` | node | — | `showcase` | this record |
There are two additions. `npm-package-registry` is [ADR 0109](0109-a-package-registry-seat-is-one-per-ecosystem.md)'s
seat, named after the provision it delivers, as 0079 named the foundation's seats after what they are.
A gitea assignment holds it. verdaccio provides the same provision and cannot hold the seat, so it is
the second provider this record exists to make harmless. Moving npm to it would take a definition
saying it can hold the seat, and then an assignment saying it does.
`mesh-vault` answers [issue 106](../04-ISSUES/106-the-vault-claims-no-seat/00-report.md). The vault is one
per mesh ([ADR 0085](0085-a-secret-is-a-provision.md), as amended), and until now that was enforced by
nothing. The seat is named after its server, by the 0079 convention.
`the-dns-port` is listed as delivering nothing, although `dnsmasq` provides `wildcard-resolution`.
That provision is node-scoped and answered on the machine, so no preference between providers arises.
## What this changes in earlier records
On acceptance, each of these is amended by this record, not edited:
- [ADR 0009](0009-modules-and-the-graph.md): a claim in a definition says a module *can* hold a seat;
the assignment says it does.
- [ADR 0079](0079-the-foundation-seats-are-named-after-their-servers.md) and
[ADR 0078](0078-the-store-and-broker-are-modules.md): "a mesh runs one postgres and one
lavinmq" becomes one holder of `mesh-store` and one of `mesh-broker`. The store and broker modules may
run on other nodes.
- [ADR 0084](0084-which-provider-serves-a-consumer.md): a requirement may name a seat, which its holder
answers; and where several providers remain and none is local, the choice is asked when the module is
assigned and recorded as a pin, rather than refused until someone pins it.
- [ADR 0109](0109-a-package-registry-seat-is-one-per-ecosystem.md): one provision per package
ecosystem stands. Where 0109 says *seat*, it means that provision. Only `npm-package-registry` is
also a seat in this set. A cargo or docker registry becomes one by a record, as any seat does.
"Gitea may hold several seats" reads: gitea may provide several ecosystems, and hold the seat of
each one that is a seat. Moving npm to verdaccio is not "assigning `npm-package-registry` to
verdaccio". It takes verdaccio's definition saying it can hold the seat, and then an assignment
holding it.
- [To-be 23](../03-DESIGN/01-to-be/23-choosing-a-provider.md): the same two changes, in the design
that describes choosing a provider.
- [To-be 21](../03-DESIGN/01-to-be/21-the-installation-in-full.md): genesis assigns the foundation's
store, broker and controller holding their seats, where their definitions claim them today.
## Consequences
- The controller carries the set in code. A test asserts its size, and that every entry names the
record that made it a seat, so changing the set means finding the argument rather than a number.
- An assignment gains the seats it holds. Genesis assigns the foundation's store, broker and
controller holding their seats, where today their definitions claim them.
- Manifest validation refuses an unknown seat, a seat named at the wrong scope, and a delivering seat
named by a module that does not provide the provision. Resolution refuses a second holder, and an
assignment holding a seat its module cannot hold.
- Resolution answers a requirement naming a seat with its holder. Assignment asks a person where
several providers remain, suggesting the seat's holder first, and records the answer as a pin. A
provider record gains the module it came from.
- A `seats` command lists the set with each seat's holder, derived from assignments.
- **What got harder:** a module wanting a new singular role can no longer invent a name. It needs a
record. And an assignment has one more thing to say. Both are the point.
- **Not changed:** the controller's seat placeholder stays as it is. It exists so the controller can
reach a foundation it made before any module existed.
## How it is checked
| Rule | Checked by |
|---|---|
| The set is closed, and every entry names its decision | A controller unit test asserts the set's size and a non-empty decision for every entry. |
| A seat outside the set is refused | Manifest-validation tests for an unknown seat, the wrong scope, and a delivering seat named by a module that does not provide it. |
| Every module in use names a seat in the set | A controller test parses every catalogue manifest and fails on any refused seat. The lab's beds read the same manifests ([ADR 0089](0089-a-bed-reads-the-catalogue-it-proves.md)). |
| A seat is held by one assignment, not by a module | A resolution test: the store module assigned to two nodes resolves, with one assignment holding `mesh-store`; a second assignment asking to hold it is refused. |
| A requirement naming a seat is answered by its holder | Resolution tests: two providers with the seat held; a second provider on the consumer's own node, where the holder still answers; the seat unheld with two providers, and with **one** provider, both refused naming the seat. |
| An assignment holds only a seat its module can hold | A resolution test: an assignment holding a seat its definition does not name is refused. |
| Every seat is listed with its holder | A `seats` command test: every seat in the set is listed with its scope, what it delivers and its holder, and an unheld seat is listed as unheld. |
| Several providers and none local is a person's choice | An assignment test: the candidates are listed with the delivering seat's holder first; the answer is recorded as a pin; with no answer the module is not assigned. |
| The foundation's seats route nobody | A resolution test: with the store module on two nodes, a database consumer is served by the one on its own node, whichever holds `mesh-store`; a requirement naming `mesh-store` is refused, because it delivers nothing. |
| A reserved provision has no other provider | The parser refuses a definition providing `secret` that cannot hold `mesh-vault`; resolution refuses an assignment providing it without holding the seat, and a pin on a `secret` requirement. |
## References
- [ADR 0009](0009-modules-and-the-graph.md): claims, scopes, and "refused, never guessed"
- [ADR 0079](0079-the-foundation-seats-are-named-after-their-servers.md): a seat named after what it is
- [ADR 0109](0109-a-package-registry-seat-is-one-per-ecosystem.md): the per-ecosystem registry seats
- [ADR 0084](0084-which-provider-serves-a-consumer.md), [to-be 23](../03-DESIGN/01-to-be/23-choosing-a-provider.md):
pins, co-location and refusal
- `mesh-controller internal/catalogue/resolve.go` (`checkClaims`, the brokered branch of `Resolve`),
`internal/catalogue/manifest.go` (claim validation), `cmd/mesh-controller/plan.go` (holdings)
@@ -0,0 +1,105 @@
---
topic: building it
status: accepted
date: 2026-09-25
deciders: jochen
reconstructed: false
extends: 0069-a-module-is-a-repository-and-a-path.md
---
# 111. A build source is on the mesh's git seat, or it is an external repository
## Context
[ADR 0069](0069-a-module-is-a-repository-and-a-path.md) made a module a repository, a path and a
ref, and the controller records all three against the module so it can rebuild it and say when
its source has moved ahead. **The repository is recorded exactly as a person typed it.** `build
<repository>` hands the string to a build machine, which runs `git clone` on it, and the same string
becomes the module's recorded source.
**So a self-hosted forge's address is written into every module built from it.** The mesh runs its
own forge, and most of what it builds lives there. Every one of those modules carries the forge's
scheme, host and port in its recorded source. Move the forge to another machine, or change the port
it is published on, and every recorded source is stale at once. Nothing notices until a rebuild fails
to clone.
**And nothing names the mesh's git at all.** gitea serves git over HTTP and over SSH, and the mesh's
vocabulary contains neither. No provision, no `serves`, no seat, as the forge survey
([research 013](../01-RESEARCH/013-the-forge-and-the-registries/the-survey.md)) found. The only trace
is a label on its public route, which the mesh is explicitly not meant to interpret.
**External repositories are ordinary, and must stay so.** An application the mesh hosts may live on a
public forge. Building it from its URL works today and must keep working unchanged.
## Considered Options
**1. Keep recording literal URLs.** Rejected. It is the problem: the forge's address copied into
every module built from it.
**2. Recognise a self-hosted source by matching its URL against the forge's current address.**
Rejected. It infers the kind of source from the shape of a string, and the inference fails in the
one case it exists for: after the forge moves, old URLs no longer match anything.
**3. Two explicit forms: a repository on the holder of the `git` seat, or an external URL.** Chosen.
## Decision
**The mesh has a `git` seat.** It is mesh-scoped and delivers the `git` provision, per
[ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md). Its holder provides `git`,
serving how a repository on it is cloned: the scheme and the port. A gitea assignment holds it.
**A source is on the git seat, or it is external, and the mesh records which.**
- `build --self <owner>/<repository>` builds from a repository on the seat's holder. The recorded
source is the repository's path on that holder, and the seat it is on. **It never contains an
address.** At the moment of building, the controller composes the clone URL from where the
holder runs and what it serves for `git`, so a moved forge changes nothing recorded.
- `build <url>` is unchanged: an external repository, recorded and cloned exactly as given. GitHub
and GitLab are the ordinary cases.
**An unheld seat refuses self-hosted builds and nothing else.** With nobody holding `git`, `build
--self` is refused, naming the seat and saying what would hold it. External builds are unaffected. A
mesh without a forge of its own builds from external repositories only, and says so rather than
failing to clone.
**The build machine is not told the difference.** It receives a URL either way. Composing the URL is
the controller's job, because only the controller knows where the seat's holder runs.
## What this changes in earlier records
On acceptance, each of these is amended by this record, not edited:
- [ADR 0069](0069-a-module-is-a-repository-and-a-path.md): a module's repository is recorded either as
a path on the `git` seat or as an external URL, never as an address of the mesh's own forge.
- [ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md): the closed set gains `git`,
mesh-scoped, delivering `git`, held by a gitea assignment.
## Consequences
- The controller's inventory gains a column saying which seat a source is on. It is empty for
every module recorded before this, which is correct: they were all recorded as literal URLs.
- `build`, `build --behind` and `build --dry-run` resolve a seat source before asking a builder. The
recorded source keeps the seat form; the build log keeps the URL that was actually cloned, because
that is what happened.
- gitea can hold `git` and provides it, serving HTTP clone on its web port; the forge's assignment holds the seat.
- **Not decided: credentials for private repositories.** The mesh's own repositories are public, and
clone without one. A private repository still works only if the build machine's own git
configuration authenticates, exactly as before. Delivering a clone credential through the `git`
provision's grant is the obvious next step, and it is its own decision.
- **Not changed:** modules already recorded from the forge keep their literal URLs until they are
rebuilt with `--self`. Rewriting them in place would be the URL-matching this record rejects.
## How it is checked
| Rule | Checked by |
|---|---|
| A seat source records no address | A controller test resolves a seat source and asserts the recorded repository is the path alone. |
| The URL comes from the holder | A test composes the clone URL from a holder's node address and served `git` scheme and port, and a second where the port was moved on the node. |
| An unheld seat refuses self-hosted builds only | A test with no holder: `--self` is refused naming the seat; an external URL passes through unchanged. |
## References
- [ADR 0069](0069-a-module-is-a-repository-and-a-path.md): a module is a repository, a path and a ref
- [ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md): seats, and a seat delivering a provision
- [Research 013](../01-RESEARCH/013-the-forge-and-the-registries/the-survey.md): git is served and declared nowhere
- `mesh-controller cmd/mesh-controller/build.go`, `internal/builder/builder.go`
@@ -0,0 +1,192 @@
---
topic: what runs on it
status: proposed
date: 2026-09-25
deciders: jochen
reconstructed: false
extends: 0046-a-module-configuration-is-its-assignments-not-its-manifest.md
---
# 112. A module definition names no node, no mesh and no path: everything it needs is a requirement the mesh resolves
## Context
[Issue 119](../04-ISSUES/119-a-module-definition-decides-where-its-files-live/00-report.md) found
**789 host-path strings in 70 of the catalogue's 71 module definitions.** Each definition chooses
where on the machine its directories, mounts, bindings, secrets, env-files and received files live,
and often repeats that path in an environment variable or in code. Mounts are checked against what
the definition declares ([ADR 0091](0091-a-mount-is-declared-three-ways.md)); nothing checks the
copies. The issue records what that has already allowed:
- a provider that would provision nobody without a word;
- a contributions file that carries host paths into containers, so every provider must mount its
grants directory at the identical path;
- defaults in code that disagree with their own manifests;
- every identity keyed by the module's name, which is why one module cannot be assigned to one node
twice. This record keeps that, and says so below.
**Paths are one case of a wider pattern.** A module gets what it needs through at least six separate
mechanisms today, each with its own syntax and its own failure modes:
- provisions, read through bindings;
- settings on the assignment ([ADR 0046](0046-a-module-configuration-is-its-assignments-not-its-manifest.md));
- ports the mesh assigns ([ADR 0038](0038-the-mesh-assigns-the-port.md));
- machine facts a manifest asks for;
- secrets, either minted or accepted from an operator;
- literals carried in the definition itself.
The mesh has already unified parts of this. Ports became the mesh's rather than the module's (0038),
configuration became the assignment's (0046), and which provider serves a consumer became the
assignment's choice ([ADR 0084](0084-which-provider-serves-a-consumer.md)). What remains is the
concept that joins them.
## Considered Options
**1. Keep host paths in definitions, and check that every copy agrees.** Rejected. It checks the
agreement of something that should not be there. A definition still could not follow its data to
another disk, or be adopted onto a machine whose data is already somewhere.
**2. Keep the separate mechanisms, and add directories as a seventh.** Rejected. It fixes paths and
keeps the pattern that produced them: each mechanism is resolved, validated and refused differently,
so a module author learns six systems and a reviewer checks six kinds of gap.
**3. One concept: a module requires, and the mesh resolves every requirement against a contract.**
Chosen.
## Decision
**A module definition is node-agnostic and mesh-agnostic.** It names no node, no mesh and no host
path. **Everything a module needs is a requirement**: a name, a contract saying what the module may
read from it, and which kind of provider answers it.
**Installing a module on a node resolves every requirement, or refuses.** A refusal names each
unresolved requirement and what could answer it, all at once.
**There are four kinds of provider, and the set is closed:**
| provider | answers | today's mechanism it replaces |
|---|---|---|
| **another module** | a database, a bucket, a vhost, a secret, a route | provisions and bindings |
| **the node's host** | a directory, a port, facts about the machine | resource paths, `${port:}`, `${machine:}`, facts |
| **the mesh** | the module's identity and names, and the delivery of every answer | derived logins and generated names; the controller's delivery |
| **the operator, through the assignment** | a value a person chooses that is not secret: a public name, a greeting, a number of workers | settings, carried literals |
A module provider is chosen as [ADR 0084](0084-which-provider-serves-a-consumer.md) and
[ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md) say. A requirement naming a seat is
answered by its holder. Otherwise it is a pin, then co-location, then the only provider, and where
several remain, a person chooses at assignment and the choice is recorded as a pin.
A host provider is always the module's own node, because a host path or a port means nothing on any
other. An operator value is the assignment's, or the requirement's default, or unresolved.
**A person's value stays cheap.** An operator requirement's contract is a type and, optionally, a
default. It needs no provider module, no grant and no credential.
**Every shared secret is a `secret` requirement, answered by the vault**, with no exception by kind
([ADR 0113](0113-the-vault-makes-every-secret.md)). A private key is made where it is used and is not
a requirement. An external API key an operator chooses is no
different: the operator delivers it to the vault ([ADR 0092](0092-an-operator-delivers-a-pair-credential.md)),
and the module requires a `secret` like any other. A provider that needs a secret for a consumer
requires it from the vault, like any consumer, and answers with resources and data. The mesh carries
every answer back.
**A directory is a host provision.** Its contract is the owner and mode the module needs, including
the owner its image expects ([ADR 0107](0107-persistent-data-is-a-directory-bind-never-a-named-volume.md)).
It carries no persistence flag. A directory is kept while it holds anything
([ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md), which refused a `keep` flag for good
reason), and data that is disposable is not a directory but a named volume (0107). *Where* it is on
the machine is the assignment's. A node has a default layout, and an assignment may place a directory
elsewhere: on a second disk, or where an adopted machine's data already is
([ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md)).
**An operator's shared data stays an `access`** ([ADR 0051](0051-shared-data-is-the-operators.md)),
not a directory. 0051 rejected giving a directory an operator owner, because the mesh must never
create, chown or remove such data, and that stands. Only where its location is written changes: the
module requires read or read-write access, and the assignment says where the data is.
**Inside a container, a module sees its own paths.** The definition says where the image expects each
directory. The mesh mounts the assignment's location there. No host path is ever a value a process
reads, and the mesh's own files (answers, contributions) name nothing by host path, so a provider needs
no mount at a machine-identical path.
**A module is assigned at most once to a node.** An assignment is a module on a node, and that pair is
its identity: its directories, containers, login, broker account and settings are keyed by it, as
they are today. A module may run on many nodes, and one of those assignments may hold a seat
([ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md)). Running the same module twice
on one machine is not supported. The cases that seemed to need it, such as two stores of one engine or
two stages of one application, are different modules, or the same module on different machines. The
line is drawn because every identity in the mesh is already a module on a node, and a second
instance would have to rename all of them.
**What must stay singular stays so** by a seat, or by an operator value colliding: a public name
already held by another assignment is refused like any other singular thing.
**The foundation's first secrets are delivered, then adopted.** Genesis generates them before the vault
can run and hands them to the vault once it is installed, and from then on they are answered the same
way as every other secret ([ADR 0113](0113-the-vault-makes-every-secret.md)).
## What this changes in earlier records
On acceptance, each of these is amended by this record, not edited:
- [ADR 0046](0046-a-module-configuration-is-its-assignments-not-its-manifest.md): settings become
operator requirements on an assignment.
- [ADR 0038](0038-the-mesh-assigns-the-port.md): a port becomes a host requirement. What 0038 decided
is unchanged; it is the first case of this rule.
- [ADR 0051](0051-shared-data-is-the-operators.md): an access keeps its shape and its semantics; its
path moves from the definition to the assignment.
- [ADR 0091](0091-a-mount-is-declared-three-ways.md): a mount's host side is a resolved requirement,
checked as resolved rather than as a path the definition declares.
- [To-be 21](../03-DESIGN/01-to-be/21-the-installation-in-full.md): the step that builds and runs a
store module as a database provider, beside the foundation's store on the same node, would run the
store module twice on one node. The adopted store module ([ADR 0078](0078-the-store-and-broker-are-modules.md))
holds `mesh-store` and serves that node's database consumers by co-location, so there is no second
one.
- The [glossary](../00-META/glossary.md): *provision* widens from "a service one module provides" to a
requirement answered by any of the four providers, and *requirement* and *contract* are added. None
of it lands while this record is only proposed, because the glossary is the authority on the words
in use, not on words under review.
## Consequences
- **Every definition changes.** 70 of 71 name host paths today, and most use at least three of the
mechanisms this replaces. The change is mechanical for most. The design has to say how existing
modules move without their data moving: an adopted or already-running assignment is placed where
its data already is.
- The controller resolves every requirement at assignment and refuses unresolved ones. The host
answers directories and ports. The settings, placeholders, facts and bindings that exist today
are retired as separate mechanisms, once nothing uses them.
- Identity stays a module on a node. Nothing is renamed, and a login still fits the tightest backend
as [ADR 0049](0049-a-consumers-identity-fits-the-tightest-backend.md) arranges.
- **What got harder:** one module cannot run twice on one machine; a second stage or a second store
of one engine is a different module or a different machine. And a definition no longer says where
a module's data is on a machine, or what a setting's value is. The assignment does, and `plan` shows
it. That is the point, and it is also a real loss of at-a-glance legibility, which the overview has
to give back.
- **Not decided here:** the syntax a definition reads a requirement's fields with; a node's default
layout; the order in which the mechanisms are retired. [To-be 27](../03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md)
proposes all three.
## How it is checked
| Rule | Checked by |
|---|---|
| A definition names no host path | A catalogue test: the host side of every mount, and every resource location, is a requirement rather than an absolute path. A declared list of exceptions shrinks to empty as definitions move. |
| No host path is a value a process reads | A catalogue test: every absolute path in a container's environment or env-files lies on the container side of one of its mounts, or is declared the image's own. A second test finds literal paths in module code used as fallbacks for an environment variable. |
| A definition names no node and no mesh | The parser has no field that names a node; a node is named only in an assignment. A catalogue test finds no domain name in any definition value. |
| Every requirement has one of the four providers | The parser refuses a requirement whose provider kind is not one of the four. |
| Installation resolves every requirement | A resolution test with one requirement unanswered: refused, naming it and what could answer it. |
| A host requirement is answered on its module's own node | A resolution test: an assignment placing a directory or a port on another node is refused. |
| A module is assigned at most once to a node | A resolution test: assigning a module to a node that already runs it is refused, naming the existing assignment. |
| A public name already taken is refused | A resolution test: a second assignment asking for a public name another holds is refused, naming the holder. |
| An adopted assignment is placed where its data is | An adoption test: the directory resolves to the data's existing location, and nothing is moved. |
## References
- [Issue 119](../04-ISSUES/119-a-module-definition-decides-where-its-files-live/00-report.md): the evidence
- [ADR 0038](0038-the-mesh-assigns-the-port.md), [ADR 0046](0046-a-module-configuration-is-its-assignments-not-its-manifest.md),
[ADR 0084](0084-which-provider-serves-a-consumer.md): the parts already unified
- [ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md): which module provider answers
- [ADR 0113](0113-the-vault-makes-every-secret.md): the vault makes every secret, and how answers travel
- [ADR 0092](0092-an-operator-delivers-a-pair-credential.md): the operator as a provider
- [ADR 0051](0051-shared-data-is-the-operators.md), [ADR 0107](0107-persistent-data-is-a-directory-bind-never-a-named-volume.md),
[ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md): what a directory's contract carries, and what it must not
@@ -0,0 +1,277 @@
---
topic: what runs on it
status: proposed
date: 2026-09-25
deciders: jochen
reconstructed: false
---
# 113. The vault makes every shared secret, a provider makes resources and data, and the mesh carries both
## Context
**A shared secret comes into being many different ways today**, counted across the catalogue and the
controller on 2026-09-25:
| kind | made by | used by |
|---|---|---|
| a credential between a consumer and a provider | the controller | 16 modules, and 3 more for model access, counted below |
| a module's own secret (`own-secrets`) | the controller, as a random value nothing owns | 54 modules |
| a module's broker account | the controller, but only when a person runs a separate command; otherwise the random value above, which cannot work ([issue 095](../04-ISSUES/095-a-module-assigned-after-genesis-has-no-broker-account/00-report.md)) | 49 modules |
| a node's and the builder's broker accounts | the controller, each in its own code path | every node, the builder |
| an enrolment token | the controller | every node joining |
| a `secret` from the vault | the controller mints it, and the vault only records it ([ADR 0085](0085-a-secret-is-a-provision.md), as amended) | 6 modules |
| a value an operator accepts | a person ([ADR 0092](0092-an-operator-delivers-a-pair-credential.md)) | where accepted |
| a licence for model access | a separate controller context with its own store | model consumers |
| the foundation's root secrets | genesis, sealed to the operator key | the foundation |
**The vault was built to end the second row, and did not.** ADR 0085 says a module's own secret
*"stops being a generated value that nothing owns"*. 54 modules still use one, and 6 use the vault.
The replacement was added and the old path was never retired.
**ADR 0085 considered and rejected making the vault the only maker**, because *"the controller must
mint in order to deliver any provision — the vault's own credential among them"*: the vault cannot
make the credentials that exist before it does. That objection is real, and this record has to answer
it rather than step around it.
**Rotation has gaps.** To-be 13 makes rotation one command, all-or-nothing, with a stated window in
which a consumer cannot authenticate. A consumer restarts only if its definition remembered to say so;
a container fed by an env-file was not recreated when that file changed
([issue 103](../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md),
since fixed in the host); and some secrets are read only when a service first initialises, where a restart changes nothing.
**And providers cannot answer with data.** [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)
left *"delivering provider-generated data back to a consumer"* to a separate decision. The analytics
provider's site id and the DNS provider's record have no way back, and say so in their code.
## Considered Options
**1. Keep the controller minting, and tidy the paths.** Rejected. The paths are the problem: each is
made, kept, rotated and audited differently, and tidying keeps them all.
**2. Every provider mints its own secrets, with one shared function in the SDK.** Rejected. Generation
becomes uniform, but custody stays spread over every provider's machine, so rotation, audit and the
operator's break-glass copies cover only some secrets. Each SDK language needs its own implementation.
**3. Raise the vault first at genesis, so it makes even the first secrets.** Rejected. The vault is
built on the shared runtime base, which the installation makes only after the store, the broker and
the controller exist, and the vault learns what to answer from the controller over the bus. Running
it first means reordering the whole installation and giving the vault a second way of being asked.
**4. The vault makes every shared secret; genesis delivers the first ones to it.** Chosen. It answers
0085's objection with a mechanism the mesh already has: a value delivered to the vault.
## Decision
**There are two kinds of secret, and each has one rule.**
- **A shared secret** is a value more than one party must hold: a password, a token, an API key. **The
vault makes every one.** Nothing else in the mesh generates a shared secret.
- **A private key** is made where it is used and never leaves: a node's sealing key, the operator's
key, the mesh's certificate authority. This is not a second way of making secrets. A private key any
other party ever held would no longer be private.
**Every shared secret is a `secret` requirement, answered by the vault:**
- a **credential between a consumer and a provider**. A provision's contract declares *for each
consumer, one secret*, and resolution expands it into one requirement per consumer. So gitea
requiring a database makes the database's provider require a secret for gitea, and the vault
answers it. The provider's own code does not change: it is handed a login and a password, as today;
- a module's **own secret**. `own-secrets` is retired;
- every **broker account** on the mesh's bus: a module's, a node's host's, the builder's, the
controller's. The broker holding `mesh-broker` carries the mesh's bus
([ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md)), and its own provisioner creates
each account from the vault's secret, like any provider. The controller no longer creates accounts,
and there is no separate command to forget;
- an **enrolment token**. The vault makes it; the operator receives the token, sealed to the operator
key, to hand to the joining machine; the controller receives only what it needs to verify it, never
the token itself;
- a **secret operator value**, such as an external API key, which the operator delivers to the vault
([ADR 0092](0092-an-operator-delivers-a-pair-credential.md)). A licence's credential is one of these.
What the licences context adds, refreshing a token, is provider behaviour, decided in its own record;
- a **secret a backend issues itself**, such as an API token a forge hands out exactly once when asked.
The vault cannot make that value. The module that received it delivers it to the vault, which keeps
it and provides it like any other; rotating it means asking the backend again.
**The controller is a module, and takes the same path.** Its store logins (inventory, identity and
licences) and its bus accounts are own secrets of its definition today, and become `secret` requirements
of that definition like any module's. **A node's host is the one party with no definition.** Its bus
account is a requirement the mesh makes for each enrolled node, answered by the vault, sealed to that
node and carried like any other. It is the only requirement not written in a definition, because the
host is what runs definitions.
**Only the vault may provide `secret`.** An assignment providing it must hold the `mesh-vault` seat.
The parser refuses a definition that provides it and cannot hold the seat, and a pin cannot route a
`secret` requirement anywhere else, because there is nowhere else.
**A secret has recipients, and the vault delivers to each.** The database credential has two: the
provider, which *applies* it by creating the login, and the consumer, which *reads* it and presents it
when it connects. The vault hands the value to the mesh sealed to each recipient's node. The controller
and the broker carry sealed values they cannot open.
**Genesis delivers, and the vault adopts.** The vault is built on the shared runtime base, which the
installation makes only after the store, the broker and the controller are running
([to-be 21](../03-DESIGN/01-to-be/21-the-installation-in-full.md)). So the vault is installed **as soon
as that base exists**, before any other module built on it, and everything needed before that moment is
generated by genesis:
- the store's superuser, and the broker's admin in the hashed form the broker needs;
- the bus accounts of the temporary and permanent controller (its account and the broker-management
login), the control-node's host, the builder, the broker's own provisioner and the vault;
- the controller's three store logins (inventory, identity and licences), and the first enrolment
token.
Until the broker's provisioner runs, genesis creates the bus accounts it generated, with the broker's
admin, as the controller does today. Genesis seals each value twice: to the control-node's key, so that when the
vault is installed the controller **delivers the values to the vault, recorded as the mesh's own**, not
as an operator's, with nobody present; and to the operator key, as the break-glass copy
[ADR 0085](0085-a-secret-is-a-provision.md) keeps of every root secret. The first enrolment token reaches
the operator the same way.
That distinction matters: an operator's value is never replaced ([ADR 0092](0092-an-operator-delivers-a-pair-credential.md)),
and these are, because the vault can make their replacements. The broker's provisioner then adopts the
accounts genesis created. From then on the vault makes every shared secret, and genesis has made its
last one.
**Raising the vault or the broker again is a genesis act.** Moving the `mesh-vault` or `mesh-broker`
seat to a new assignment, or recovering either after it is lost, is done the way genesis did it: the
values it needs are delivered, not made by a vault that is not there. They come from the operator-sealed
copies, which the operator opens. The vault keeps a copy of every secret sealed to the operator key
(0085), so nothing the mesh relies on exists only inside the vault. That is a break-glass procedure,
stated and checked, never an ordinary assignment.
**A provider makes resources and data, and the mesh carries data back.** A provider's adapter may
answer with its contract's non-secret fields: a site id, a registered name. The mesh delivers them to
the consumer as resolved values. Who a consumer is stays the mesh's: a provider makes what a consumer
is *given*, never what it is *called* ([ADR 0049](0049-a-consumers-identity-fits-the-tightest-backend.md)).
### Rotation
**Who asks and who makes are decided here; the mechanism is not.** A rotation is asked of the vault,
by an operator or by the vault's policy, such as a maximum age in the requirement's contract, and the
vault makes the new value. A delivered value the vault cannot replace, such as an external API key, is
not rotated by the vault: rotating it means an operator delivering a new one. A secret a backend
issued is rotated by the module that holds the backend asking it again and delivering the new value to
the vault.
**Each recipient takes a new value one of two ways, marked per recipient:**
| recipient takes it by | example | what happens on rotation |
|---|---|---|
| **applying** it | a provider setting a login's password; the broker's provisioner updating an account; a store's provisioner changing its own superuser | its provisioner applies the new value; it is never restarted for it |
| **reading it at start** | a consumer reading its password when it starts | the host recreates it, because a file it read at creation changed |
The marking is per recipient, not per secret, because one secret has recipients of both kinds. A
provision's contract marks its provider's side, which applies. A consumer's side is read at start
unless its requirement says otherwise. The broker's contract marks the host's bus account the same
way: the broker's provisioner applies it, and the host reads it.
Every module in the catalogue reads its secrets at start, and none watches them
([research 016](../01-RESEARCH/016-how-a-credential-can-be-rotated/02-the-readers.md)). A secret a
backend takes only when it first initialises is marked applied, and its provider's provisioner makes
the change using the old value. Where no provisioner can make it, the requirement is marked **not
rotatable by the mesh**, and a rotation request is refused, saying why, rather than restarting a service
that would carry on with the old value.
**How old and new change over is decided in [ADR
0114](0114-a-shared-credential-rotates-over-two-credentials.md).** Three mechanisms were measured against
every provider's code in [research
016](../01-RESEARCH/016-how-a-credential-can-be-rotated/00-overview.md): in place, as the controller's
`rotate` does today; two secrets on one login; and two logins over one resource. Its findings bound the
choice:
- every credential provider already re-applies a password in place, so today's rotation works, with a
window in which a consumer cannot authenticate;
- eight of nine name the consumer's resource after its login, and five destroy the consumer's data when
they remove the login. **No mechanism may retire a login through today's remove**, because in those
five it deletes the consumer's data;
- one backend holds two passwords on one login, and two more hold several tokens;
- every provider can hold two credentials over one resource, eight as two logins and one as two tokens,
once the adapter separates the resource from the credential.
On those facts, 0114 rotates a credential two parties hold over two credentials, rotates one a single
party holds in place, staged, and separates retiring a credential from removing a consumer.
## What this changes in earlier records
On acceptance, each of these is superseded or amended by this record, not edited:
- [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md) is superseded: the controller
no longer mints a provider's credential; the vault makes it. That a provider is handed its
credential and seals nothing stands.
- [ADR 0085](0085-a-secret-is-a-provision.md) is amended: the vault makes every shared secret, own
secrets are retired, and its rejection of vault-only minting is answered by genesis delivering the
first secrets. Genesis seals its values to the control-node's key as well as to the operator key, so
the controller can deliver them unattended. "The vault stores no plaintext, ever" and the
operator-sealed break-glass copies stand.
- [ADR 0092](0092-an-operator-delivers-a-pair-credential.md) is amended: an operator delivers a secret
to the vault. Genesis's values reach the vault by delivery too, but are recorded as the mesh's own,
so 0092's rule that an operator's value is never replaced does not apply to them.
- [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) is amended: a broker
account is created by the broker's provisioner, not the controller. Its scoping stands.
- [To-be 13](../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md),
[to-be 21](../03-DESIGN/01-to-be/21-the-installation-in-full.md) and
[to-be 24](../03-DESIGN/01-to-be/24-the-secrets-vault.md) are amended: the vault makes a rotated value and each
requirement says whether its recipient applies it or reads it at start, and the changeover is
[ADR 0114](0114-a-shared-credential-rotates-over-two-credentials.md)'s; the vault is installed as soon as the
shared runtime base exists, and genesis delivers its secrets to it; the vault is the only maker.
- [To-be 07](../03-DESIGN/01-to-be/07-the-foundation.md) is amended: genesis seals its values to
the control-node's key as well as the operator key.
- [To-be 12](../03-DESIGN/01-to-be/12-a-module-repository.md), [to-be 16](../03-DESIGN/01-to-be/16-module-coverage.md)
and [to-be 18](../03-DESIGN/01-to-be/18-building-a-module.md) are amended: `own-secrets` is retired from the
manifest they describe.
- The [glossary](../00-META/glossary.md) gains *shared secret*, *recipient*, *applies* and *reads at
start*, and *reserved provision*, once this record is accepted.
- [Issue 103](../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md)
is a prerequisite, and its fix is in the host: a container is recreated when a file it read at
creation changes. The issue is to be recorded as fixed, and derived restarts rest on it.
## Consequences
- **The vault is on the path of every new or rotated shared secret.** Today the controller holds that
place, on the same node. A secret can no longer be made while the vault is down.
- Resolution expands per-consumer requirements from a provision's contract. The contract declares
them, never the provider's code, so what a provider requires stays predictable from the catalogue.
- A data provider's adapter gains a return value. What a credential provider's adapter must change for
rotation is [ADR 0114](0114-a-shared-credential-rotates-over-two-credentials.md)'s.
- The broker's provisioner gains every bus account, and the controller loses five separate places it
generates a secret today.
- 54 modules move from own secrets to vault requirements. Six provider clients export a password
generator nothing uses any more; it is removed, so no module can quietly start minting again.
- The installation changes order: the vault is installed as soon as the shared runtime base exists,
before any other module built on it.
- **What got harder:** a secret some services read only at first start can no longer be "rotated" by
a restart that quietly changes nothing; it is refused instead, or applied by its provisioner. And
moving the vault or the broker is a procedure, not an assignment.
## How it is checked
| Rule | Checked by |
|---|---|
| Only the vault generates a shared secret after genesis | A controller test: no code path generates a shared secret. A catalogue test: no module's code generates one, found by scanning for generation calls. Exempt are the vault itself, and randomness that is not a secret any other party holds, such as a password hash's salt, each named in a declared list. An installer test: genesis generates exactly the list above and delivers it to the vault, recorded as the mesh's own. |
| A private key is made where it is used | A test per key: a node's sealing key never leaves the node, the operator's private key never enters the mesh, and the certificate authority's private key never leaves the controller's identity store. |
| Only the vault provides `secret` | The parser refuses a definition providing `secret` that cannot hold `mesh-vault`, and resolution refuses a pin on a `secret` requirement. |
| The controller and each node's host take the same path | A catalogue test: the controller's definition declares no own secret, only requirements. A controller test: a node's bus account is made by the vault and delivered sealed to that node; an enrolment token reaches the controller only as what verifies it. |
| Genesis's values reach the vault unattended, and the operator keeps a copy | An installer test: each of genesis's values is sealed to the control-node's key and to the operator key; the controller delivers the first to the vault when it is installed, with no operator step; the operator's copy opens only with the operator key. |
| Moving the vault or broker is a procedure | A resolution test: an ordinary assignment moving `mesh-vault` or `mesh-broker` is refused, naming the procedure. |
| A backend-issued secret enters through the vault | A vault test: a value delivered as issued is provided like any other, and rotating it is refused as the vault's act. |
| A secret with no provisioner to apply it is not rotated by restart | A vault test: rotating a secret whose requirement is marked not rotatable by the mesh is refused, naming why. |
| A rotation never destroys a consumer's data | A provider test per credential provider: rotating a consumer's credential leaves its resource and data intact. It fails today for no provider, because rotation is in place; it guards whichever mechanism replaces it. |
| A provider's per-consumer secret comes from the vault | A resolution test: a consumer requiring a database expands to a secret requirement for it, answered by the vault and delivered to both recipients. |
| Values are carried sealed | A controller test: each recipient's copy opens with that recipient's node key and no other; neither the controller nor a message on the broker can open one. |
| Own secrets are retired | A catalogue test: no definition declares an own secret, with a declared list of exceptions that shrinks to empty. |
| Bus accounts come from the broker's provisioner | A resolution test: assigning a module that speaks on the bus yields its account, created by the broker's provisioner with no separate command. |
| An operator's value is never rotated by the vault, and genesis's values are | Vault tests: a rotation request on an operator's external key is refused, naming the operator; the same request on a value genesis delivered makes a replacement. |
| Restarts are derived from how a secret is read | A host test: a secret read at start recreates the container that read it at creation, through an env-file or a direct mount; an applied secret restarts nothing. A catalogue test: a secret that reaches a process, or a file in a mounted directory, has `restart-on` naming it. |
| A provider answers data back | A lab test with a consumer requiring analytics: the provider's site id reaches it as a resolved value. |
## References
- [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md): the decision this supersedes,
and the return path it left open
- [ADR 0085](0085-a-secret-is-a-provision.md): the vault, and the objection this record answers
- [ADR 0092](0092-an-operator-delivers-a-pair-credential.md), [ADR 0049](0049-a-consumers-identity-fits-the-tightest-backend.md),
[ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md): delivered values,
identity, and broker accounts
- [ADR 0112](0112-a-module-definition-names-no-node-mesh-or-path.md), [to-be 27](../03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md):
everything a module needs is a requirement
- [Issue 095](../04-ISSUES/095-a-module-assigned-after-genesis-has-no-broker-account/00-report.md),
[issue 103](../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md): what fails today
@@ -0,0 +1,245 @@
---
topic: what runs on it
status: proposed
date: 2026-09-26
deciders: jochen
reconstructed: false
extends: 0113-the-vault-makes-every-secret.md
---
# 114. A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached
## Context
[ADR 0113](0113-the-vault-makes-every-secret.md) decides who asks for a rotation (an operator, or the
vault's policy) and who makes the new value (the vault). It leaves open how old and new change over.
[Research 016](../01-RESEARCH/016-how-a-credential-can-be-rotated/00-overview.md) read every
credential provider in the catalogue against its code. There are nine:
- **all nine re-apply a password in place**, on the same login, every time they run. The controller's
`rotate` command relies on that, and states the window it leaves: between the provider applying the
new value and the consumer restarting with it, the consumer cannot authenticate;
- **eight of nine name the consumer's resource after its login**: a database, a bucket, a virtual host,
a key prefix, a topic prefix, a mailbox. Only the forge's npm registry keeps them apart, because an
organisation owns the packages;
- **five of nine destroy the consumer's data when they remove its login**: postgres, mssql and mongodb
drop the database, lavinmq drops the virtual host with its queued messages, and mailu deletes the
mailbox with its mail. In those adapters, *retire a login* and *delete the consumer's data* are one
call. minio drops a bucket only if it is empty. The provisioner harness makes it worse: a consumer
whose derived login changed is removed under the old login and created under the new one, in one pass;
- **one backend holds two passwords on one login** (redis), and two hold several tokens beside one
password (the forge and mailu);
- **eight of nine can give two logins the same rights over one resource**. mailu cannot, because a mail
user *is* its mailbox. It can give one user several tokens. In postgres, a second login is not enough
on its own: objects belong to whichever login created them, so the resource must be owned by a role of
its own;
- **an administrative credential has one party and a fixed name.** The provider module both applies it
and reads it. Five backends take it only at first initialisation: postgres, mssql, mongodb, mosquitto
and lavinmq. Their credential file is mounted directly into both the server and the provisioner, so
replacing the file recreates the provisioner holding only the new value, which the backend does not
know yet. The provisioner is then locked out;
- **no module watches a secret.** Every reader reads at start, and the host recreates a container when
a file it read at creation changes
([issue 103](../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md)).
A first draft of 0113 chose to overlap old and new "through the adapter's existing create and
remove". In five providers, that remove deletes the consumer's data. The mechanism has to be chosen on
what the providers do, and the danger has to be closed whichever mechanism is chosen.
## Considered Options
**1. In place for everything, as today.** Works with every provider unchanged. Rejected for credentials
two parties hold. The window cannot be closed, only shortened, and the two ends are on different
machines with nothing ordering them. For an administrative credential, it locks the provisioner out.
**2. Two secrets on one login.** Rejected as the mechanism. It works for three providers out of nine,
and using it there and something else elsewhere would put the difference in the mesh instead of in the
adapter.
**3. Two logins over one resource.** Rejected as the mechanism. It works for eight of nine, and not for
mailu.
**4. Two credentials over one resource, with the adapter choosing what a credential is.** A credential
is what a consumer presents, a login and a secret. The mesh alternates between two of them. Each adapter
makes the second one the way its backend can: a second login for eight providers, a second token on the
same login for mailu. A credential a single party holds is staged in place instead, and retiring a
credential is separated from removing a consumer before either is used. Chosen.
## Decision
### Retiring a credential never removes what it reached
**A provider's adapter keeps two things apart that today are one:** the consumer's *resource* (its
database, bucket, virtual host, key or topic prefix, mailbox) and a *credential* that reaches it.
They get separate operations:
- **ensure the resource**, named after the consumer;
- **ensure a credential** with a value, holding the consumer's rights over its resource;
- **retire a credential**. Anything it owns moves first to the resource's owner, and any session it has
open is ended. Then the credential is removed, and nothing else;
- **remove the consumer**, which is what removes the resource, and retires every credential it has.
**Remove the consumer runs only when the consumer no longer requires the provision from this provider.**
That happens when its assignment goes, when its definition drops the requirement, or when re-resolution
sends it to another provider. It never runs because a login or a value changed. The harness keys what it
applied by the consumer, not by the login, so a changed login is a credential change and never a removal.
What removing a resource does with the data in it stays
[ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md)'s, and re-resolving to another provider
moves no data.
**The resource is named after the consumer, and owned by the resource, not by a login.** A consumer's
identity is derived from its assignment ([ADR 0049](0049-a-consumers-identity-fits-the-tightest-backend.md)),
and today its login is that same string, so **no existing resource is renamed**. Where a backend makes
whatever a login creates the login's own, as postgres does, the resource is owned by a role that cannot
log in, and each credential works as that role. Ownership of an existing resource moves to it once. A
credential is retired by handing what it owns to that role, never by dropping what it owns.
### A credential two parties hold rotates over two credentials
**Two parties** means an applier and a reader that are different modules, or a module and a node's
host. The vault's custody copy does not count, because the vault holds every secret. So this covers a
credential between a consumer and a provider, and every bus account: a module's or a host's, applied by
the broker's provisioner and read by its owner. **Each consumer has two credentials, one in use at a
time**, both holding the same rights over the one resource. For eight providers the second is a second
login, derived by the mesh as the consumer's identity with a short fixed suffix. For mailu it is a
second token on the same login.
**The vault drives each rotation and records every step durably.** A provisioner learns which
credentials to hold from what it receives: both of them, for as long as a rotation is under way. It
never learns them from its own memory, so a provisioner restarted mid-rotation resumes from the step the
vault has recorded.
1. **The vault makes the new value.**
2. **Each applier ensures the unused credential with it**, with the consumer's rights, and leaves the
one in use untouched. It verifies that the new credential authenticates and the old one still does,
and confirms. It repeats the confirmation on every reconcile pass until the vault acknowledges it, so
a lost message costs one pass.
3. **Only then does the vault release the new credential to the readers.** The mesh delivers it and the
value together, and the host recreates each reader, because a file it read at creation changed. A
node's host is its own reader: it reconnects to the bus with the new login, and confirms over it.
4. **Each reader confirms by authenticating with the new credential.** It shows this through its
health check, where its definition declares one, or the applier sees the new credential in use,
where its backend reports that. A reader for which neither is possible is confirmed by an operator.
It is never assumed from the reader having restarted.
5. **Only when every reader has confirmed is the old credential retired**, as above, and verified to no
longer authenticate.
**A reader that goes away leaves the rotation.** A reader unassigned, or re-resolved to another
provider, is no longer waited for. A consumer removed mid-rotation has both of its credentials retired
with it.
**A rotation can be abandoned until the old credential is retired.** An operator abandons it. Readers
that moved are given the old credential back, and recreated. The new credential is retired. Nothing is
lost, because the old one was never removed.
`status` shows a rotation as waiting on whichever applier or reader has not moved, and it is not done
until the old credential is gone. A reader that cannot be reached keeps working on the old credential
until it can, and the rotation waits for it. That wait is shown, never hidden.
**Queues and permissions belong to the consumer, not to a login.** A module's queue on the bus is named
for the module on its node, and both of its logins get the same permissions over it
([ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md)). An MQTT client
identifier is chosen by the consumer and is independent of its login. A reader recreated with a new
login keeps it, and the broker hands the session over.
### A credential a single party holds rotates in place, staged
This covers a provider's administrative credential and a module's own secret, which only that module
reads. The vault makes the new value, and the one party takes it:
- **applied**: the vault delivers the new value **staged, beside the current one**, and the current file
is left as it is. The party's provisioner changes the backend using the current value, verifies the
new one, and confirms. Only then does the vault make the new value current. This is the only form for
a backend that takes its administrative credential only at first initialisation. Replacing the file
first would lock the provisioner out;
- **read at start**: the vault delivers the new value as current, and the host recreates the party.
There is no window between two parties, because there is only one. Where neither form can change the
value, the requirement is marked not rotatable by the mesh, and a rotation is refused, saying why
([ADR 0113](0113-the-vault-makes-every-secret.md)).
### One rule decides which
**The number of parties decides, never the provider.** The resolver knows it from the requirement's
recipients, leaving out the vault's custody copy, so no definition declares it.
### Until an adapter can
**An adapter that cannot yet ensure a second credential says so.** The two-party credentials it applies
rotate in place, as today, and the window is stated when the rotation is asked for. So does a
two-party credential whose backend has one fixed name and no second credential for it. These are listed
by a check, and the list is meant to shrink. Separating *retire a credential* from *remove the
consumer*, and keying the harness by consumer, come first. They close a data-loss path that exists
today, whatever rotation does.
## What this changes in earlier records
On acceptance, each of these is amended by this record, not edited:
- [ADR 0113](0113-the-vault-makes-every-secret.md): the changeover it left open is decided here.
- [ADR 0049](0049-a-consumers-identity-fits-the-tightest-backend.md): a consumer's identity leaves room
for the second login's suffix within the tightest backend it reaches, and both logins are checked
against it.
- [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md): a module's broker
account is two logins with the same permissions over the same queue, one in use at a time. Its scoping
is unchanged.
- [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md), already superseded by 0113:
a provider now ensures and retires credentials over a resource it owns separately.
- [To-be 13](../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md): rotation of a two-party
credential is no longer all-or-nothing with a window. It overlaps, with each step confirmed.
A single-party credential is staged, not replaced.
## Consequences
- **Every credential provider's adapter changes**, in two steps. The first separates *retire a
credential* from *remove the consumer*. It names and owns the resource after the consumer, which
keeps the name it has but moves ownership once in postgres and mssql, and the harness is keyed by
consumer. The second ensures a second credential with the same rights.
- **The vault gains rotation state**: each rotation's step, per applier and reader, recorded durably.
Staged delivery is added for single-party secrets. The SDK harness carries the alternation and the
repeated confirmation, so no adapter implements them.
- **No consumer module changes.** It reads one credential at start, as today, and is recreated by the
host when it changes. The exception is a reader that has neither a health check nor a backend that
reports use: its rotations wait for an operator until it declares one.
- **The derived identity is two characters tighter** in the tightest backend, a minio access key of 20
characters.
- **What got harder:**
- a provider briefly holds two credentials per consumer;
- a rotation lasts until its slowest reader moves, so an unreachable reader keeps the old credential
valid until it is reached;
- an adapter has four operations where it had two;
- retiring a login in mssql has to end its sessions first.
## How it is checked
| Rule | Checked by |
|---|---|
| Retiring a credential never removes a resource | A provider test per credential provider: retiring one of a consumer's credentials leaves its resource and data intact, reachable through the other. |
| What a retired login owned survives it | A postgres and an mssql test: objects created under login A, tables included, are still there and alterable under login B after A is retired. |
| A changed login is not a removal | A harness test: changing a consumer's derived login ensures a credential and never calls remove. |
| Remove runs only when the requirement goes | Harness tests: unassigning, dropping the requirement and re-resolving each remove the consumer once; a rotation and a login change never do. |
| No existing resource is renamed | A provider test: a consumer created before the change keeps its resource, with ownership moved to the resource's own role where the backend needs one. |
| Both credentials hold the same rights | A provider test per credential provider: data and structure created under one credential are read, changed and altered under the other. |
| Readers move only after the applier confirms | A rotation test: readers receive nothing until both credentials authenticate at every applier. |
| A reader confirms by authenticating | A rotation test: a reader recreated but failing to authenticate with the new credential does not confirm, and the old credential is not retired. |
| The old credential is retired only after every reader confirms | A rotation test with one reader's node unreachable: it keeps authenticating with the old credential, the rotation shows waiting on it, and it completes when the reader returns and confirms. |
| A reader that goes away leaves the rotation | A rotation test: unassigning a waiting reader lets the rotation complete; removing the consumer mid-rotation retires both credentials. |
| A rotation can be abandoned | A rotation test: abandoning after readers moved gives them the old credential back and retires the new one. |
| Rotation state survives a restart | A test restarting the applier's provisioner, and then the vault, between steps: the rotation resumes from the recorded step. |
| A single-party applied secret is staged | A rotation test on a first-initialisation administrative credential: the provisioner receives the new value beside the current one, applies it, and only then does the new value become current. At no point does it lose its connection. |
| The number of parties decides | A resolution test: a secret with an applier and a reader in different parties is marked for two credentials, and one held by one module for in place. The vault's copy is not counted. |
| Both logins fit the tightest backend | A controller test: both derived logins for the longest node and module names fit the limit ADR 0049 sets. |
| A host rotates its bus login | A rotation test on a node's bus account: the host reconnects with the new login and confirms over the bus before the old one is retired. |
| What still rotates in place is listed | A catalogue test lists every adapter that cannot yet ensure a second credential, and every two-party credential with one fixed name. A rotation of these states its window. |
## References
- [Research 016](../01-RESEARCH/016-how-a-credential-can-be-rotated/00-overview.md): the survey this
rests on, provider by provider
- [ADR 0113](0113-the-vault-makes-every-secret.md): who asks and who makes
- [ADR 0049](0049-a-consumers-identity-fits-the-tightest-backend.md), [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md),
[ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md): identity, bus accounts, and data outliving
its declaration
- [To-be 13](../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md): rotation as implemented
- [Issue 103](../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md):
why a reader's restart can be derived
@@ -0,0 +1,49 @@
---
topic: what runs on it
status: proposed
date: 2026-09-26
deciders: jochen
extends: 0112-a-module-definition-names-no-node-mesh-or-path.md
---
# 115. One assignment of a module per node: the module's name is the assignment's identity
## Context
Everything an assignment owns is named after its module: the database user
(`mesh_<node>_<module>`), the broker account (`<node>-<module>`), the containers, the sealed
secrets, and — since ADR 0112 — the placed directory (`<root>/<module>`). A second assignment
of the same module on the same node would collide on every one of those names at once, which
is why the mesh has never allowed it.
[Issue 119](../04-ISSUES/119-a-module-definition-decides-where-its-files-live/00-report.md)
recorded this as a kept limitation, and 0112 deliberately did not fix it — giving assignments
identities of their own would have touched every naming recipe in one already-large change.
The question stayed open: is multi-assignment a requirement deferred, or a requirement at all?
The original motivation was real — one module serving two tenants on one machine, a second
photo site, a second mail domain. The operator held that requirement once and has now weighed
it against what it costs.
## Decision
**Dropped.** One assignment of a module per node is the rule, not a limitation. The module's
name IS the assignment's identity on a node, permanently, and every naming recipe may rely on
it.
Wanting the same software twice on one node has a spelling the mesh already supports: **two
modules.** A module definition is cheap — two photo sites are two modules sharing artifacts
(the build's images are content-addressed; nothing is built twice), each with its own name,
its own directory, its own grants and its own routes. The tenant boundary lands where every
other boundary already is: the module name.
## Consequences
- The naming recipes stay as simple as they are. No instance suffixes, no assignment ids
threaded through six systems, no migration of every existing name.
- `<root>/<module>` is the assignment's directory with nothing left open (0112's placement
language stands unchanged).
- The controller may refuse a second assignment *plainly* — "novox already runs mailu, and one
node runs one of each (ADR 0115)" — instead of failing on whichever name collides first.
- Multi-tenant asks are answered in the catalogue (a second module definition), not in the
control plane.
@@ -0,0 +1,175 @@
---
topic: the mesh
status: accepted
date: 2026-09-26
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0106-the-bus-is-nats.md
---
# 116. The bus is built in five steps, and the protocol moves with it
## Context
[ADR 0106](0106-the-bus-is-nats.md) decided the bus is NATS and described the change as one thing:
built beside the migration, cut over in one rollout after its core. The architecture was written as
[design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md) and revised once after review. What neither
says is how the work is divided, and three gaps follow from that.
**The whole of the build is one point.** Design 25 §9 numbers four items. The first — "the `nats`
module, the controller's and host's link on NATS, the runtime's client — built and proven in the
lab" — is every line of code the change requires; the other three are the rollout. A step of that
size ends at nothing provable until it ends at everything, which is the failure
[design 22](../03-DESIGN/01-to-be/22-the-work-ahead.md) opens by naming: *a phase that ends at a
claim is a phase that went missing without anything complaining.*
**There is no adoption path.** Design 25 §5 says the broker "is raised at genesis like the store,
adopted as a module in the same phase" — which describes a mesh being raised from nothing. The mesh
this is for is already running, and a running node gets a foundation module by adoption in place
([ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md)), not by genesis. §9 goes
straight from "built beside" to "one rollout" and never crosses that gap.
**The wire changes and the protocol specification does not know it.** Design 25 §8 says a module
sees nothing new. That is true of the SDK's contract — `request`, `handle`, `publish`, `subscribe`
— and false of the wire underneath it.
[ADR 0074](0074-the-wire-is-specified-not-the-types.md) settled that what an SDK implements is a
*specified wire*, checked by fixtures that must match byte for byte, precisely because two
implementations that disagree about an envelope do not fail to compile. That specification is
[design 19](../03-DESIGN/01-to-be/19-the-module-protocol.md), and it is written entirely in AMQP:
exchanges, a durable per-consumer queue named `<node>.<module>.events`, a shared `serve.<key>`
queue — sixteen occurrences of an AMQP term across the document. Design 25 does not cite design 19
anywhere, and design 19 does not cite ADR 0106. So the record that says *what two implementations
may not disagree about* still describes the bus being replaced.
## Considered Options
1. **Keep ADR 0106's shape — build it all, cut over once.** Rejected: not for its rollout, which
is right, but because it leaves the build a single step of unknown length with no intermediate
anyone can run. The three gaps above were found by dividing it; they were invisible while it
was one item.
2. **Cut over incrementally by traffic kind** — events to NATS first, then tools, then control,
the mesh on two buses meanwhile. Rejected: ADR 0106 already rejected two buses, and this is
that with extra steps. The store-window guarantee
([ADR 0083](0083-one-push-leaves-the-mesh-consistent.md)) is exactly the one that cannot cross
a seam, and control is exactly the traffic that carries it.
3. **Five steps, each ending at something provable, the cutover still one rollout.** The build is
divided; the bus still moves once. Adopted.
## Decision
**The bus is built in five steps. Each ends at something a lab bed proves, and no step's proof
waits for the one after it. The cutover remains a single rollout** — dividing the build does not
divide the bus.
**Step 1 — genesis raises the broker.** The `nats` module, and genesis placing it in the
foundation. This is built and proven even though the mesh it is for will never travel this path,
because
genesis is where the foundation is *defined*: the only place the mesh comes from nothing, and the
definition every other path is measured against. A genesis path that exists only on paper is one
nobody discovers is wrong until there is a second mesh.
**Step 2 — adoption puts the broker in the seat.** A mesh already running receives the broker by
adoption in place, and the seat it claims is **`mesh-broker`** — unchanged.
[ADR 0079](0079-the-foundation-seats-are-named-after-their-servers.md) named the foundation seats
after the server's *role* rather than the product for exactly this case, and ADR 0106 restated it:
the broker module changes, the seat does not. A seat named after the product would have to be
renamed by every change the seat exists to survive.
**Step 3 — the protocol gets a NATS binding.** ADR 0074 stands unamended: the mesh defines a module
protocol, an SDK is an implementation of it in one language and nothing more, the protocol is split
per capability, and conformance is executable fixtures per capability rather than prose. What
changes is what the specification specifies. Design 19's wire section is rewritten from exchanges
and queues to subjects and streams; the conformance suite is built — first against the bus the
mesh has, then restated on NATS; every SDK claims the capabilities it passes, and a language may
arrive with connection and events alone.
**The SDK gains no conveniences in the process.** [ADR 0039](0039-what-the-sdk-holds-and-refuses.md)
refuses frequent-and-cascading code in the shared library, and ADR 0074 restates it — an SDK is
"not a convenience layer, not a place for helpers to accumulate." A new transport is the moment
that pressure is highest and the reason to hold hardest: the predecessor's shared library is the
cautionary tale ADR 0039 opens with, and it did not become that in one decision. Code shared among
a module's own features stays in that module.
**Step 4 — the core speaks NATS.** The controller's link, the host's link and the tool runtime's
client, and with them the flows that are today carried by something other than the bus: a build
source's change reaching the builder, an installation, a module's own reports. Each is a
conversion with a named before and after, not a rewrite. Observation — heartbeats, conditions,
key-value state — is [research 017](../01-RESEARCH/017-a-mesh-that-heals-itself/00-overview.md)'s
and stays there; that effort already reserves it for after the move, and this step does not
pre-empt its design.
**Step 5 — deployment.** The rollout ADR 0106 decided, unchanged: the controller, every host and
every runtime move together, each node confirmed to have heard before AMQP stops. What this step
adds is that steps 1 to 4 *ship ahead of it* without moving any node's bus — the module exists in
the catalogue, the protocol is specified, the code is written and beds pass, and the running mesh
is still on AMQP throughout. The bus moves on one day, at the end, once.
## Consequences
- **Design 25 §9 is replaced by these five steps** and §10's beds are attributed to the step each
proves, so no proof waits for the last step.
- **Design 19 is stale in its wire section from today** and says so in its own frontmatter and
opening until step 3 rewrites it. A specification that describes the bus being replaced is worse
than an absent one, because it reads as current.
- **`nats-broker` is not a seat.** The module is `nats`; the seat is `mesh-broker`.
- **Steps 1 to 4 leave every node on AMQP.** A step can be abandoned, or reordered after step 2,
without a rollback — the cost of being wrong is bounded until step 5.
- **What got harder:** five steps mean five proofs rather than one, and step 3 rewrites a design
other designs cite, so their references are checked when it lands. Dividing the work does not
reduce it.
## How it is checked
- **Per step, a bed, and the bed named in design 25 §10 against the step it belongs to.** Step 1:
a mesh raised from nothing has the server standing, its streams asserted and every account
composed from the manifests — no mesh traffic on it yet. Step 2: the broker adopted into a mesh
already running, a second holder of `mesh-broker` refused at resolution, and nothing routed to it.
Step 3: the conformance suite passing per capability, on NATS, for every implementation that
claims it — and a module built before the binding serving its tools unchanged. Step 4: each
converted flow proved against the behaviour it replaced, **and the full genesis bed** — a mesh
enrolling, holding a push, and rolling out an upgrade on NATS. Step 5: the cutover bed, then the
rollout with every node reporting.
- **Step 3 is not done when the code runs.** It is done when the fixtures match byte for byte
across implementations, which is ADR 0074's own test and the only one that catches two SDKs
quietly ignoring each other.
- **A step that cannot name what its bed proves is not a step**, and is divided further before it
is started.
## Progressive insights
Corrections of fact made in this record after it was accepted, under the rule in
[`README.md`](README.md). The decision — five steps, each proved, one rollout — is untouched by
both; each corrects something this record asserted about the *state of the code*, which nobody had
measured when it was written.
> **Progressive insight — 2026-09-26.** *There is no conformance suite to recapture.* This record
> said step 3's "fixtures are recaptured on NATS", and design 19's wire was described as specified
> and conformed. Measuring the four repositories found no conformance fixtures in any of them:
> [design 22](../03-DESIGN/01-to-be/22-the-work-ahead.md)'s Phase 1.2, which would have built the
> suite, is still open. Step 3 therefore **builds** it, and builds it first against the bus the
> mesh has — a suite born on the new bus certifies whatever the new bus happens to do. The step's
> place in the order, and ADR 0074's model it implements, are unchanged.
> **Progressive insight — 2026-09-26.** *The full genesis bed belongs to step 4, not step 1.* This
> record attributed "a mesh raised on NATS from genesis" to step 1, and described that step as "the
> `nats` module and a mesh raised on it from nothing". A bed that enrols a node, holds a push and
> rolls out an upgrade needs the controller and the host to speak NATS — which is step 3's
> implementations and step 4's flows. A step whose proof cannot run is exactly the failure this
> record was written to prevent, so step 1 now ends at the server standing from genesis, correctly
> configured and carrying nothing, and the full bed is named under step 4. The five steps, their
> names, their order and the single rollout are unchanged; only where two beds run has moved.
>
> Recorded here rather than superseded because the decision this record makes — that the work is
> divided and each division is proved — is what *produced* the correction: the dependency was
> invisible while the build was one item.
## References
- [ADR 0106](0106-the-bus-is-nats.md) — the decision this divides.
- [ADR 0074](0074-the-wire-is-specified-not-the-types.md) — the protocol and its conformance suite;
step 3 is its NATS binding, not a replacement.
- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — why step 3 adds no helpers.
- [ADR 0079](0079-the-foundation-seats-are-named-after-their-servers.md) — why the seat is
`mesh-broker`.
- [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md) — the adoption step 2 uses.
- [design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md), [design 19](../03-DESIGN/01-to-be/19-the-module-protocol.md) — the two documents this changes.
@@ -0,0 +1,135 @@
---
topic: what runs on it
status: accepted
date: 2026-09-26
deciders: jochen
reconstructed: false
extends: 0110-a-seat-is-a-module-assignment-from-a-closed-set.md
---
# 117. A machine's uplink is a seat: the mesh configures the manager, never the link
## Context
The mesh installs on top of a machine's own networking. The private network's generator says
so in as many words: a machine has an address and a route to the broker *before* the mesh
exists, the broker's address travels in the enrolment token rather than being resolved, and the
private network is something the mesh installs on top, like anything else. Nothing in the mesh
says who manages that uplink, or what the mesh needs from whoever does.
Adopting the first workstations showed that the mesh does need something from it, and gets it
by accident:
- **The resolver the mesh owns depends on a file the mesh does not.** `resolv-conf` writes
`/etc/resolv.conf` and names the mesh's resolver. On a machine running NetworkManager, the
manager rewrites that file on every connectivity change unless it is told `dns=none`; on one
running dhcpcd, every lease renewal rewrites it unless it is told `nohook resolv.conf`. On
the adopted machines both settings exist only because the predecessor wrote them. No module
declares them. Remove the predecessor's file and the mesh's resolver is silently replaced the
next time a laptop changes network, while every surface of the mesh still reads green.
- **`resolv-conf` cannot declare them itself.** Which setting is needed depends on which
manager runs, and a `service` resource for a manager that is not installed fails the
declaration. A resolver module that knew about network managers would be the wrong module
knowing the wrong thing.
- **The private network's interface is exposed to the manager.** A manager that considers
every interface its own may try to configure `mesh0`, or tear it down on a profile change.
Nothing tells it not to.
- **Two managers on one machine go unnoticed.** Among the machines adopted so far, one runs
NetworkManager *and* dhcpcd at once: two programs that each believe they own the machine's
addresses and its resolver file.
Nothing detected it, because nothing in the mesh knows the role exists.
The machines differ in a way that matters: servers are wired and never move, while
workstations join wireless networks, captive portals and phone hotspots wherever they are.
## Considered Options
**1. The mesh manages the uplink: links, addressing, wireless networks and their
credentials.** Rejected. The mesh reaches a machine only over that link. A declaration that
gets it wrong — a mistyped network, a stale credential, a manager that fails to start — takes
the machine off the network, and with it the only channel a fix could arrive on. That is the
one failure the sshd module's `listens` rule forbids the firewall to arrange; a mesh that owned
the link could arrange it with any push. And a wireless network is joined at the machine, by
the person using it, in the moment. A declaration composed elsewhere cannot answer a captive
portal.
**2. Leave the uplink unmanaged; accept the implicit dependency.** Rejected. It keeps the
resolver working only for as long as a predecessor's file survives, and it leaves two managers
on one machine undetectable.
**3. The uplink is a seat. The module holding it configures the manager's relationship to the
mesh, and never the link.** Chosen.
## Decision
**`the-uplink` is a node-scoped seat** in the closed set ([ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md)).
It delivers no provision. It is held by the module for the program that manages the machine's
own network, one per manager: `networkmanager`, `systemd-networkd`, and `dhcpcd` for a machine
with nothing more. Assigning a second is refused, naming the first.
**What a holder declares** — only what keeps the manager and the mesh from contradicting each
other:
- the manager's package, present — and its service **with no state**: the manager's lifecycle is
the machine's. The mesh never starts, stops, enables or disables it, because stopping it takes
the link down, and a holder unassigned by mistake — or the wrong holder assigned — must not be
able to do that, nor start a second manager beside the one the machine runs. The service is
declared only so a change to the holder's settings reaches a *running* manager;
- the manager's own configuration that leaves the resolver file to the mesh (`dns=none` for
NetworkManager, `nohook resolv.conf` for dhcpcd, and nothing for systemd-networkd, which
never writes the resolver file);
- the manager's own configuration that leaves the private network's interface alone
(NetworkManager's `unmanaged-devices` naming `mesh0`; dhcpcd's `denyinterfaces mesh0`; for
systemd-networkd a network file of the module's matching `mesh0` as `Unmanaged=yes`);
- each as a drop-in beside the manager's main file where the manager reads one, and written
*into* a shared file otherwise, as a marked region the host owns
([ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md)'s idea for text files),
placed where the manager reads it as global — at the start of `dhcpcd.conf`, above any
`interface` line, because every line after one belongs to that interface;
- the service **reloaded** when a drop-in changes, never restarted — a restart drops the link,
and the link is the mesh's own channel to the machine. **A manager that cannot reload is not
restarted instead:** its setting takes effect at the manager's next start. Measured on the
adopted machines: NetworkManager (1.58) and systemd-networkd (systemd 261) both report
`CanReload=yes`; dhcpcd (10.3) reports `CanReload=no`, so its module declares no trigger at
all. Whether each setting is actually *applied* by a reload is confirmed on a machine before
the module is taken there, not assumed.
**What a holder never declares:** a link, an address, a route, a connection profile, a
wireless network or its credentials. Those are the operator's, in the sense of
[ADR 0051](0051-shared-data-is-the-operators.md): the mesh does not create, change or delete
them, and the module's `access`, if it needs one, is read-only.
## Consequences
- `resolv-conf` stays generic. The condition it could not express — "only if NetworkManager
runs" — is expressed by assigning the module for the manager that does.
- The dependency on the predecessor's `dns=none` file becomes a declared resource. On an
adopted node the holder's drop-in arrives beside the predecessor's; both say the same thing,
and the predecessor's is retired by hand after the take, like any other file the mesh
replaced under another name.
- A setting a manager reads only at its start is not in force until then. On an adopted machine
the predecessor's identical line normally already is; on a machine that was not adopted,
dhcpcd's resolver hook keeps rewriting the resolver file until dhcpcd next starts, and the
operator restarts it once, in a window of their choosing.
- A machine running two managers is found at assignment: the second holder is refused, and the
operator decides which manager the machine keeps before either module is taken.
- Workstations keep joining networks the way they always have. Under NetworkManager and
systemd-networkd the host already cooperates with the manager — its dispatcher hook wakes it
on every connectivity change — and nothing here changes that. A dhcpcd-only machine has no
such hook, and nothing here adds one.
- The seat table gains one entry: `the-uplink`, node scope, delivering nothing, decided here.
- **Not decided here:** whether the mesh should ever *offer* known networks to a machine — a
sealed, add-only list the operator curates once for all workstations. That is a different
question (the mesh holding credentials for links it must never be able to break) and gets its
own record if it is wanted.
## References
- [ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md): the closed set this seat joins;
[to-be 26](../03-DESIGN/01-to-be/26-the-seats.md): the seat table
- [ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md): written into, never over
- [ADR 0051](0051-shared-data-is-the-operators.md): what is the operator's stays the operator's
- mesh-controller `internal/catalogue/seats.go` (the seat), `internal/overlay/generator.go` (the
mesh installs on top of the machine's own networking)
- mesh-catalog `modules/networkmanager`, `modules/systemd-networkd`, `modules/dhcpcd`
- mesh-host `internal/apply/block.go` (a file written into a marked region, `at` start or end)
@@ -0,0 +1,119 @@
---
topic: what runs on it
status: accepted
date: 2026-09-27
deciders: jochen
reconstructed: false
extends: 0102-the-mesh-writes-into-a-shared-file-never-over-it.md
---
# 118. Undeclaring removes what the mesh made, and gives a unit back the state it was found in
## Context
When a resource stops being declared — its module unassigned, the node sent a
deliberately-empty declaration ([issue 127](../04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)),
or a new catalogue version renaming its id — the host undoes it. The host's own code states
the rule it means to follow: **it removes what it made and leaves what it merely configured.**
For almost every resource it does exactly that:
- a container, a network, a process's unit, a directory it created: removed;
- a file it created: removed; a file it replaced: its kept original put back
([ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md));
- keys and list members it wrote into a shared file: given back as they were
([ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md));
- a package: left installed — the host cannot know it is unused;
- an operator's path it was given access to: never touched
([ADR 0051](0051-shared-data-is-the-operators.md)).
**A service is the exception.** A `service` resource never installs a unit: it puts one that
already exists — the distribution's, the operator's — into a state. Undeclared, the host stops
it. That contradicts the rule above, and in practice it is the most dangerous thing an
undeclare can do. Found reviewing the uplink modules
([issue 130](../04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md)):
- the private network declares the container runtime's unit only so a change to the registry
trust reloads it — unassigning the private network stops the runtime, and every container on
the machine, the mesh's and not;
- the sshd module declares the ssh daemon — unassigning it stops ssh, the lockout that module's
own `listens` rule forbids;
- the uplink modules would have stopped the network manager, taking the machine off the only
link the mesh reaches it by.
[ADR 0125](0117-a-machines-uplink-is-a-seat.md) answered that for its own modules with a
service declared with no `state`. Every other module that declares a unit it did not make is
exposed in the same way, and relying on each author to remember an opt-out is how the next one
is missed.
## Considered Options
**1. Undeclaring touches nothing on the machine.** Rejected. What the mesh made would outlive
the module that made it: a container nobody manages keeps serving and stops being patched; a
unit the mesh wrote keeps running a bundle nothing updates; a name collides when the module
is assigned again. An undeclare that leaves the mesh's own work behind is an orphan factory.
**2. Keep stopping services; make "leave it running" an opt-in per resource.** Rejected. It
keeps the dangerous behaviour as the default for exactly the units that matter most — the
runtime, the ssh daemon, the network — and each new module is one forgotten field away from a
machine that goes dark when it is unassigned.
**3. Never stop a unit the mesh did not create.** Rejected, found while implementing it. The
mesh's packet filter is a unit the distribution installed and the mesh started at converge;
returning a node to adopted ([ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md))
unloads it by undeclaring it. Never stopping it would leave the mesh's filter loaded beside the
predecessor's firewall re-enabled — the one rollback a converge promises, broken. Who wrote the
unit file is not the line; what the mesh *did* to the unit is.
**4. Give the unit back the state it was found in.** Chosen.
## Decision
**Undeclaring removes what the mesh made and gives back what it changed.** For a unit the mesh
did not create, what it changed is the unit's state, so that is what is given back: **the host
records the state it first found the unit in, and undeclaring returns the unit to it.**
- **Recorded once**, the first time the host applies the service — whether it was running, and,
where the declaration sets it, whether it was enabled at boot — and carried in the host's
record from then on. Later applies never overwrite it: by then the unit's state is the mesh's
doing.
- **A unit found running is left running.** The container runtime, the ssh daemon, a network
manager: running before the mesh arrived, running after it leaves.
- **A unit the mesh started is stopped again**, and one it enabled is disabled again — the packet
filter a converge loaded, which returning to adopted unloads.
- **Never started on the way out.** A unit the mesh stopped is not started again when its
declaration goes; starting something is a decision, and the operator makes it.
- **Unknown is left alone.** A record written before the host kept what it found says nothing
about the unit before the mesh; the unit is left exactly as it is. A unit left running can be
stopped by the operator; one stopped by mistake may be the link the operator needed to do it.
- A unit the mesh *did* create — a `process` resource's unit and bundle — is stopped and removed
with its declaration. That is the mesh's own code. (Before this record there was no way to
remove one at all: an undeclared process failed every apply on its node.)
- The service's settings the mesh wrote are given back by their own resources (a kept original
restored, a region or keys removed). A running service keeps running on what it read until it
next reads its configuration; the mesh does not restart it to make it notice.
- A service declared with no `state` (ADR 0125) remains the way to say the mesh must not
**start** a unit either; undeclared, it is forgotten.
## Consequences
- Unassigning the private network no longer stops the container runtime; unassigning sshd no
longer stops ssh; no uplink module can take a machine's network down on its way out.
- The host's removal report says what it gave back — "restored: stopped again, as the host
found it" — or "forgotten: it was running before the mesh; left as it is" where it used to say
"stopped". Its plan names each unit an undeclare will stop, before it does.
- On a fresh machine where the mesh installed and started a service, unassigning its module
stops it again — the mesh gave, the mesh takes back. An operator who wants it kept declares it
in a module of their own, or starts it themselves after.
- A daemon can keep running after its module is gone, on configuration that was taken back from
under it. That is a visible, running process the operator can see and stop; the alternative
was an invisible outage.
- **Not decided here:** an unassign preview that lists what an undeclare will remove and what it
will leave running. Issue 130 asks for it; it is the controller's to build.
## References
- [issue 130](../04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md): the finding
- [ADR 0125](0117-a-machines-uplink-is-a-seat.md): the uplink modules, and a service with no state
- [ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md), [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md):
what is given back, and how
- mesh-host `internal/apply/apply.go` (`remove`, the service case)
@@ -0,0 +1,87 @@
---
topic: the mesh
status: accepted
date: 2026-09-27
deciders: jochen
reconstructed: false
extends: 0105-the-mesh-adopts-the-predecessors-tunnel-in-place.md
---
# 119. A taken tunnel's predecessor is retired once the take is proven
## Context
[ADR 0105](0105-the-mesh-adopts-the-predecessors-tunnel-in-place.md) has the private network take
over the tunnel it finds: the found unit stopped and disabled, never flushed, and **its
configuration left on disk, kept like any held file.** That was the right caution for the take
itself — if the mesh's interface failed to come up, the host starts the found unit again and the
peers never notice — and every apply since stops the found unit again should anyone start it.
What it leaves is a predecessor that never finishes leaving. On every machine that has enrolled,
the tunnel is the mesh's and has been proven so — its interface up with the found key, the peers
handshaking, the machines resolving and reaching each other over it — and still the predecessor's
configuration sits where its unit reads it, held for a module that has long since replaced it.
The predecessor itself is being deprecated. A tunnel that can be started again by one command, with
a configuration nothing maintains any more, is not a rollback path; it is a second way onto the
network that nobody is watching. And the hold never ends, so every node report keeps listing it.
## Considered Options
**1. Keep it, as 0105 says.** Rejected: the caution it bought is spent once the take is proven, and
what remains is a live, unmaintained way back onto the network.
**2. Delete it at the take.** Rejected: the take is exactly the moment the fallback is needed. If
the mesh's interface does not come up, the host must still be able to raise the found one.
**3. Retire it once the take is proven.** Chosen.
## Decision
**Once the mesh's interface has proven it carries the tunnel, the found interface's configuration
is removed from where its unit reads it.**
- **Proven means:** the tunnel's state is *taken* — the found unit down and disabled, the mesh's
interface up with the found key — and the mesh's interface has completed a handshake with at
least one peer. Not before: until then, a failed take still falls back to the found unit.
- **Retired means:** the configuration file the found unit reads is removed. Its original was
already kept, before anything happened to it
([ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md)), and stays kept; that copy
is the record of what the predecessor was, and a person's way back if one is ever wanted.
- The found unit stays disabled. Without its configuration it cannot raise the interface, so the
every-apply stop that guarded against it becomes a check that finds nothing to do.
- **The hold ends.** What was held for the private network has been replaced; the node stops
reporting it.
- **The mesh never brings it back.** Undeclaring the private network does not restore the found
tunnel: the mesh stopped it, and nothing is started on the way out
([ADR 0126](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md)). A machine whose
private network is unassigned has no tunnel until it is assigned again — which is what
unassigning it means.
## Consequences
- On every machine that took a tunnel, the predecessor's tunnel configuration disappears at the
first apply after the take is proven. Nothing a peer sees changes; the mesh's interface already
carries the same key, port, address and peers.
- A take that is never proven — no peer ever handshakes — keeps the found configuration, and the
node says so, so a broken take is visible rather than silently retired.
- Rolling back to the predecessor's tunnel becomes a deliberate act, in this order: **unassign the
private network first**, then copy the kept original back and start its unit. The mesh does
neither. While the private network is still assigned, the tunnel is the mesh's: a restored
configuration is held and retired again at the next proven apply, and the found unit cannot
bind the port the mesh's interface holds. The node says so when it happens.
- A configuration something keeps writing back — the predecessor's own tooling, say — is retired
again each time it appears, but the first original stays the one kept; a different content is
kept once beside it, and the node reports that the configuration came back.
- The host retires only the found interface's own configuration file (`/etc/wireguard/<iface>.conf`),
never a path the mesh writes, and never a link: a configuration that is a link to somewhere else
is left, with its target, for a person to retire.
- 0105's "its configuration stays on disk, kept like any held file" holds until the take is proven,
and not after.
## References
- [ADR 0105](0105-the-mesh-adopts-the-predecessors-tunnel-in-place.md): the take, and why it keeps
the found configuration during it
- [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md): kept originals
- [ADR 0126](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md): nothing is started on the way out
- mesh-host `internal/apply/takeover.go`
@@ -0,0 +1,136 @@
---
topic: what runs on it
status: accepted
date: 2026-09-27
deciders: jochen
reconstructed: false
extends: 0112-a-module-definition-names-no-node-mesh-or-path.md
---
# 120. A roster fact carries its format as a template: the mesh owns the data, the module owns the format
## Context
A **fact** is a thing only the mesh knows — which machines exist, what they are called, where they
are — written into a file where a module asks for it. The mesh computes it from the graph; a module
loads it, restarts on it, does what its software does with it. Facts replaced three modules that
existed only because computed output needed somewhere to live and ran no software of their own
([ADR 0040](0040-what-a-module-is.md)).
But the *format* lived in the control plane. A fact was a name from a closed list, and each name had
a formatter written in Go beside the others: `node-names` wrote the roster as an `/etc/hosts` file,
`node-zones` wrote it as a dnsmasq resolver's `local=`/`address=` lines. Adding a consumer meant
adding a formatter — in the consumer's own configuration language — to the mesh.
The ssh work made the cost plain. An operator's `~/.ssh` wants three roster projections — a
`known_hosts`, an ssh `config` of `Host` blocks, an `authorized_keys` — each in ssh's syntax. Under
the closed list that is three more formatters in the control plane, teaching it ssh's configuration
language. And it does not stop at ssh: every daemon that reads the roster in its own file format
would put its grammar here. The control plane was accreting the configuration languages of software
it does not run — the exact thing [ADR 0040](0040-what-a-module-is.md) says is a
module's and not the mesh's.
The shape underneath is one shape. WireGuard's `[Peer]` blocks, `/etc/hosts`, dnsmasq's zones, an
ssh `known_hosts` — all of them are *the roster, projected into a file*. Only the projection differs,
and the projection belongs to whoever runs the software that reads it.
## Considered Options
**1. Keep the closed list; add a formatter per consumer.** Rejected. The control plane learns the
configuration language of every daemon any module might run, without bound, and each format lives in
the mesh rather than in the module that owns the file. A module cannot change how its own file is
written without a control-plane change.
**2. A general placeholder vocabulary over `content`, like `${machine:address}` but for the
roster.** Rejected. The mesh's other substitutions each resolve to *one* scalar — this machine's
address, one provider's port. The roster is inherently a *repetition*: one block per machine. A flat
`${…}` vocabulary cannot iterate, and a mechanism that could would be a template in all but name.
**3. The module gives a path and a template over the roster; the mesh renders it.** Chosen. The mesh
owns the data — who exists, their names and addresses — and hands it to a Go `text/template` the
module wrote. The mesh renders and reads neither the template's intent nor the file's meaning.
## Decision
**A fact is a path and a template.** In a module's manifest, `facts` maps a name the module chooses
to a `{ path, template }`. The template is a Go `text/template` over a fixed **roster view**:
- `.Node` — this machine's bare name.
- `.Suffix` — what a mesh name ends in (`internal`, or the operator's choice), as composed.
- `.Names` — every name the mesh serves: the machines *and* the names it was told to route.
- `.Machines` — only the machines that are nodes of this mesh.
Each of `.Names` and `.Machines` is a list of `{ Name, FQDN, Address }`. A machine the mesh has a
record for but cannot yet place has no address and is left out of both — a name that resolves to
nothing is a connection that hangs, so it is omitted rather than written (the same rule as before).
**The mesh owns the data; the module owns the format.** The control plane holds **no** formatter.
The two built-in projections render through the same path any module uses:
- **`/etc/hosts`** is a template on the mesh's own network module. The mesh writes `/etc/hosts`
because being on the private network is what gives a machine a name — but the *layout* is a
template like any other, shipped with the control plane because that module ships with it, not
because the control plane knows the hosts-file format.
- **dnsmasq's zones** move into dnsmasq. The `local=`/`address=` grammar is dnsmasq's configuration
language, and it now lives in dnsmasq's manifest, where the module that runs dnsmasq owns it.
**A fact also says whether its file is the mesh's whole or a region of the machine's.** A hosts file
is the machine's — its `localhost`, the operator's lines, another tool's marked blocks — so
`node-names` is `shared`: the mesh owns only its region and keeps the rest byte for byte, the host
laying it down `into: block`
([ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md), hq issue 128). A resolver's
zones file is the mesh's whole, and is not shared. The template renders the content either way;
`shared` decides how the host writes it. This composes with hq 128 rather than replacing it: the
region *mechanism* is the host's, the region's *format* is the module's template.
**The names-vs-machines distinction is the template's choice** ([04-ISSUES/111](../04-ISSUES/111-the-resolver-is-told-names-the-mesh-serves-not-only-machines/00-report.md)):
a container's hosts ranges `.Names`, so a routed name resolves to the machine serving it; a resolver
told the mesh's suffix is its own ranges `.Machines`, or a routed name written there with the suffix
appended is a name nobody will ever ask for.
**A template that will not render is refused at composition, not on a machine.** A template that does
not parse, or reads a field the roster does not have, fails where the manifest is — the closed-list
safety, moved from the fact's *name* to the roster's *shape*. A daemon that starts, reads a file the
mesh could not render, and answers nothing is a much worse way to find out.
**WireGuard stays a computed generator, and that is the line.** Its `mesh0.conf` is not a pure roster
projection — it carries topology the control plane decides: which peers are reachable, endpoints, hub
forwarding, keepalive for a NAT'd node. And it is *foundational*: the overlay must be up before any
module can be delivered, so the thing that writes it cannot itself be a delivered module. The line
this draws: **the substrate that delivery rides on is the control plane's; everything layered on a
working overlay is a roster template.** DNS, hosts, and ssh are layered; the overlay is the floor.
## Consequences
- **ssh is two templates and no control-plane change.** Once the roster view carries a machine's ssh
host key and its operator account ([to-be 29](../03-DESIGN/01-to-be/29-a-node-has-operator-accounts.md)),
`known_hosts`, the ssh `config`, and `authorized_keys` are templates on the ssh modules — the mesh
gains no knowledge of ssh's syntax. This ADR is what makes that work land without touching the
controller.
- **A new roster projection never touches the control plane.** Any module that reads the roster in
its own format ships its own template.
- **A module can change how its own file is written** without a control-plane change — it is editing
its own manifest.
- **The schema changed and is not backward compatible.** A fact was a string (a path); it is now
`{ path, template }`. The old string form has no template and cannot be auto-upgraded, because the
format it implied was the formatter this ADR deletes. The controller and every catalogue module
using facts — only dnsmasq — land together. A controller and a catalogue that disagree cannot
compose the module: the running daemon on a machine is unaffected, but the mesh will not send it a
new declaration until both sides agree.
- **The output did not change.** The `/etc/hosts` and dnsmasq zones a machine receives are
byte-for-byte what the deleted formatters wrote, pinned by tests that render the built-in template
and compose the real dnsmasq manifest.
## References
- [ADR 0040](0040-what-a-module-is.md): a module is software the mesh runs — a
format the mesh knows for software it does not run was the accretion this stops
- [ADR 0112](0112-a-module-definition-names-no-node-mesh-or-path.md): a module definition names no
path; this is its sibling for content — a module definition names no format the mesh must know
- [to-be 29](../03-DESIGN/01-to-be/29-a-node-has-operator-accounts.md): the ssh consumer this
unblocks, and the roster fields it will add
- [04-ISSUES/111](../04-ISSUES/111-the-resolver-is-told-names-the-mesh-serves-not-only-machines/00-report.md): every served name is not a
machine — now the template's choice of `.Names` or `.Machines`
- mesh-controller `internal/catalogue/roster.go` (the mechanism), `internal/overlay/generator.go`
(the built-in `/etc/hosts` template), `internal/catalogue/manifest.go` (`RosterFile`)
- mesh-catalog `modules/dnsmasq/module.json` (the zones template, dnsmasq's own)
@@ -0,0 +1,132 @@
---
topic: what runs on it
status: accepted
date: 2026-09-27
deciders: jochen
reconstructed: false
extends: 0110-a-seat-is-a-module-assignment-from-a-closed-set.md
---
# 121. A system seat is named for its scope, and a module may define its own
## Context
[ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md) made seats a closed set the
control plane defines: a well-formed name no longer becomes a seat by being claimed, so a person can
read what a mesh can have and who fills each role. It left two things unsettled that the growing set
now exposes:
- **The names carry no rule.** `mesh-controller`, `mesh-store`, `mesh-broker` are named for the mesh;
beside them sit `the-artifact-store`, `the-build-machine`, `the-dns-port`, `the-showcase`,
`the-uplink` — a second naming style with no principle behind it. A reader cannot tell a seat's
scope from its name, and the mesh's own roles do not look like the mesh's.
- **The set is the *only* place a seat may be defined.** A module claiming any name not in the
control plane's set is refused. That is right for *system* roles — one broker, one packet filter
per node — but it means a module can never define a role of its own: a demo module's
`the-showcase`, a future application's coordination role, must be smuggled into the control plane's
set or not exist. The control plane ends up holding roles that are not the mesh's to define.
Reviewing the set against these also found seats whose *scope* or *membership* is wrong, not just
their name — the review is the occasion to fix those too.
## Decision
**A system seat — one the control plane defines — is named for its scope:**
- **`mesh-*`** for a mesh-scoped seat: one holder in the whole mesh, a role the mesh has once
(`mesh-controller`, `mesh-store`, `mesh-broker`, `mesh-git`, …). A `mesh-*` seat is always held by
a module **on a named node** — `mesh-git` is gitea *on novox*, not "gitea"; another node running
gitea does not hold `mesh-git` unless it is the holder. The seat is the mesh's single answer for
the role, and which node answers is part of what the seat records.
- **`node-*`** for a node-scoped seat: one holder per node, a role each machine has at most once
(`node-packet-filter`, `node-intrusion-prevention`, `node-uplink`, …).
The three already-`mesh-*` seats keep their names; the rest are renamed by this rule. The scope a
name declares must match the seat's actual scope — a `mesh-*` seat at node scope, or the reverse, is
a contradiction the reader is entitled to trust is impossible.
**The control plane defines only system seats. A module may define its own.** A seat named `mesh-*`
or `node-*` is the control plane's, and claiming one the control plane does not define is refused as
before. Any *other* name is a **module-defined seat**: valid when the module declaring the claim also
declares the seat (its name, scope, and — if any — the protocol its holder speaks). The control plane
enforces one-holder-per-scope for it exactly as for its own, but does not otherwise know what it
means. So an application can coordinate its own instances through a seat of its own, and the mesh's
closed set stays what its name says it is: the *system's* roles, not everyone's.
**Specific seats this settles:**
- **`the-build-machine` → `mesh-build-machine`, and its scope becomes mesh.** There is one build
machine in the mesh (the builder on novox), not one per node. Node scope said the opposite. It
delivers no provision; it is the mesh's single build machine.
- **`the-private-network` → `mesh-private-network`, held by the network *server* on one node.** Today
it is node-scoped and held on every node, with a stated (untested) story that a different VPN could
hold it per machine — which would force every provider module to independently implement receiving
and applying the controller-composed configuration. The mesh does not work that way and should not
pretend to: **one mesh decides one private network.** The seat is mesh-scoped, held by the server
module (WireGuard on the hub, novox). A machine that joins is given a **client module** that
receives the composed configuration and applies it; when a node joins, the mesh emits each node's
configuration so all of them know each other at once. This drops per-node VPN choice deliberately —
the private network is nox-mesh's own, and it defines the nodes' configuration rather than being
assembled from each node's opinion. (Implementation: the overlay generator's per-node computation
is unchanged; what changes is the seat's scope and the server/client split of the module.)
- **`the-showcase` → removed from the set; it becomes a module-defined seat.** It is a demo module's
own coordination role, claimed by nothing else and held nowhere. It is the first module-defined
seat, and the reason the rule above is needed rather than hypothetical.
- **`the-dns-port` → `node-dns-resolver`** (the daemon that binds `:53`), kept distinct from
**`the-resolver-configuration` → `node-resolver-config`** (what writes `resolv.conf`). Two roles,
two seats; the rename must not blur them.
- **`the-packet-filter` → `node-packet-filter`** and **`the-intrusion-prevention` →
`node-intrusion-prevention`** — names kept as-is but for the prefix. "Packet filter" stays distinct
from "firewall", which would swallow intrusion-prevention too.
- **`the-uplink` → `node-uplink`** ([ADR 0125](0117-a-machines-uplink-is-a-seat.md)). Unheld, so it
renames with no migration.
- **The registry seats — `the-artifact-store`, `npm-package-registry` (→ `mesh-artifact-store`,
`mesh-npm-package-registry`) — and `git` (→ `mesh-git`) — are decided but deferred.** They each
*deliver* a provision, so renaming them is a delivering-seat migration: a holder that stops
resolving mid-flight takes a provision away from every consumer. That risk is not worth carrying in
the same pass as the node-* renames, so they keep their names until done deliberately.
**`distribution` stays the mesh's registry; only `verdaccio` is retired.** An earlier draft of this
record had the registry consolidating onto gitea and `distribution` retired — that was reversed:
`distribution` is the standalone OCI registry serving every `artifact-store://…@sha256` image (the
control plane's own included), and the mesh keeps it. `verdaccio` was a *second* npm registry;
gitea already provides `npm-package-registry`, so verdaccio is redundant and is removed. It is only
in the catalogue (never registered in the running mesh), so removing it is deleting the module — no
migration, nothing to strand.
## Consequences
- **A reader learns a seat's scope from its name.** `mesh-*` is mesh-wide and one; `node-*` is
per-machine. The mesh's own roles finally look like the mesh's.
- **Applications get their own seats** without the control plane learning their meaning. The closed
set shrinks to what it should be — the system's roles — and stops being where unrelated roles hide.
- **The renames are a coordinated migration, not a rename.** A held seat's name lives in three places
that must move together: the control plane's set (`seats.go`), every claiming manifest, and what
each node reports it holds (re-derived by re-registering the manifest and re-pushing). A seat
renamed in one place and not the others stops resolving to its holder — and for a *delivering* seat
(`mesh-store`→postgres, `mesh-broker`→amqp, the registry seats) that is a mesh-wide provision
outage, the same failure mode as a schema change hitting an old manifest. So: the non-delivering
`node-*` seats and `mesh-build-machine` migrate as one tested controller+catalogue change;
`node-uplink` is free (unheld); the delivering registry seats are deferred to their own pass.
- **The node-* migration was done as one controlled step, and it froze briefly.** Deploying the new
controller made it reject the still-old-named claims in the stored manifests, so composition stopped
for the affected nodes until each manifest was re-registered under its new name; running services
were untouched, and the window was seconds. This is the coordinated-migration cost named above,
paid once — and the reason the *delivering* registry seats, whose freeze would be a provision
outage rather than a compose pause, are not folded into the same pass.
- **`distribution` is not retired.** It stays as the registry; only `verdaccio` (a redundant second
npm registry) is removed. The mesh keeps one OCI registry (`distribution`) and gitea for npm/git —
the "one registry, on gitea" idea was considered and dropped.
- **The private network stops pretending to be swappable per node.** The gain is a coherent
server/client model matching how the controller already composes configuration; the cost is that
choosing a different VPN is now a mesh-wide change, not a per-node one — accepted.
## References
- [ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md) — the closed set this refines
- [ADR 0125](0117-a-machines-uplink-is-a-seat.md) — `the-uplink`, renamed here to `node-uplink`
- [ADR 0079](0079-the-foundation-seats-are-named-after-their-servers.md) — the original `mesh-*` seats
whose naming this generalises
- [to-be 26](../03-DESIGN/01-to-be/26-the-seats.md) — the seat table, updated by this
- mesh-controller `internal/catalogue/seats.go` (the set and claim validation),
`internal/overlay/generator.go` (the private network as server + client)
@@ -0,0 +1,107 @@
---
topic: what runs on it
status: accepted
date: 2026-09-27
deciders: jochen
reconstructed: false
supersedes-in-part:
- 0110-a-seat-is-a-module-assignment-from-a-closed-set.md
- 0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md
---
# 122. A seat is data the controller owns, and a rename is a database update
## Context
[ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md) made the seats a closed set the
control plane defines, and [ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md)
named them by scope. Both were right about *what* a seat is. Both left it defined the wrong *way*:
**the set is a hardcoded Go slice compiled into the controller, and everything references a seat by
its name as a string literal.** Renaming `the-packet-filter` to `node-packet-filter` this session
took, in one pass:
- an edit to the Go slice in `internal/catalogue/seats.go`, recompiled into a new controller image;
- an edit to a `const gitSeat = "git"` in *production* control-plane code (`source.go`), because a
seat's name was hardcoded where a repository's home is resolved;
- edits to every claiming manifest in the catalogue, each re-registered;
- a controller **rebuild and redeploy**, which — because the running controller then refused the
still-old-named claims in stored manifests — **froze composition** for the affected nodes until
each manifest was re-registered under its new name;
- the same coupling in the **build machine**, which embeds the same seat set and refused to build
anything claiming a name it did not yet know;
- a **deadlock** when the build machine's own seat was renamed, since the old builder could not
build the new builder whose manifest claimed a name it rejected.
None of that is what a rename should cost. A rename is the operator changing a label. It should be a
single write, and nothing should have to be rebuilt, refused, or unfrozen. The set being *closed*
(0110) and *named by scope* (0121) are good rules; **the set being code is the mistake.** When
adhering to the design means twenty steps and a `const` in the resolver, the design is what to fix.
## Decision
**The seat set is data the control plane owns, not code it is compiled from.** The seats live in a
table in the controller's store — one row per seat: a **stable id**, a `name`, a `scope`, what it
`delivers` (a provision, or nothing), and the record that decided it. The rows are seeded by a
migration (the closed set 0110 defines still ships with the mesh), and thereafter they are ordinary
data the control plane reads and writes.
**A seat is referenced by its stable id, never by its name.** A claim, a held-seat record, and any
control-plane code that must name a seat (the git-seat resolver, the artifact-store guard) hold the
**id**. The `name` is a label for people and for what a manifest writes; it is resolved to an id
once, when a claim is registered. So:
- **A rename is one `UPDATE seats set name = … where id = …`.** Nothing is recompiled, nothing is
re-registered, nothing is refused, nothing freezes. Held records and claims already point at the
id, so they follow the rename for free. The build machine is not involved, because the build
machine validates a claim against the set it reads from the mesh, not one baked into its image.
- **Adding or removing a seat is an `INSERT`/`DELETE`** (within the closed-set discipline: a change
to the set is still a decision with a record — the record is now a row's `decided` column and an
ADR, not a line of Go). No controller release is needed to change the roster of roles.
- **Production code stops hardcoding names.** `const gitSeat = "git"` becomes a lookup of the seat
that delivers the `git` provision (or a well-known id), so renaming its label cannot break the
code that finds a repository's forge.
**What does not change** (0110 and 0121 still hold): a seat is still a module assignment from a
closed set; there is still one holder per scope; a delivering seat is still the single answer for
its provision; system seats are still `mesh-*`/`node-*` and a module may still define its own. Only
their *storage and reference* change — from a compiled slice keyed by name to a table keyed by id.
**A manifest still claims by name, and that is fine.** A manifest is written by a person and names
the seat in words; the mesh resolves the name to an id at registration and stores the id. If a
seat's name changes, manifests written against the old name are updated in the catalogue like any
other edit (and the mesh can keep the old name as an alias row during a transition so nothing breaks
in the window) — but the *control plane* never has to change or redeploy for it, which is the whole
point. The heavy, mesh-wide, freeze-prone half of a rename disappears; only the ordinary catalogue
edit remains.
## Consequences
- **A rename, and a set change, become operations, not releases.** The pain this session paid —
three freezes, a builder deadlock, hand-resolved manifests — is designed out. The seat migrations
still outstanding (the delivering registry seats, and the private network's scope change) should
wait for this: done as data, each is a write, not a coupled multi-repo deploy.
- **The controller gains a small table and a seed migration**, and its seat lookups change from
slice scans to id-keyed reads. `SeatNamed`, `SeatDelivering`, `claimProblems` read the table.
- **The build machine reads the set from the mesh** (it already talks to the control plane), rather
than embedding it — which removes the controller/builder seat coupling that made every breaking
seat change a two-sided deadlock (see [to-be 30](../03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md)).
- **The closed set is still closed.** Data being editable is not the set being open: changing it is
still a decision, still recorded. What changes is that recording it no longer means shipping a
binary.
- **This is a real refactor**, touching the store schema, the seat lookups, claim registration
(name→id resolution), and the held-seat records. It is worth its own build; until it lands, the
current compiled set stands and further renames are held rather than forced through the heavy path.
- **Config on a seat is still the module's** (the question that surfaced this): a seat row carries
the seat's own metadata (scope, delivers, protocol), not a module's configuration — that stays in
the holding module's manifest ([ADR 0046](0046-a-module-configuration-is-its-assignments-not-its-manifest.md)). Making
seats data does not make them a config store.
## References
- [ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md),
[ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md) — the seat
rules this keeps, whose *storage* it changes
- [to-be 30](../03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md) — the controller/builder
seat coupling and the breaking-change freeze this removes for seat changes
- mesh-controller `internal/catalogue/seats.go` (the compiled slice this replaces),
`cmd/mesh-controller/source.go` (`const gitSeat`, the hardcoded name this removes)
@@ -0,0 +1,133 @@
---
topic: the mesh
status: superseded
superseded-by: 02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md
date: 2026-09-26
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0106-the-bus-is-nats.md
---
# 125. The bus is the only broker
## Context
[ADR 0106](0106-the-bus-is-nats.md) moved the mesh's bus to NATS and kept the AMQP broker "as a
module with one purpose — the predecessor's clients", retiring with the last of them.
[Design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md) repeats that: a compatibility module with
a single purpose and a retirement condition.
**It is not single-purpose, and was not when that was written.** Two modules of the *new* mesh
declare `requires: ["amqp"]` and are answered by the broker module's own provisioner:
- `amqp-ping`, whose source says it "exists to PROVE the grant end to end: the mesh gave it a
scoped login and a vhost of that name on the lavinmq provider";
- `amqp-email-forwarder`, which uses it for work.
What that provisioner answers is **not the mesh's bus**. Its own comment draws the line: a
consumer gets "its own message broker, isolated from every other consumer's by the vhost
boundary… a broker of its own, not a shared account on the mesh's control-plane broker" —
vhost-per-login, "the exact analog of postgres's database-per-login."
So two different things wear the word *broker*: the mesh's nervous system, and a private message
broker handed to a module as a resource, the way a database is. The first is being replaced. The
second was never examined, and on the retirement condition ADR 0106 sets, it disappears with no
successor and nothing notices — a module of the new mesh left requiring something no provider
answers.
The operator's direction, asked at the point this surfaced: **NATS is the heart of the
application** — not a component it contains, and not a thing to reproduce the predecessor's
shapes on.
## Considered Options
1. **Carry the private broker forward onto NATS** — each requiring module gets its own NATS
account, provisioned like a database. Rejected on three counts. It reproduces the
predecessor's shape on the new bus, which is the thing this whole move exists to stop. It
gives the mesh two messaging models, so "how does a module send a message" has two answers
depending on a manifest line. And NATS accounts isolate subject spaces *entirely*: a module
inside its own account cannot reach the mesh's bus at all, so it would hold two connections
and two identities to do one job.
2. **Keep the compatibility broker indefinitely** for the mesh's own modules. Rejected: its
retirement condition is the point of it. A module of the new mesh depending on the retired one
keeps the predecessor alive permanently, which is the opposite of a compatibility module.
3. **One bus. A module's messaging is subjects on it, scoped by what it declares.** Adopted.
## Decision
**The bus is the only broker.** NATS is the mesh's one messaging system, and every module's
messaging is subjects on that bus under its own account, scoped by its `emits` and `consumes`
([ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md)). There is no second
broker, and none is handed to a module as a resource.
**The `amqp` interface is not carried forward.** It leaves the set of things a module may require
and retires with the compatibility broker rather than gaining a successor.
Concretely, in the controller's seat table: **the `mesh-broker` seat delivers nothing.** It
currently reads `Delivers: "amqp"` — the seat's holder answers a requirement for a broker — and
under this decision it joins `mesh-controller` and `the-catalogue`, the foundation seats that
deliver no provision at all. The bus is not something a module asks for; it is what a module is
reached through.
- `amqp-email-forwarder` moves to the bus like any module: what it emits and consumes, declared,
and the account follows.
- `amqp-ping`'s *purpose* is kept and its mechanism is not. Proving end to end that a module
receives scoped messaging it did not configure itself is worth a probe; it becomes a probe of
the bus, and its assertion changes from "I reached my own vhost" to "I reached exactly my
subjects and was refused the rest."
**A module that wants a queue of its own has one already**: a subject nothing else may publish to
and a durable consumer of its own, both derived from its declaration. What it does not get is a
server of its own.
**The mesh's own streams are the controller's, created at genesis, not provisioned** — and
`EVENTS` is one stream, closing the question [design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md)
§11 left open. The reason is not preference but **bootstrapping**: a provisioner is a module, and
a module needs a bus account before it can run at all. Anything the bus itself is made of must
exist before the first module starts, so it is composed as configuration
([ADR 0106](0106-the-bus-is-nats.md): never through a management API) rather than provisioned by
something that could not yet be running.
## Consequences
- **Design 25 gains the distinction and loses the "single purpose" claim**; its §11 question about
the `EVENTS` stream closes here.
- **Nothing in [design 27](../03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md)'s
model changes** — the four provider kinds, the contract, resolution all stand, and it never
enumerated interfaces, so there is nothing to strike from it. What changes is that messaging
leaves the set of things resolved at all: every module has it by existing.
- **One line of the controller's seat table changes**, and it is the load-bearing one:
`mesh-broker` stops declaring what it delivers. A requirement for `amqp` then resolves to
nothing and is refused at assignment, which is how the two modules below are found rather than
discovered at runtime.
- **Two modules have conversion work**, and it belongs to step 4 of
[ADR 0116](0116-the-bus-is-built-in-five-steps.md), with the flows. Neither blocks step 1.
- **The compatibility broker becomes what ADR 0106 already called it** — single-purpose — once
those two have moved. That record's claim was wrong when written and is made true by this one.
- **What got harder:** a module that genuinely wanted an isolated server — a tenant boundary at
the broker rather than at the subject — no longer has that option, and would have to argue for
it as a new decision. That is the intended cost: one bus is the point.
## How it is checked
- **A module's messaging works with no `requires` line for it.** A lab bed: a module declaring
only `emits` and `consumes` reaches its subjects, and is refused every other — which is
[ADR 0116](0116-the-bus-is-built-in-five-steps.md) step 1's permission bed, already required.
- **Nothing requires `amqp`.** With the seat delivering nothing, a module still declaring it is
refused at resolution — the existing "requirement no provider answers" path, not a new check. A
catalogue test asserts no module declares it once the two have moved.
- **The probe proves the claim it is named for.** `amqp-ping`'s successor fails if a module can
reach a subject outside its declaration, not merely if it cannot reach its own.
- **The compatibility broker's retirement condition can actually be met.** A check that no module
of the mesh — as opposed to a predecessor client — holds a connection to it.
## References
- [ADR 0106](0106-the-bus-is-nats.md) — the bus is NATS; corrected here on what the compatibility
broker serves.
- [ADR 0116](0116-the-bus-is-built-in-five-steps.md) — the steps; the conversions land in step 4.
- [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) — the scoping that
makes one bus safe.
- [design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md),
[design 27](../03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md) — the two documents
this changes.
@@ -0,0 +1,145 @@
---
topic: the tiers
status: accepted
date: 2026-09-26
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md
---
# 126. A module declares its own seats; the mesh reserves its own
## Context
[ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md) closed the set of seats. Its
evidence was strong and still is: nothing could answer *which seats does this mesh have, and who
holds each*. Answering it meant reading every manifest in two repositories and then the
controller's own code, and when that enumeration was done by hand while writing the record, **it
reported eleven claims where there were thirteen.** The fix was a table in the controller, and
adding a seat became a decision.
What that table cannot express is the architecture [ADR 0125](0125-the-bus-is-the-only-broker.md)
opened. With one bus and no private brokers, a module offering a service to other modules offers
it as **a role on the bus**: a set of subjects, exactly one holder, addressed by what it does
rather than by which module or node provides it. A telegram sender, a licensing master, anything
a mesh might want one of. Under a closed table, adding any of those means editing the controller
— so a capability contributed by a module would require a change to the mesh itself, which is the
coupling the module system exists to prevent.
**The two requirements look opposed and are not.** 0110 needs the set *enumerable*. The
architecture needs it *extensible*. Those conflict only if enumerable means *written down in one
place by hand* — which is exactly the property that let the count drift in the first place.
## Considered Options
1. **Keep the closed table, add each new seat by decision.** Rejected. Every capability a module
contributes would need a change to the controller and a record before it could be offered, and
the mesh would carry the names of services it does not itself implement.
2. **Free-form seats, as before 0110.** Rejected for 0110's own reason, unchanged: nothing can
say what a mesh has, and a name invented at a claim site is a name nobody can explain later.
3. **A set that is closed at any moment and derived rather than maintained**, with the mesh's own
seats reserved by name. Adopted. 0110 weighed options 1 and 2 and never considered this one.
## Decision
**A seat may be declared by a module, and the set of seats a mesh has is derived: the mesh's own,
plus those declared by every module it has registered.** The set is still closed — a seat named
nowhere is refused — but it is computed from the catalogue rather than written in the controller.
Everything 0110 decided about what a seat *is* stands untouched: one holder at its scope; a
definition says which seats a module *can* hold and an assignment says which it *does*; holding
one may deliver a provision; a seat makes a role singular, never a module.
**Enumeration is a query, not an inventory.** The catalogue knows every registered manifest, so
"which seats does this mesh have, and who holds each" is answered by asking it. This is a
stronger answer than the table gave, not a weaker one: a derived list cannot drift from reality,
and drift is how the hand-made count came out at eleven of thirteen.
**The mesh's own seats are reserved by prefix.** Every seat the mesh itself defines is named
`mesh-*`, and a module declaring any `mesh-*` name is refused at registration. The prefix *is*
the reservation rule — no list of reserved names to maintain, and no way for the mesh's own
namespace to be colonised by a manifest. This requires renaming the seats that drifted from
[ADR 0079](0079-the-foundation-seats-are-named-after-their-servers.md)'s convention: `the-catalogue`
becomes `mesh-catalog`, `git` becomes `mesh-git`, and the node-scoped `the-build-machine`,
`the-dns-port`, `the-intrusion-prevention`, `the-packet-filter`, `the-private-network`,
`the-resolver-configuration`, `the-showcase` take the same prefix.
The mesh's seats stay the mesh's for a reason that does not apply to a module's: **the mesh's own
code looks them up by name.** The resolver *is* the thing that finds the store. `mesh-store` is
not a convention the controller follows, it is an identifier the controller dereferences.
**A declared seat carries a protocol.** A module declaring a seat says what may be sent to it,
what it emits, and what it serves. The holder must satisfy it; a module may not claim a seat whose
protocol it does not implement. Callers declare that they use the *seat*, never the module, so
replacing the implementation changes nothing for any caller.
**A seat is for a role; an event stays addressed to its emitter.** The two are not
interchangeable and the choice is not stylistic. An event is *this happened to me* — the emitter's
identity is the meaning, which is why the envelope carries source, node and time
([ADR 0042](0042-the-shape-of-an-event-on-the-wire.md)); routing it through a role would erase the
provenance an audit needs. A seat is *this capability, whoever provides it* — where not knowing
the holder is the point. Publish an event when the fact is about you; declare a seat when you are
offering something another module could offer instead.
**Two modules declaring the same seat name is refused at registration**, second one loses.
Registration is the last moment the mesh can still say no, and a seat name meaning two different
protocols is the failure nobody could diagnose afterwards.
## Consequences
- **The controller's seat table stops being the set** and becomes the mesh's own reserved entries.
Resolution reads the catalogue for the rest.
- **Ten seats are renamed.** A rename is a migration, not an edit: existing assignments hold the
old names, so the change carries a mapping and is applied once, and the lab beds that name seats
are updated with it.
- **A `uses` naming an undeclared seat is refused at registration**, which is where 0110's
guarantee lands under this model — the same refusal, at the same moment, from a derived set.
- **Adding a capability stops requiring a decision record.** That is a real loss of governance and
the intended trade: the argument for a seat's existence moves into the module that declares it,
where it is reviewed as part of the manifest. The mesh's own seats keep the old bar.
- **[ADR 0041](0041-events-are-a-relationship.md)'s machinery claim is already stale** for a
different reason, and is corrected in place there under the rule in
[`README.md`](README.md) — a progressive insight: on JetStream a subscription is a durable
consumer, a real object someone must create.
- **What got harder:** a seat's protocol is now a compatibility surface between modules that do
not know each other. Changing one breaks callers already bound to it, and nothing here says how
that is versioned. It is the first thing to answer in the design, and the thing most likely to
hurt later rather than now.
## How it is checked
- **The overview answers, and is right.** A command lists every seat, its scope, its protocol and
its holder, derived from the catalogue — and a test asserts the count against a fixture mesh,
because an enumeration nobody checks is how thirteen became eleven.
- **`mesh-*` is refused to a module.** A registration test: a manifest declaring `mesh-anything`
is refused, naming the prefix as the reason.
- **An undeclared seat is refused.** A registration test on `uses`, and a resolution test that
nothing reaches runtime unresolved.
- **A second declarer loses.** A registration test: two manifests, same seat name, the second
refused and the first untouched.
- **A holder must satisfy the protocol.** A claim whose module does not serve what the seat
declares is refused at assignment, not discovered when a caller times out.
## Progressive insight
> **Progressive insight — 2026-09-26.** *A seat rename is not a data migration.* This record's
> consequences say "a rename is a migration, not an edit: existing assignments hold the old
> names, so the change carries a mapping and is applied once". Implementing it showed there is
> nothing stored to migrate: a seat's holding is **derived at resolution** from the claims in
> manifests (`resolve.go` builds it each time), never written down, so no recorded name is left
> pointing at the old one. What exists is source — the controller's seat table, the manifests
> that claim them, and a manifest that may be registered later from its own repository. So the
> change is an edit plus a **kept** rename table, which tells a manifest written against an old
> name what it became rather than refusing it as unknown.
>
> The decision — that modules declare seats, that the mesh reserves `mesh-*`, and that the ten
> are renamed — is unchanged. Only the shape of the work was wrong.
## References
- [ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md) — superseded here; its
requirement is kept and only its mechanism replaced.
- [ADR 0125](0125-the-bus-is-the-only-broker.md) — one bus, which is what makes a role addressable.
- [ADR 0079](0079-the-foundation-seats-are-named-after-their-servers.md) — the naming convention
the reserved prefix restores.
- [ADR 0041](0041-events-are-a-relationship.md) — the event half of the boundary drawn here.
@@ -0,0 +1,104 @@
---
topic: the mesh
status: superseded
superseded-by: 0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
date: 2026-09-26
deciders: jochen
reconstructed: false
supersedes: 02-DECISIONS/0125-the-bus-is-the-only-broker.md
---
# 127. AMQP is a provision, not the bus
## Context
[ADR 0125](0125-the-bus-is-the-only-broker.md) decided that the bus is the only broker, and went
one step further than it had grounds for: it also decided that the `amqp` **interface** — a module
requiring a message broker of its own — "is not carried forward" and "retires with the
compatibility broker rather than gaining a successor", with the two modules declaring it converted
to the bus in step 4.
The operator's correction: **AMQP is deprecated as the mesh's transport, not abolished as a
service.** The broker module keeps running and keeps answering `amqp` requirements. It is no
longer a core part of the mesh — *"it's just a module like mssql now."*
**What 0117 conflated** is two different reasons a module might ask for a broker, which look
identical in a manifest:
1. **To talk to other modules.** Wrong under one bus, and the thing 0117 was right to refuse: a
private broker used as inter-module transport is a second bus, with every guarantee crossing a
seam and no scoping the mesh can see.
2. **Because it genuinely needs an AMQP broker**, the way something needs a database — a queue for
its own internals, or interop with software that speaks AMQP and nothing else. That is a
backing service, and the mesh has a word for backing services already.
0117 saw the first and legislated against both. The second is ordinary, and forbidding it would
make the mesh unable to run a large class of perfectly normal software while claiming that as
architecture.
## Considered Options
1. **Keep 0117 as written** — retire the interface, convert the two modules. Rejected by the
operator, and wrongly reasoned besides: it treats "needs an AMQP broker" as always a mistake.
2. **Keep the broker as the predecessor's compatibility module**, as ADR 0106 framed it, with a
retirement condition. Rejected: it is not single-purpose and its clients are not only the
predecessor's, so the retirement condition describes a day that will not come.
3. **The broker is an ordinary provider module of an ordinary provision.** Adopted.
## Decision
**The mesh's bus is NATS and only NATS.** Everything 0117 decided about *the bus* stands: one bus,
a module's messaging is subjects on it scoped by what it declares, no module is handed a bus of
its own, and the `mesh-broker` seat is the NATS server's.
**`amqp` remains a provision a module may require**, answered by the broker module the way
`postgres-database` is answered by the store module or a database is answered by mssql. It is not
deprecated as an interface; the software behind it is simply no longer the mesh's nervous system.
**The broker module stops being foundation.** It claims no seat — `mesh-broker` is the NATS
server's — it is not raised at genesis, nothing in the mesh requires it, and a mesh that never
installs it is a complete mesh. It is installed when something wants it, like any other provider.
**The rule that survives, stated so it can be applied:** *inter-module communication goes over the
bus.* A module may hold a broker, a database or a cache as a backing service; it may not use one
as a channel to another module. The line is not which software is involved, it is whether a second
module is on the other end.
**Neither `amqp-ping` nor `amqp-email-forwarder` needs converting.** 0117 put that work in step 4;
it is removed. They require a backing service and a provider answers.
## Consequences
- **The "compatibility broker" framing is wrong and goes.** There is no `lavinmq-compat`, no
single purpose and no retirement condition. Design 25 §5 is corrected.
- **[ADR 0106](0106-the-bus-is-nats.md)'s progressive insight was itself wrong** and is corrected
by a second one there. It said 0117 would make 0106's "one purpose — the predecessor's clients"
sentence true by moving the mesh's modules off. Nothing moves off; the sentence is simply not
what the broker is.
- **The seat change stands**, for a better reason than 0117 gave: not because a broker cannot be
provisioned, but because *this* broker is not the mesh's bus. The broker module drops its
`mesh-broker` claim and the `nats` module takes it.
- **Step 4 loses two conversions**; step 1 and the WBS are otherwise unaffected.
- **What got harder:** the rule is now a judgement rather than a prohibition. "Is this a backing
service or a channel to another module?" has to be asked in review, where 0117 could have
answered it with a parser. That is the honest cost of allowing the legitimate case.
## How it is checked
- **A module's own messaging needs no `requires`.** The check from 0117, unchanged: a module
declaring only `emits` and `consumes` reaches its subjects and is refused every other.
- **The broker holds no seat.** A manifest test: the broker module claims nothing, and a mesh
raised without it is complete — genesis names it nowhere.
- **`amqp` resolves like any provision.** A resolution test: a module requiring it is answered by
the provider, refused when none is assigned, and neither case touches the bus.
- **What cannot be checked mechanically**, and is said rather than implied: that a module holding
a broker is not using it to reach another module. Review, not a parser.
## References
- [ADR 0125](0125-the-bus-is-the-only-broker.md) — superseded; its ruling on the bus is kept
whole and only its ruling on the interface is reversed.
- [ADR 0106](0106-the-bus-is-nats.md) — the bus is NATS; its compatibility-broker framing is
corrected here.
- [ADR 0126](0126-a-module-declares-its-own-seats.md) — seats, including the one the NATS server
now holds alone.
@@ -0,0 +1,123 @@
---
topic: the mesh
status: accepted
date: 2026-09-26
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
---
# 128. The mesh bus is required, not ambient
> **Pointer repointed, 2026-09-27.** This record was written extending
> [ADR 0127](0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)), which
> [ADR 0131](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md) has since superseded — AMQP is
> not a provision at all. Nothing decided here changes; the frontmatter now rests on the live record,
> and the citations below are read with that in mind.
## Context
[Design 29](../03-DESIGN/01-to-be/32-what-a-module-declares.md) opened by saying the bus is
*ambient*: "No module requires it, the way no module requires a filesystem. Every module gets a
connection and an identity whether it asks or not."
**Two counts say that is wrong.** Of the 72 modules in the catalogue, **49 declare an own-secret
named `broker` and 23 do not.** So the bus is not universal — nearly a third of the catalogue
never speaks to it — and an ambient connection would mint an account, a password and a permission
set for every one of those 23, each a credential nothing uses and everything must rotate.
And the 49 that do take one **each hand-write the path it lands at**
(`own-secrets: { broker: "/var/lib/<module>/broker" }`). That is a special case doing badly what
provisioning already does well: a consumer names where a credential lands, the mesh seals it
there, and rotation and removal follow the same path as every other credential.
**The argument that made the bus ambient was narrower than it looked.**
[ADR 0125](0125-the-bus-is-the-only-broker.md) reasoned that bus accounts cannot be provisioned
because a provisioner is itself a module that needs an account before it can run. That is true of
a **provisioner process**, and it is not true of a provision: the mesh's bus accounts are composed
by the *controller*, into configuration, and the controller is not waiting on a bus account to
exist. The circularity is real for one mechanism and absent for the other, and the earlier record
applied it to both.
## Considered Options
1. **Keep the bus ambient.** Rejected on the counts above: it over-grants to 23 modules and keeps
a hand-written path in 49.
2. **Derive the requirement** from whether a module declares any `emits`, `consumes`, `serves` or
`uses`. Rejected: it is the ambient model with extra inference. A reader of a manifest still
cannot see that the module holds a bus credential, and the rule would have to be re-derived
every time the set of bus-facing declarations grew.
3. **The mesh bus is a provision a module requires**, delivered by the seat that holds it.
Adopted.
## Decision
**A module that speaks to the mesh requires `mesh-bus`, and receives what it needs to connect.**
The contract is an address, a credential sealed to the module, and the trust to verify the
server. It lands where the module's manifest says, like any provision. A module that does not
require it gets no account, no password and no permissions — and 23 modules in the catalogue
should get none.
**The `mesh-broker` seat delivers `mesh-bus`.** Its holder is the mesh's own bus, and what
holding it delivers is the connection to that bus — which is what a seat delivering a provision
has always meant ([design 26](../03-DESIGN/01-to-be/26-the-seats.md)).
**The requirement delivers the connection; the declarations shape the authority.** They are two
different things and both stay explicit. `requires: mesh-bus` says *this module talks to the
mesh*; `emits`, `consumes`, `serves`, `uses` and a declared seat say *what it may say and hear*,
and the permission set is derived from those and nothing else
([ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md)). Requiring the bus
grants no subject; declaring a subject without requiring the bus is refused at registration as
incoherent.
**The `mesh-bus` provision is answered by the controller, not by a provisioner.** This is the
surviving kernel of ADR 0125's bootstrap argument, narrowed to what it actually supports: the
bus's accounts are configuration the controller composes and the server reloads
([ADR 0106](0106-the-bus-is-nats.md) — never through a management API), so there is no provisioner
process in the path and nothing waiting on a bus account to create bus accounts. It is a provision
whose provider is the mesh itself.
**A module may also provide a NATS server of its own, and that is a different interface.** Exactly
as the AMQP broker provides `amqp` ([ADR 0127](0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md))), a module
may run its own NATS and offer it as a backing service. That interface is **`nats`**; the mesh's
own bus is **`mesh-bus`**; the two are never the same name, because a manifest that said `nats`
could mean either and the difference is the whole architecture. The rule from 0119 decides which
is legitimate: a private bus is a backing service, never a channel to another module.
## Consequences
- **Design 29's opening is reversed.** The bus is not ambient; it is required, and the document's
first paragraph says the opposite of this.
- **The seat's `Delivers` is `mesh-bus`** — corrected twice in one day, which is worth recording
rather than tidying: it read `amqp`, which was the old broker's interface; ADR 0125 emptied it,
on the reasoning that a bus cannot be provisioned; and it is neither. The seat delivers the
mesh's bus.
- **`own-secrets: { broker: ... }` is retired** in favour of the provision's own delivery, across
49 manifests. That is a mechanical change, and it belongs with the conversions in step 4 rather
than step 1.
- **23 modules lose a credential they never used.** Not a regression — an over-grant removed, and
the smallest honest statement of what this buys.
- **What got harder:** one more line in most manifests. The trade is that the line is true, and
its absence is also true.
## How it is checked
- **A module with no `requires: mesh-bus` has no account.** A composition test: the derived user
list contains exactly the modules that require it, and the 23 that do not appear nowhere in it.
- **Declaring a subject without requiring the bus is refused.** A registration test on a manifest
with `emits` and no requirement, naming the contradiction.
- **Requiring the bus grants no subject on its own.** A composition test: a module that requires
`mesh-bus` and declares nothing else gets a connection and an empty permission set.
- **`nats` and `mesh-bus` are distinct interfaces.** A resolution test: a module requiring `nats`
is answered by a module providing it, never by the seat holder, and vice versa.
## References
- [ADR 0125](0125-the-bus-is-the-only-broker.md) — superseded by 0119; its bootstrap argument is
narrowed here to the case it supports.
- [ADR 0127](0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)) — a broker as a backing service; this
applies the same shape to the mesh's own bus and separates the two names.
- [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) — authority from
declarations, which this leaves untouched.
- [design 26](../03-DESIGN/01-to-be/26-the-seats.md) — a seat delivering a provision.
@@ -0,0 +1,102 @@
---
topic: the mesh
status: accepted
date: 2026-09-27
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0122-a-seat-is-data-a-rename-is-a-database-update.md
---
# 129. A seat carries the protocol of its role
## Context
[ADR 0126](0126-a-module-declares-its-own-seats.md) let a module declare a seat with its protocol:
what work the role accepts, what it emits, what it serves. A module's own seats work that way today.
**The mesh's own seats — the `mesh-*` set — carry no protocol at all**, only a name, a scope and the
provision they deliver. They say who does a job and nothing about what may be said to them or by
them.
That gap surfaced three times in one day, each time as a different-looking problem.
**A build machine.** On the bus the mesh runs on today a builder has its own account kind, created by
its own command, with permissions written by hand: read the build queue, write to two exchanges. One
publish to a shared exchange reached all three audiences a finished build has — whoever asked, the
controller that records it, and the catalogue that places it in the module graph. On a bus where
permissions are per subject those are three separate grants, and nothing derives them, because a
builder is not a module and holds a seat that promises nothing.
**An event about a role rather than about a module.** The module holding the artifact-store seat
declared an event named after a *different* module
([issue 127](../04-ISSUES/127-a-module-event-derives-a-subject-nothing-publishes/00-report.md)). The
bus refuses that, because a namespace belongs to who it is named for. The event is genuinely about the
role — "the artifact store accepted an image" — and a consumer written against whichever module holds
that role today breaks when the holder changes. There was nowhere else to put it.
**A catalogue catching up.** The controller answers a request for builds it may have missed by
re-publishing them under its own name, which no consumer of the builder's subject hears. Publishing
them under the builder's name would be the controller signing an event as another module. Answering
into the asker's inbox needs a grant over every inbox in the mesh, which
[design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md) §4 refuses.
Three symptoms, one cause: **the mesh has roles it cannot describe.**
## Decision
**A seat carries the protocol of its role, whether the seat is a module's or the mesh's own.** The
`mesh-*` set gains the same three fields a declared seat has — what it accepts, what it emits, what it
serves — and the holder's authority, its work queue and its consumers are derived from them by the
machinery that already does this for a module's seats.
**Builds become work submitted to a role.** The build machine seat accepts a build and emits an
outcome. The dedicated `mesh.build.*` branch and the stream behind it retire: a work queue shared by
several build machines is exactly what a seat's `accepts` already is, and keeping a second mechanism
for it means two things to reason about and two places for a permission to be wrong.
**One publish still reaches three audiences, and now the mesh derived the subject.** A build's outcome
is the seat's own event. Whoever asked matches it by the id their request carried; the controller
records it; the catalogue places it. That is the fan-out the shared exchange gave for free, expressed
as a subject rather than as a topology, and it means no holder needs permission to publish into
anybody's inbox.
## Alternatives considered
**A dedicated principal kind for a builder**, mirroring the account the old bus issues it. Smaller: one
addition to the composer, no change to seats, and it matches how a builder is treated today. Not taken
because it answers one of the three symptoms and leaves the other two, and because "the builder is
special" is a claim nobody could justify from the design — a build machine is a role the mesh has, and
the mesh has a word for a role.
**Leaving the outcome as a reply to the asker's inbox.** Rejected on authority: a holder able to answer
any asker needs a grant across the whole inbox space, which is the one grant design 25 §4 refuses by
name. The seat's event costs the asker a filter and costs the mesh nothing.
## Reconciled with 0122, which landed in parallel
*Added 2026-09-27, on merging.* [ADR 0122](0122-a-seat-is-data-a-rename-is-a-database-update.md) moved
the seat set out of compiled code and into a table the controller owns. This record was written against
the slice, and says the `mesh-*` set "gains the same three fields a declared seat has".
**The decision is unaffected and the mechanism is better for it.** What a seat accepts, emits and serves
becomes three columns beside its name and scope, so giving a role a protocol is a write rather than a
rebuild — which is the whole argument of 0122 applied to the thing this record adds. Where this text
says the set gains fields, read: the table gains columns.
## Consequences
**A seat is now the mesh's unit of "a role that talks".** A role that accepts work, announces outcomes
or answers questions says so where it is defined, and everything about permissions, queues and
consumers follows. Nothing hand-writes a grant for a role again.
**The shared library cannot yet publish on a seat, and that is now the blocking gap rather than a
curiosity.** A module holding a seat has the authority and no way to use it; the build machine is
written in Go and reaches the bus directly, so it is unaffected, but the artifact-store event stays
under its module's own name until the library has a surface for this. That is a task, and this record
is what makes it one.
**A second mechanism disappears.** `mesh.build.*`, the BUILDS stream and the builder's hand-written
account all retire. Fewer things, and the ones left are derived.
**The catch-up question is not settled by this**, only made answerable: a seat that serves something
gives the controller a way to be asked, which the mesh did not have. Whether catch-up should be a
question at all remains open.
@@ -0,0 +1,71 @@
---
topic: the mesh
status: accepted
date: 2026-09-27
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
---
# 130. The predecessor is ending, and its broker goes with it
> **Pointer repointed, 2026-09-27.** This record was written extending
> [ADR 0127](0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)), which
> [ADR 0131](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md) has since superseded — AMQP is
> not a provision at all. Nothing decided here changes; the frontmatter now rests on the live record,
> and the citations below are read with that in mind.
## Context
[ADR 0127](0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)) settled that the old broker is an ordinary
provider of the `amqp` provision rather than a compatibility module with an end date. It rejected
giving it a retirement condition, and said why: *"its clients are not only the predecessor's, so the
retirement condition describes a day that will not come."*
**The operator has said that day is coming.** The predecessor is deprecated. Some of it is still
running, and it is not being migrated — it is being left to stop. Its broker may be shut down.
That is a fact about this installation, not a change of mind about what a broker is. It is recorded
because three documents reason from the premise it overturns:
[design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md) §5 and §9, and
[design 28](../03-DESIGN/01-to-be/28-building-the-bus.md)'s closing note that the predecessor's world
"does not need to move: its broker is the compatibility module until its last client is gone."
## Decision
**The predecessor's broker retires when nothing requires `amqp`, by being unassigned like any other
provider.** No retirement condition, no end-date machinery, no special case — which is ADR 0127 being
paid off rather than revised. Because that record made the broker an ordinary provider, ending it
needs nothing that does not already exist: a provision with no consumers has its provider unassigned,
and the module system has done that since it existed.
**So step 5.3 has an ending.** "The mesh's own accounts removed from the deprecated broker" was
written as the last thing that could be said, because the broker itself was going to outlive the
question. It now finishes: once the mesh's own traffic has moved and the predecessor's remnants have
stopped, the module is unassigned and the port is free.
**And the transitional doubling has a date.** The build outcome is announced under both the module's
name and the role's on the old bus, so that a catalogue deployed before the rename and one deployed
after both hear it. That exists only while the old bus does, and goes with it.
## Consequences
**The remote access path goes with it, and that is the one practical consequence worth planning
around.** The predecessor's own mesh communicates over that broker — so shutting it down ends the
tooling that reaches this installation's machines remotely. Work on the node after that point is done
from the node. **This matters most for the rollout**, which is the step that would otherwise be driven
from a workstation: it has to be driven locally, or driven before the broker stops.
**What is still running on it stops when it stops.** Some of the predecessor's services are live and
are not being moved. That is the operator's decision and it is recorded here so that nobody later reads
a broker with clients as an accident.
**Nothing in a served request's path is affected.** Modules serve from their own containers; the mesh's
bus carries the mesh's own traffic — declarations, reports, events, tool calls. This was checked rather
than assumed when the question came up, and it is why the operator's position (*"as long as my services
keep running"*) is a bounded risk rather than a gamble.
**One reason to keep the broker survives**: `amqp` remains a provision a module may require, and a
module that genuinely needs an AMQP broker can be given one. What retires is *this* broker's role as
the predecessor's, not the mesh's ability to provide the thing.
@@ -0,0 +1,94 @@
---
topic: the mesh
status: accepted
date: 2026-09-27
deciders: jochen
reconstructed: false
supersedes: 0127-amqp-is-a-provision-not-the-bus.md
---
# 131. Everything on the mesh speaks to the broker seat, and AMQP is not a provision
## Context
[ADR 0127](0127-amqp-is-a-provision-not-the-bus.md) settled the old broker as an ordinary provider
of an ordinary provision, `amqp`, kept for whatever wanted a message broker of its own. The day the
bus moved was the day that framing was tested, and it failed in a way that took the control plane
down for an evening.
Three things came out of the wreckage. **The protocol had leaked into the seat's contract**: for a
module to hold `mesh-broker`, it had to provide what the seat delivers, and what it delivered was
`amqp` — so the module that will carry the bus on NATS could not hold the seat that names the bus,
while the module the mesh was leaving could. **A consumer of `amqp` is not asking for AMQP.** The two
modules requiring it wanted the mesh's messaging — to emit an event, to hear a topic — and named the
wire protocol only because that was the word available. **And AMQP and NATS are not interchangeable
at the wire.** A provision named after a protocol can only ever be answered by that protocol, so once
the bus is NATS an `amqp` provision has one possible provider, and it is the thing being retired.
The operator's position, stated during the outage: modules depend on the broker *seat*, not on a
protocol; AMQP is obsolete as anything the mesh's core knows about; a module that depends on `amqp`
is wrong; and everything should reach the mesh's bus and be able to emit events and consume topics
through it.
## Decision
**A module that needs messaging uses the mesh's bus, and the mesh's bus is whatever holds
`mesh-broker`.** Emitting an event and consuming a topic go through the sdk, which is handed the
bus by the mesh with the module's own credential. No manifest names a wire protocol to get it.
**`amqp` is neither a provision nor a requirement.** Registration refuses a manifest that provides
it or requires it. The `mesh-broker` seat delivers `mesh-bus`, and its holder is the module that
provides `mesh-bus` — today the nats module, and only it.
**The old broker's module and the two modules that required it leave the catalogue.** They are
removed, not converted: one was a proof that a grant worked end to end, the other forwards mail off a
queue, and both are re-done against the bus if wanted, as new modules under this record.
**The controller's AMQP transport is deleted once every node reports on the new bus**, and the
switch that selects a transport goes with it — one bus, so nothing to select.
The predecessor's own broker is outside the mesh and not this record's concern
([ADR 0130](0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md)): what the predecessor's
tooling loses when it stops is accepted there.
## Options considered
1. **Keep 0127: AMQP stays an ordinary provision with the old broker as its provider.** Rejected. It
is what put the protocol into the seat's contract, it is why the seat could be left with no valid
holder mid-change, and it keeps two transports in the control plane indefinitely for the benefit of
two modules that did not want AMQP in the first place.
2. **Bridge it: the old broker's module also provides `mesh-bus`, so both can hold the seat during the
change.** Rejected. It makes the retiring broker a legitimate mesh bus for exactly as long as
nobody removes the line, which in practice is forever, and it leaves `amqp` as a thing the core
still knows the name of.
3. **The seat is the dependency; the protocol is nobody's business but the holder's.** Adopted.
## Consequences
- **The change of holder is a handover, and it needs a command.** Nothing today moves a seat from
one assignment to another as one act, and a seat the control plane dereferences cannot be empty
in between — that emptiness is the outage this record comes from. The command takes a seat and the
assignment taking it over. Designed and built before the cutover, under
[28 — Building the bus](../03-DESIGN/01-to-be/28-building-the-bus.md).
- **The seat's row moves to `mesh-bus` before the new holder registers, and that is safe.** The
control plane composes its own bus address through the seat *by name*
(`${seat:mesh-broker:…}`), and the overview derives holders by name; only registration and the
provision-to-seat resolution read what a seat delivers. So the row can change under the current
holder without unseating it, the new holder can then register its claim, and the handover happens
when both are running. Verified in the code during the outage, not assumed.
- **Registration gains two refusals**: a manifest providing `amqp`, and one requiring it.
- **The `rollout check` stops saying the old broker stays.** It said so under 0127; it now lists
unassigning it as the last step of the move.
- **What got harder**: a third party that genuinely wants an AMQP broker on a mesh node runs one as
any application module, with no provision and no seat, and nothing on the mesh routes to it. That
is the cost of the mesh not knowing the word.
## How this is checked
| Rule | Checked by |
|---|---|
| No manifest provides or requires `amqp` | a registration test refusing each, naming this record; and a whole-catalogue test asserting no registered manifest names it |
| `mesh-broker` delivers `mesh-bus`, and only a `mesh-bus` provider may hold it | the existing registration test for a delivering seat, with the row's value read from the store (mesh-controller#89) |
| The seat's row can change without unseating the holder | a test composing the control plane's own address and the overview under a row that the current holder does not satisfy |
| The rollout does not leave the old broker running | `rollout check` output, asserted in its test |
| The AMQP transport is gone | the package does not compile with it referenced; the switch variable is refused as unknown at start |
+63 -3
View File
@@ -8,8 +8,45 @@ the decision is recorded here, and only then is the design written. Following th
numbers walks the process in the order it happens.
One file per decision, numbered, never deleted. A superseded record has its `status:` changed
and gains a pointer to what replaced it — **its text is never edited**. The reasoning that was
rejected is the expensive half to rediscover.
and gains a pointer to what replaced it — **its reasoning is never rewritten**. The reasoning that
was rejected is the expensive half to rediscover.
## Progressive insight
A record is a decision, not a snapshot of everything that was true the day it was written, and
those two fail differently. **A fact a record asserted can turn out to be wrong while the decision
it supports stays right** — a count taken before anyone measured, a file named that does not
exist, a proof attributed to a step that cannot run it. Superseding a record for that buries a
correct decision under a second one, and teaches every reader to first work out which of two
records is live. Done a few times, the reading order stops being one.
So: **a correction of fact that leaves the decision standing is made in the record, in place,
marked and dated.**
> **Progressive insight — YYYY-MM-DD.** What was found, what the record said before, and what it
> says instead.
Three conditions, all of which hold:
- **It corrects a fact, not a judgement.** That a suite does not exist is a fact. That building it
is the wrong order is a judgement, and judgements supersede.
- **It adds; it never quietly replaces.** Where body text changes, the note says what stood there
before, so a reader who followed a citation to the old wording can find out what happened to it.
A correction nobody can see is indistinguishable from a record that was always right, which is
the failure the immutability rule exists to prevent.
- **The decision, the options weighed and the consequences stand untouched.** If the correction
changes what was decided, which alternatives were rejected, or a consequence another record
relies on, it is not an insight — write the superseding record.
**What still supersedes**, without exception: reversing a decision, changing its scope, rejecting
an option it accepted, or making a consequence false that a later record cites. The test is not
how large the edit looks in a diff; it is whether a reader who acted on the old text would now be
wrong about *what was decided* rather than about *a detail the decision did not rest on*.
**How this is checked.** `00-META/checks/records.py` requires every insight to be marked in the
form above and dated no earlier than the record's own `date:` — an unmarked edit is a rule
violation the reviewer looks for in the diff, and a marked one is legible in the record itself.
The git history is the backstop, not the record of intent; the note is the record of intent.
The records run in the order the decisions were taken, oldest first.
@@ -95,6 +132,14 @@ python3 00-META/checks/index.py fail if stale
- **0104** — [A provision may be answered by an adapter to the predecessor](0104-a-provision-may-be-answered-by-an-adapter-to-the-predecessor.md)
- **0105** — [The mesh adopts the predecessor's tunnel in place](0105-the-mesh-adopts-the-predecessors-tunnel-in-place.md)
- **0106** — [The bus is NATS](0106-the-bus-is-nats.md)
- **0116** — [The bus is built in five steps, and the protocol moves with it](0116-the-bus-is-built-in-five-steps.md)
- **0119** — [A taken tunnel's predecessor is retired once the take is proven](0119-a-taken-tunnels-predecessor-is-retired.md)
- **0125** — [The bus is the only broker](0125-the-bus-is-the-only-broker.md) *(superseded)*
- **0127** — [AMQP is a provision, not the bus](0127-amqp-is-a-provision-not-the-bus.md) *(superseded)*
- **0128** — [The mesh bus is required, not ambient](0128-the-mesh-bus-is-required-not-ambient.md)
- **0129** — [A seat carries the protocol of its role](0129-a-seat-carries-the-protocol-of-its-role.md)
- **0130** — [The predecessor is ending, and its broker goes with it](0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md)
- **0131** — [Everything on the mesh speaks to the broker seat, and AMQP is not a provision](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)
### Its tiers, from the bottom up
@@ -123,6 +168,9 @@ python3 00-META/checks/index.py fail if stale
- **0094** — [A module may hold several secrets from one provider, each a pair of its own](0094-a-module-may-hold-several-secrets-from-one-provider.md)
- **0095** — [The control plane is the way to ask a module](0095-the-control-plane-is-the-way-to-ask-a-module.md)
- **0098** — [A fact a provider makes at first start is fetched from it, not carried in its manifest](0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md)
- **0108** — [A route carries the policy applied to a request, and names a secret rather than holding one](0108-a-route-carries-the-policy-applied-to-a-request.md)
- **0109** — [A package registry seat is one per ecosystem, not one for all of them](0109-a-package-registry-seat-is-one-per-ecosystem.md)
- **0126** — [A module declares its own seats; the mesh reserves its own](0126-a-module-declares-its-own-seats.md)
### What runs on them, and how it gets there
@@ -132,7 +180,7 @@ python3 00-META/checks/index.py fail if stale
- **0026** — [The mesh has a session of its own, and it is the node session's mechanism](0026-the-mesh-has-a-session-of-its-own.md)
- **0027** — [A provision names what the consumer is coupled to, not the role it plays](0027-a-provision-names-what-the-consumer-is-coupled-to.md)
- **0035** — [One implementation, several surfaces, and what that costs](0035-one-implementation-several-surfaces.md)
- **0038** — [The mesh assigns the port, and a module does not care](0038-the-mesh-assigns-the-port.md) *(proposed)*
- **0038** — [The mesh assigns the port, and a module does not care](0038-the-mesh-assigns-the-port.md)
- **0040** — [What a module is](0040-what-a-module-is.md)
- **0041** — [Events are a relationship, the lighter sibling of provisioning](0041-events-are-a-relationship.md)
- **0042** — [The shape of an event on the wire](0042-the-shape-of-an-event-on-the-wire.md)
@@ -154,6 +202,16 @@ python3 00-META/checks/index.py fail if stale
- **0087** — [A seeded file is created once, and what grows in it is not the mesh's](0087-a-seeded-file-is-created-once.md)
- **0091** — [A mount is declared, and there are three things it can be](0091-a-mount-is-declared-three-ways.md)
- **0099** — [A step that runs once names what it reads, and runs again when it changed](0099-a-step-that-runs-once-names-what-it-reads.md)
- **0110** — [A seat is held by one assignment, from a closed set, and it may deliver a provision](0110-a-seat-is-a-module-assignment-from-a-closed-set.md)
- **0112** — [A module definition names no node, no mesh and no path: everything it needs is a requirement the mesh resolves](0112-a-module-definition-names-no-node-mesh-or-path.md) *(proposed)*
- **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md) *(proposed)*
- **0114** — [A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached](0114-a-shared-credential-rotates-over-two-credentials.md) *(proposed)*
- **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md) *(proposed)*
- **0117** — [A machine's uplink is a seat: the mesh configures the manager, never the link](0117-a-machines-uplink-is-a-seat.md)
- **0118** — [Undeclaring removes what the mesh made, and gives a unit back the state it was found in](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md)
- **0120** — [A roster fact carries its format as a template: the mesh owns the data, the module owns the format](0120-a-roster-fact-carries-its-format-as-a-template.md)
- **0121** — [A system seat is named for its scope, and a module may define its own](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md)
- **0122** — [A seat is data the controller owns, and a rename is a database update](0122-a-seat-is-data-a-rename-is-a-database-update.md)
### How it is built
@@ -172,6 +230,8 @@ python3 00-META/checks/index.py fail if stale
- **0086** — [A secret reaches a process as a file, and an exception is declared](0086-a-secret-reaches-a-process-as-a-file.md)
- **0096** — [An upstream image is copied between registries, never through a machine's image store](0096-an-upstream-image-is-copied-between-registries.md)
- **0097** — [A vendor image is a declared build input, and a recipe fetches nothing undeclared](0097-a-vendor-image-is-a-declared-build-input.md)
- **0107** — [Persistent data is a directory bind, never a named volume](0107-persistent-data-is-a-directory-bind-never-a-named-volume.md)
- **0111** — [A build source is on the mesh's git seat, or it is an external repository](0111-a-build-source-is-on-the-git-seat-or-external.md)
### How it is checked
+91
View File
@@ -0,0 +1,91 @@
---
layer: as-is
status: implemented
code: [mesh-controller, mesh-catalog]
updated: 2026-09-26
decisions:
- 02-DECISIONS/0109-a-package-registry-seat-is-one-per-ecosystem.md
- 02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md
- 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md
---
# The seats, as they run
Written from the controller's code and the catalogue's manifests on their main branches. This is
the second piece of the new shape to exist, after [the lab](11-the-lab.md), and the first that the
predecessor has no equivalent of: there, any well-formed name became a seat by being claimed.
## The set is closed, and it lives in the controller
Fourteen seats are defined in the controller — seven at mesh scope, seven at node scope. Each entry
carries a name, the scope at which there may be only one holder, the provision its holder answers
for (or none), and the record that made it a seat. A person reads the set to learn what a mesh can
have; nothing else can add to it.
Five seats deliver a provision: the mesh's store delivers the relational database, the mesh's broker
the message transport, the artifact-store seat the registry, and two more the package registry for
one ecosystem and the git service. The remaining nine — the controller and catalogue seats, and the
node-scope ones for the build machine, the DNS port, the packet filter, intrusion prevention, the
private network, the resolver's configuration and the showcase — mark a role without answering for
anything a consumer requires.
## A claim is checked against the set
Parsing a manifest refuses three things, in this order: a claim that is not a usable name or not at
node, site or mesh scope; then a claim naming a seat the mesh does not define, answered with the
whole set so the writer can see what was available; then a claim at the wrong scope for that seat.
A malformed claim is refused once, for being malformed, and not a second time for being unknown.
One further refusal ties a seat to what it promises: claiming a seat that delivers a provision is
refused unless the claiming module actually provides that provision, at the seat's scope. A seat
cannot be held by something that could not answer for it.
## A seat is held by an assignment, and nothing else is recorded
The holder is a module assignment — a node and a module together. There is no separate record of
holders: what the mesh knows about one is what it already knows about that assignment. The pair is
also what tells a seat's holder apart from another module providing the same thing on another node.
## Where a seat changes resolution, and where it does not
Resolution runs in two passes. The first learns only what each node offers and takes requirements
on trust; no declaration is ever built from it. In the second:
- **nothing provides the wanted provision** — refused, naming what would answer it ("assign X to a
node");
- **one provider** — taken, unless a pin names a different node, which is refused rather than
silently overruled;
- **several providers** — a pin wins; failing that, the holder of the seat that delivers the
provision answers; failing that, refused with the candidates listed.
So a seat decides **which** of several providers answers. It does not make a requirement optional:
before the seat is held, a requirement for what it delivers is refused like any other unanswerable
one. That is the standing condition [issue 121](../../04-ISSUES/121-builders-real-package-registry-grant-deadlocks-genesis/00-report.md)
records, where the module that builds the forge's image requires a registry only the forge provides.
## The mesh can be asked
A `seats` command prints every seat with its scope, what it delivers and who holds it, in reading
order, and the same as JSON. An unheld seat prints as an answer — this mesh has no such thing — not
as a fault. Claims held that are **outside** the set are listed separately rather than hidden, so a
mesh carrying one from before the set existed says so.
## A build source may name the git seat
A module's repository is either a URL, recorded and cloned exactly as given, or a path on the forge
holding the git seat, recorded as that path plus the seat. The clone URL is composed from wherever
the holder runs at the moment of building, so the build machine is never told an address that could
go stale. Before this, a self-hosted forge's scheme, host and port were written into every module
built from it, and moving the forge made every record stale at once — noticed when a rebuild failed
to clone.
The schema column added for this defaults to empty rather than null, because "not on a seat" is a
real answer, so every row recorded before the change keeps exactly the meaning it had.
## Where this differs from the design
**Capacity is not implemented.** The design's vocabulary has a seat with a capacity, and a
higher-capacity seat is a bench several holders share. Neither exists in the code: a seat as it runs
is one role with one holder at its scope, and nothing expresses a bench. Every seat in the set today
is exclusive, so nothing has yet needed it — but a design that says capacity and an implementation
that has none is a disagreement worth reading here rather than discovering in the type.
+1
View File
@@ -20,6 +20,7 @@ Where the two disagree, the implementation wins and the disagreement is stated.
| [`09-interfaces-and-observability.md`](09-interfaces-and-observability.md) | How the mesh is reached and watched — tools, board, proxy, health, thoughts |
| [`10-module-catalogue.md`](10-module-catalogue.md) | The catalogue's shape, and what its shape says |
| [`11-the-lab.md`](11-the-lab.md) | The lab — the first piece of the new shape that exists, and what it does not yet do |
| [`12-the-seats.md`](12-the-seats.md) | The seats the mesh defines, who holds one, and where a seat changes resolution |
## What these documents are not
+42 -1
View File
@@ -7,11 +7,13 @@ code:
- mesh-controller internal/identity/authority.go
- mesh-host internal/identity/serving.go
- mesh-host internal/apply (the service that reflects a rule set)
updated: 2026-09-23
updated: 2026-09-27
decisions:
- 02-DECISIONS/0104-a-provision-may-be-answered-by-an-adapter-to-the-predecessor.md
- 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0105-the-mesh-adopts-the-predecessors-tunnel-in-place.md
- 02-DECISIONS/0119-a-taken-tunnels-predecessor-is-retired.md
- 02-DECISIONS/0108-a-route-carries-the-policy-applied-to-a-request.md
- 02-DECISIONS/0103-what-an-adopted-node-holds-and-what-its-guard-refuses.md
- 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
- 02-DECISIONS/0099-a-step-that-runs-once-names-what-it-reads.md
@@ -442,6 +444,39 @@ exactly the per-module cost it is meant to remove. The design is the composition
stopgap until the manifest layer can carry a label and a domain separately
([ADR 0066](../../02-DECISIONS/0066-public-routing-is-name-agnostic.md)).
### A route also carries what a request arriving at it may do
*2026-09-25, from comparing the mesh's proxy against the ingress it would replace
([issue 116](../../04-ISSUES/116-route-proxy-has-no-auth-or-ip-restriction/00-report.md)),
decided in [ADR 0108](../../02-DECISIONS/0108-a-route-carries-the-policy-applied-to-a-request.md).*
A grant hands back a name. It did not say what the name admits, and the proxy admitted everything —
its request path was a host lookup and a forward. Measured against what the replaced ingress
actually relies on, four things were missing: **authentication**, **refusal scoped to a path**,
**path-scoped routing with priority**, and **redirect**. Three modules depend on the first, each to
gate an admin surface that has no login of its own; one dependent of the second is a live incident
mitigation.
**Policy belongs to the route, not beside it.** A contribution carries it along with the name, the
port and the location. The alternative — a proxy-side settings layer keyed by route name — keeps the
grant literally clean but makes *"what protects this route"* a question answered from two files that
nothing keeps in step. A route's protection is part of what a route is.
**The set is closed at those four.** A fifth is an amendment, so each addition is earned by a
dependent that exists rather than added because a middleware surface was open. An open surface would
recreate the thing being replaced, and is far harder to narrow later than a closed one is to widen.
**Where policy needs a credential, the declaration names a secret; it never carries one.** The mesh
already mints and holds credentials, and that machinery stays the only thing that does — so a hash
never reaches anything regenerated, synced or committed.
**This re-keys the table.** Two of the four need one host routed more than one way, so the proxy
matches on host **and path**, with priority, rather than mapping a host to a single target. Equal
priorities must resolve identically every time, or the proxy stops being reproducible.
The proxy remains a reference implementation: the contract is the file the mesh writes, not the
program that reads it, and another proxy may implement the same file.
## 4 — Filtering
**Derived from what is assigned here, and from the overlay's shape** — a node's open ports are a
@@ -734,6 +769,12 @@ where a found tunnel is left running beside the mesh's; where it is adopted ther
The guard admits the mesh's ports from that one interface, and the predecessor's peers arrive on
it.
*2026-09-27, [ADR 0127](../../02-DECISIONS/0119-a-taken-tunnels-predecessor-is-retired.md).* The
found configuration is kept only until the take is proven — the found unit down, the mesh's
interface up and handshaking with a peer. Then it is removed from where the found unit reads it
(its original stays kept), the hold ends, and the predecessor's tunnel cannot be raised again by
anything but a person restoring it by hand. Undeclaring the private network does not bring it back.
## The bus is NATS
*2026-09-23, [ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md). Architecture to be written
+3 -2
View File
@@ -5,8 +5,9 @@ code:
- mesh-controller cmd/mesh-builder
- mesh-controller internal/builder
- mesh-catalog modules/builder
updated: 2026-09-21
updated: 2026-09-25
decisions:
- 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md
- 02-DECISIONS/0097-a-vendor-image-is-a-declared-build-input.md
- 02-DECISIONS/0096-an-upstream-image-is-copied-between-registries.md
- 02-DECISIONS/0091-a-mount-is-declared-three-ways.md
@@ -38,7 +39,7 @@ controller's again. The builder's whole responsibility is the middle.
| term | is |
|---|---|
| **source** | a repository, a path within it, and a ref — resolved to one commit ([ADR 0069](../../02-DECISIONS/0069-a-module-is-a-repository-and-a-path.md)) |
| **source** | a repository, a path within it, and a ref — resolved to one commit ([ADR 0069](../../02-DECISIONS/0069-a-module-is-a-repository-and-a-path.md)). The repository is either on the forge holding the `git` seat, recorded by its path there and cloned from wherever that forge runs at build time, or external, recorded and cloned exactly as given ([ADR 0111](../../02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md)) |
| **recipe** | how *one* artifact is produced from that source |
| **toolchain** | what a recipe runs inside — a compiler, a runtime, the SDK |
| **artifact** | what a recipe produced, named by the digest of its content |
+104 -36
View File
@@ -3,11 +3,14 @@ layer: to-be
status: proposed
code:
- mesh-sdk src
- mesh-tools src/broker-amqp.ts
- mesh-tools src/broker-nats.ts (and broker-amqp.ts until the rollout)
- mesh-controller internal/link
updated: 2026-09-21
updated: 2026-09-26
decisions:
- 02-DECISIONS/0095-the-control-plane-is-the-way-to-ask-a-module.md
- 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0128-the-mesh-bus-is-required-not-ambient.md
- 02-DECISIONS/0116-the-bus-is-built-in-five-steps.md
- 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md
- 02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md
- 02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md
@@ -23,6 +26,18 @@ language and nothing more ([ADR 0074](../../02-DECISIONS/0074-the-wire-is-specif
This is a specification, so it says what is required rather than how anything is arranged. Where it
describes current behaviour that is *not yet* specified-and-conformed, it says so.
> **Rewritten onto NATS, 2026-09-26** (step 3 of
> [ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md)). What
> [ADR 0074](../../02-DECISIONS/0074-the-wire-is-specified-not-the-types.md) decided is untouched:
> a floor plus independent capabilities, an implementation legitimate when it claims less,
> identity from the sealed credential, at-least-once with dedup on `x-event-id`, and conformance
> as executable fixtures rather than prose. What changed is the transport beneath all of it —
> exchanges and queues became subjects and streams. The envelope keeps its shape
> ([ADR 0042](../../02-DECISIONS/0042-the-shape-of-an-event-on-the-wire.md)).
>
> Statements here marked *verified* were checked against a running server while the runtime's
> client was written, not reasoned from documentation.
## The shape of it
A **floor** every implementation needs, and three **capabilities** that are independent of each
@@ -47,22 +62,39 @@ document:
| field | is | required |
|---|---|---|
| `url` | an `amqps://` URL carrying the account's user and password | yes |
| `fingerprint` | sha256 of the certificate the broker must present | yes for a scoped account |
| `url` | a `tls://` URL for the bus, with the account's user and password | yes |
| `fingerprint` | sha256 of the certificate the bus must present | yes for a scoped account |
| `node` | the machine this account was issued for | yes for a scoped account |
| `module` | the module this account was issued for | yes for a scoped account |
A plain string rather than a document is a **bootstrap URL** — unscoped, for the moment before a
mesh can issue anything. An implementation accepts both and must not treat the second as ordinary.
**`node` and `module` are not decoration: every subject an implementation touches is derived from
them.** Its own namespace is `mesh.mod.<module>`, its consumer is `<node>_<module>`, its inbox is
its own. So a credential without them is refused rather than guessed at — an implementation that
fell back to an environment variable would let anything on the machine decide which module it is,
which is what the identity rule below exists to prevent.
The credential itself is fetched, never carried in a declaration: a declaration is persisted as
state and a sealed secret in a stream is an archive rather than a moment
([design 29](32-what-a-module-declares.md) §10).
### Connecting
- The connection **pins the fingerprint**. It does not trust a certificate authority, and it does
not skip verification. A broker presenting a different certificate is refused, whatever else is
not skip verification. A bus presenting a different certificate is refused, whatever else is
true of it.
- A scoped account **does not declare exchanges**. The foundation owns them; an account that may
declare one is an account that may create a parallel mesh by typo.
- An implementation **declares its own queue** and nothing else.
- **The certificate must also carry a name the bus is dialled by.** *Verified:* the NATS client
exposes no hook to replace hostname verification, so pinning no longer makes it redundant the
way it did on AMQP — the pin happens before dialling and the library's own name check happens
beside it. A certificate without a matching subject-alternative name is refused at connect, by
a library error rather than by anything the mesh says.
- An implementation **creates nothing on the bus**: not a stream, not a consumer, not a subject.
Streams and durable consumers are the controller's alone ([design 25](25-the-bus-on-nats.md)
§3), and a module's account cannot reach the JetStream API to make one. An implementation binds
the consumer the mesh created for it, and if it is absent that is a mesh that has not finished
assigning the module, not something for the module to fix.
### Identity
@@ -77,26 +109,49 @@ the credential disagree, the credential wins and the variable is overwritten.
## Capability: events
### The exchanges
### The subjects
| exchange | carries |
| subject | carries |
|---|---|
| `mesh.events` | every event |
| `mesh.events.dead` | what could not be handled |
| `mesh.mod.<module>.event.<key>` | an event that module emitted |
| `mesh.seat.<seat>.event.<verb>` | an event the holder of that role emitted |
### The queue
Both are captured by the `EVENTS` stream. **An event's source is enforced rather than claimed**: a
module's account may publish only into its own namespace, so `x-source` cannot disagree with where
the message arrived from.
One **durable** queue per consumer, named `<node>.<module>.events`, with as many bindings as the
module has patterns. Durable because an event emitted while a module is restarting is exactly the
one that must not be lost.
**The `event` token is load-bearing.** A module's namespace also carries its tool calls
(`mesh.mod.<module>.tool.<tool>`), and a stream is defined by a subject filter — without the token
the events stream would capture every tool invocation in the mesh, and a tool call must never be
persisted.
**A message matching two bindings is delivered once**, so an implementation must match the routing
key against its own patterns locally to decide which handlers run. An implementation that ran every
handler whose exchange binding matched would run the wrong one.
### The consumer
One **durable consumer** per module, named `<node>_<module>`, carrying one filter per pattern the
module consumes. Durable because an event emitted while a module is restarting is exactly the one
that must not be lost.
**Created by the controller, bound by the implementation.** A module declares what it reacts to
and never how delivery works, so it does not name its consumer, does not choose its ack policy or
delivery limit, and cannot misconfigure them.
*Verified, and it is a trap:* a durable name **may not contain a dot**, while the subject a
consumer acknowledges on is `$JS.ACK.<stream>.<consumer>.…` — two names joined by one. An
implementation that treats them as a single string reads correctly in a permission list and is
refused as a consumer name. Left wrong, the symptom is every message redelivered forever while
the permissions look right.
**One consumer may carry filters wider than one handler's pattern**, because a module subscribing
twice gets one consumer with both. So an implementation still matches the key against its own
patterns locally to decide which handlers run — and **acknowledges a message no handler wanted**,
or it is redelivered until it expires.
### The envelope
Headers ride as AMQP headers. The body is JSON.
Headers ride as **NATS headers**; the body is JSON, and the body alone. *Verified:* the payload is
the event's `body`, not the whole envelope re-encoded — an implementation that nested the envelope
would pass every one of its own tests and agree with no other, which is the exact failure the
conformance fixtures exist to catch. The key is recovered from the subject, not carried twice.
| header | is | required |
|---|---|---|
@@ -117,6 +172,11 @@ breaking change for everybody.
At-least-once. **Deduplication is on `x-event-id`**, which only the emitter can produce — a
consumer cannot tell a redelivery from a second event any other way.
On NATS the id does double duty: an implementation passes it as the publish's message id, so the
**server** also refuses a duplicate inside its window. That narrows the window in which a
consumer has to deduplicate; it does not remove the requirement, because the window is finite and
a redelivery after it is still a redelivery.
### What is true, checked (2026-09-16)
Go emits all five required headers; the SDK requires exactly those. `x-causation-id` and `x-schema`
@@ -131,22 +191,30 @@ version to declare.
A module's tools are its operator-facing surface.
- A tool is served from a **shared durable queue**, `serve.<key>`. Shared, so several runtimes
serving one tool compete for a call rather than each answering it.
- A call is request and reply. The reply returns through the RPC exchange `mesh.rpc`, keyed by the
caller's own reply queue — **not** through the default exchange, which would let a caller publish
into any queue on the broker.
- A caller needs a **reply queue**, and that is what a module's scoped account may not declare
([issue 049](../../04-ISSUES/049-a-module-can-serve-tools-and-nothing-can-call-them/00-report.md)).
So a module may serve tools and may not call them.
- **The control plane is the way to ask**
([ADR 0095](../../02-DECISIONS/0095-the-control-plane-is-the-way-to-ask-a-module.md)):
`ask <module> <tool> [json]` publishes on `mesh.rpc` under `<module>.<tool>` with a private reply
queue bound under its own name, and prints the answer as the module gave it. A module declares
nothing about being asked — serving a tool is being askable through the control plane. A
module-to-module call, if one is wanted, is a grant like any other and a later decision.
*How it is checked:* a tools-only bed asks a served tool through the control plane and asserts
an answer arrived, where a timeout would read differently.
- A tool is served on `mesh.mod.<module>.tool.<tool>`, with a **queue group** — so several
runtimes serving one tool compete for a call rather than each answering it.
- A call is request and reply on **core NATS, never a stream**. A tool call is not persisted: a
lost one is a timeout the caller already handles, and a stream of them would be the mesh's most
voluminous and least valuable traffic competing for retention with the messages that matter.
- The reply goes to the inbox the request carries. A responder may answer it because its account
is granted **`allow_responses`** — one reply to the subject of a message it actually received,
and nothing wider. That is what makes a per-account inbox prefix workable: no user is ever
granted `_INBOX.>`, so without it a responder could not reach the caller at all.
- **A module may now call a tool, which on AMQP it could not.** *Verified:* two modules on
separate connections, one serving and one calling, with an answer returned and a throwing
handler reaching the caller as an error rather than a timeout.
[Issue 049](../../04-ISSUES/049-a-module-can-serve-tools-and-nothing-can-call-them/00-report.md)
recorded the old limit — a scoped account could not declare the reply queue a caller needs —
and [ADR 0095](../../02-DECISIONS/0095-the-control-plane-is-the-way-to-ask-a-module.md) routed
every ask through the control plane because of it. **That constraint is gone**, and each
account's own inbox prefix replaces it.
ADR 0095 is not thereby reversed: the control plane remains *a* way to ask, and a person asking
a module should still go through it. What changes is that "a module-to-module call, if one is
wanted, is a later decision" is no longer a question about *capability*. It is a policy
question, and the answer the mesh already has is `uses`: a module declares the seat it calls,
and the permission follows the declaration.
- A module declares nothing about being asked — serving a tool is being askable.
---
+1
View File
@@ -12,6 +12,7 @@ decisions:
- 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md
- 02-DECISIONS/0075-two-stores-and-which-provides-what.md
- 02-DECISIONS/0014-no-npm-workspace.md
- 02-DECISIONS/0109-a-package-registry-seat-is-one-per-ecosystem.md
---
# The work ahead
+11 -3
View File
@@ -2,10 +2,11 @@
layer: to-be
status: designed
code: []
updated: 2026-09-20
updated: 2026-09-25
decisions:
- 02-DECISIONS/0084-which-provider-serves-a-consumer.md
- 02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md
- 02-DECISIONS/0126-a-module-declares-its-own-seats.md
---
# 23 — Choosing a provider
@@ -50,9 +51,16 @@ provider on a different node. That coupling is exactly what may not be guessed,
names the provider. Naming it is also what makes a later move safe — the mesh knows the binding is
to that provider and not to whichever one is nearest.
**A seat names the mesh's one provider of a kind.** Where a seat delivers the provision, its holder
answers for it when several providers exist and the consumer named none. That is not picking: the
choice was made once, mesh-wide, by assigning the holder, rather than once per consumer by naming it
([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md) (superseding [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md)),
[26 — The seats](26-the-seats.md)). A named provider still wins over the seat, because a consumer
coupled to particular contents has said so.
**Ambiguity is refused, never resolved by picking.** If several providers of a kind exist, none is
named, and none is co-located, the requirement is unsatisfiable and is refused with the candidates
shown — the same stance
named, none is co-located, and no seat delivers it, the requirement is unsatisfiable and is refused
with the candidates shown — the same stance
[ADR 0027](../../02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md) took
against a confidently-wrong match, applied to the instance rather than the dialect. A wrong answer
delivered quietly costs more than a refusal.
+553
View File
@@ -0,0 +1,553 @@
---
layer: to-be
status: in-progress
code:
- mesh-controller internal/link (to be replaced)
- mesh-host internal/link (to be replaced)
- mesh-tools src/broker-amqp.ts (to be replaced)
- mesh-catalog modules/nats (to be written)
- mesh-sdk src (the protocol's NATS binding, step 3)
updated: 2026-09-27
decisions:
- 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
- 02-DECISIONS/0116-the-bus-is-built-in-five-steps.md
- 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md
- 02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md
- 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
- 02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md
- 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
- 02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md
- 02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md
- 02-DECISIONS/0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md
---
# 25. The bus on NATS
**Status: proposed — a design to be reviewed before any code.** This is the architecture
[ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md) asks for. It says what rides the bus,
under which subject, with which guarantee, under whose account; how a node joins; how a person
reaches a tool; how the mesh moves from the bus it has to this one; and how each claim is checked.
Prose and diagrams only; no configuration is pasted.
## 1. What the bus is for
**The bus is where the mesh happens.** Not a transport the mesh sends things over — the place a
module is reachable at all, where a role is addressed without knowing who holds it, where the
mesh's own state lives, and where what a module may say is decided by what it declared.
| Traffic | Shape | Guarantee it needs |
|---|---|---|
| **control** — a node's report, a build's outcome, an enrolment | job | nothing lost while the store restarts; retried; in order per node |
| **heartbeat** — a node saying it is alive | fire and forget | none; a lost one is the next one |
| **declarations** — the controller tells a node what to be | state | the node gets the newest; a stale one is never applied |
| **builds** — work for the build machine | job | at least once, one worker at a time |
| **events** — a module says something happened | 1:many | delivered to every consumer that declared it; dead-lettered when it cannot be |
| **tools** — a module or a person asks another's tool | request/reply | one answer, from one server, or a timeout |
| **work to a role** — a module submits to a capability without knowing who provides it | job | exactly one holder does it; it queues while nobody does |
The last two rows are the ones worth dwelling on, because they are not messaging in the sense of
carrying bytes from A to B. **A role is addressable**, so a caller names the capability and never
the module or the node — and the implementation can be replaced under it without a caller
changing ([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md)). That is a
property of the mesh's architecture that happens to be expressed in subjects.
And more of the mesh lands here as it is built: conditions and observed state in key-value
buckets that anything may watch, the server's own advisories becoming observations like any other
([research 017](../../01-RESEARCH/017-a-mesh-that-heals-itself/00-overview.md)), and a person's
client speaking the bus directly rather than through a surface built over it (§7). None of that
is a message being moved; all of it is the bus being the mesh's centre.
**What a module sees of it is small and derived.** It declares what it emits, consumes, serves
and uses, and the subjects, streams, consumers and permissions all follow from that
([design 29](32-what-a-module-declares.md)). The sdk's contract — `request`, `handle`, `publish`,
`subscribe`, `close` — is the whole surface, and it does not change
([ADR 0039](../../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md)).
## 2. Subjects
NATS addresses everything by subject. The mesh's subject space is one tree, and every account's
permissions are expressed as which branches of it that account may publish to and subscribe from.
```
mesh.control.<node>.report a node's report (JetStream: CONTROL)
mesh.control.<node>.alive heartbeat (core, no persistence)
mesh.control.enrol an enrolment request (JetStream: CONTROL)
mesh.node.<node>.declare a declaration for a node (JetStream: NODES, last-per-subject)
mesh.mod.<module>.event.<event> an event (JetStream: EVENTS)
mesh.mod.<module>.tool.<tool> a tool invocation (core request/reply)
mesh.seat.<seat>.accept.<verb> work submitted to a role (JetStream: per-seat work queue)
mesh.seat.<seat>.event.<verb> a role's own event (JetStream: EVENTS)
mesh.seat.<seat>.tool.<verb> a role's tool (core request/reply)
mesh.ask.<node>.<command> the controller's command api (core request/reply)
```
**Revised 2026-09-27** ([ADR 0129](../../02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md)):
**`mesh.build.request`, `mesh.control.built` and the BUILDS stream are gone.** A build is work submitted to a role, and the
mesh already has a shape for that — a seat's `accept` subjects, on a work queue with a queue group of
holders, which is what a build queue shared by several machines *is*. Keeping a second mechanism for it
meant two things to reason about and two places for a permission to be wrong. A build's outcome is the
seat's own event — which is why the control branch loses its copy too: one publish reaches whoever
asked, the controller that records it and the catalogue that places it — the fan-out a shared exchange gave for free, as a derived subject rather than a
configured topology.
**Revised 2026-09-26** ([design 29](32-what-a-module-declares.md)): a module's events and tools
moved from `mesh.events.*` / `mesh.tools.*` into one namespace per module, `mesh.mod.<module>.>`,
so a module's authority over its own name is a single subject pattern the server enforces — and
each carries a **kind token**, without which an events stream's filter would capture tool calls.
Seats are the same shape, one namespace per role.
Two things this buys over the exchanges: **request/reply is native** — a tool call is one
`request` on `mesh.tools.<module>.<tool>` answered by whichever runtime serves it (a queue group per
tool, so several nodes may serve one tool); and **a declaration is last-per-subject** — the NODES
stream keeps only the newest message on `mesh.node.<node>.declare`, so a node that was away gets
exactly the current declaration and nothing older. That is the wire-level answer to
[issue 107](../../04-ISSUES/107-a-declaration-carries-no-order/00-report.md): the stream's sequence
*is* the order, and a node that sees sequence n refuses n−1 by construction.
**A reply-to travelling through a JetStream stream is carried in the payload, never in the
transport `Reply` field.** *Verified against a running server, 2026-09-27*: a caller published
asking for a reply to `_INBOX.LCr3M83q…`, and the consumer saw a `Reply` field of
`$JS.ACK.PROBE.probe_consumer.1.1.1…`. The address is replaced, not merely at risk — so the
payload-borne reply subject below is necessary rather than defensive, and the check is a test
rather than a note, because a future server that stopped doing this would leave enrolment
working and the reason for the field quietly becoming folklore.
Revision, first review: core NATS request/reply sets the requester's
ephemeral inbox as the message's `Reply` field, and a plain responder answers it directly — but a
message a JetStream consumer delivers has already had that field claimed for the consumer's own
ack address (`$JS.ACK.<stream>.<consumer>...`), so by the time the controller (§3's CONTROL
consumer) sees the message, `Reply` names where *it* must ack, not where the original caller is
waiting. `mesh.control.enrol` is the case that matters: a synchronous-feeling caller waiting on an
ephemeral inbox, over a subject the store-window guarantee may legitimately delay by several
`nak` cycles — exactly the combination that would otherwise deliver the answer to a caller who
has long since timed out and unsubscribed. So every CONTROL message that expects an answer states
its reply subject as an ordinary field of its own payload; the controller reads it from there and
publishes the answer to it explicitly, never via `Respond()`. Nothing else in this design routes
a reply through a stream — tools and heartbeats stay on core NATS, where `Reply` means what it has
always meant.
## 3. Streams, and the guarantees they carry
Core NATS is at-most-once. Everything the mesh must not lose lives in a JetStream stream:
| Stream | Subjects | Retention | Why |
|---|---|---|---|
| CONTROL | `mesh.control.>` except `alive` (a build's outcome moved to its seat, ADR 0121) | work queue, one consumer (the controller), explicit ack | the store-window guarantee ([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)): the controller `nak`s with a delay while its store is away and the message is redelivered; nothing is dropped |
| NODES | `mesh.node.>` | last per subject | one declaration per node, always the newest |
| EVENTS | `mesh.mod.*.event.>` | limits (age, size), durable consumer per subscribing module | a subscriber that was down catches up; after `max-deliver` attempts the advisory feeds `mesh.events.dead` (its own small stream) |
Tool calls and heartbeats stay on core NATS: a lost heartbeat is the next heartbeat; a lost tool
call is a timeout the caller already handles.
Streams and consumers are objects the controller creates at genesis and asserts on start; a module
declares nothing about them. The controller is the only writer of stream definitions.
### The store window, and what moving it into the server changes
The guarantee ([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)) is
that a push the controller cannot record because its store is restarting is **held and retried** —
never dropped, never falsely acknowledged. Here that is a `nak` with a delay: the server holds
the message and redelivers it, so the controller keeps no list of parked messages and one that
restarts mid-window loses nothing it was holding.
**That is a plain win, and it introduces one problem worth naming.** Holding a delivery in memory
let the controller drop an older report when a newer one for the same node arrived, because
acting on the older after the newer would undo the newer. A `nak`ed message belongs to the server
and comes back whatever happened meanwhile — so the older report is redelivered *after* the newer
was applied.
The answer was already in the message. A report carries the **digest of the declaration it is
about**, which exists because an earlier attempt to order reports by time lost the race it
invited: an apply that began under the previous declaration finishes after the next is sent, and
its report reads as newer than the send. Clocks cannot answer *which*.
So supersession stops being something the controller remembers and becomes something it checks —
a report whose digest is not the one outstanding for that node is acknowledged without being
acted on. The same shape as a node refusing a superseded declaration by sequence
([issue 107](../../04-ISSUES/107-a-declaration-carries-no-order/00-report.md)): **ordering settled
by what a message says, not by when it arrived.** And staleness is checked before the store is
waited on, so a redelivery that lost its race does not hold a slot in the window that a current
message needs.
## 4. Accounts
[ADR 0043](../../02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) says a
module's account may publish only what it `emits` and consume only what it `consumes`. NATS
expresses this exactly, per subject, and better than a vhost could:
- **One NATS account for the mesh.** Accounts in NATS isolate subject spaces entirely; the mesh
is one space, so it is one account. The predecessor's compatibility broker is not on this bus at
all.
**One account is a choice with a cost, stated plainly on revision:** none of NATS's own
isolation is free here, because there is only the one subject space, for everyone. Two
consequences that a single-account mesh must therefore grant on purpose, not by omission:
- **A durable consumer needs permission to ack, or it never really consumes.** Acking a
JetStream delivery is a publish to that consumer's own ack-reply address
(`$JS.ACK.<stream>.<consumer>.>`), a different subject from anything the consumer subscribes.
A module's user is therefore granted publish on `$JS.ACK.EVENTS.<module>.>` as well as its
emits — scoped to the one consumer name the controller derives for that module, so a module
can ack only its own deliveries. Without this, first review found, every message it receives
would be redelivered forever: refused by the permission list it already has.
- **A reply inbox needs a subject nothing else can guess or enumerate.** With one account,
inbox privacy is the permission list or it is nothing — there is no second account backing
it up. So no user is ever granted a bare `_INBOX.>`. Each user's inbox subject is derived
from its own identity (`_INBOX.<module>.<node>.>`, or `_INBOX.person.<name>.>`), and its
permissions name only that one prefix, for the reply to any request it makes and nothing
wider. First review found the account note without this and read it as "any user may
subscribe any inbox" — which was accurate against the text as it stood.
**And a scoped inbox needs `allow_responses`, or nothing can answer.** Revision, found while
composing the first real configuration: the rule above scopes each user's inbox to itself,
which is right — and leaves a responder unable to reply, because the answer goes to the
*caller's* inbox, which the responder has no permission for. The two ways out are granting
every responder `_INBOX.>`, which is exactly the blanket grant this bullet refuses, or NATS's
own `allow_responses`: the server permits one reply to the reply-subject of a message the user
actually received, within a TTL, and nothing else. So authority to answer is bounded by having
been asked, and only principals that serve something are granted it — a pure consumer gets
nothing. Without this the scoping is not merely incomplete: every tool call in the mesh times
out, and the permission list looks correct while it happens.
- **One user per module per node**, as today, with publish permissions
`mesh.events.<module>.<event>` for each emit, `mesh.tools.<module>.>` to serve its tools, its
own ack-reply subject for each durable consumer it holds, and its own inbox prefix; subscribe
permissions for each consumed event's subject, its tool subjects, and that same inbox prefix.
Nothing else. A module that tries to publish outside its emits is refused by the server, not by
convention.
- **The controller's user** owns `mesh.control.>`, `mesh.node.>` and the streams, and may submit work
to the seats the mesh's own flows use — a build, for one (ADR 0121).
**A host's user** may publish its own `mesh.control.<node>.>` and subscribe its own
`mesh.node.<node>.declare` — and nothing of any other node's.
- **A person's user** (§7) is a module-shaped user with permissions on the tool subjects it may
invoke, issued and revoked by the controller like any account.
**The server does not verify client certificates, and TLS is still required.** *Revision,
2026-09-27, found by building the module's image and connecting to it as a host would.* The first
composed configuration said `verify: true`, which makes the server demand a **client** certificate —
and nothing in the mesh presents one. A host pins this server's exact certificate and authenticates
with the password the mesh minted ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)),
and a module's runtime does the same. With it on, every connection in the mesh dies at the TLS
handshake before any password is looked at, and the error — "client didn't provide a certificate" —
reads as a fault in the client rather than in the bus's configuration. The `tls` block is what makes
TLS required; `verify` only decides whether client certificates are checked. What is given up is a
second factor the mesh has no machinery to issue or rotate — a certificate per module per node — and
what is kept is stronger than a name check in both directions: an exact pin outward, a password
scoped per user inward. **Mutual TLS is a later question and would need that machinery first.**
**The mesh composes the accounts; the module composes its server.** *Revision, 2026-09-27, while
building the composition.* An earlier reading of the paragraph below had the controller writing the
whole file. It writes only the user list. A server's ports, its TLS paths and its store directory
are properties of the container the module raises — they live in its image and its mounts and change
when it does — so the module declares its own configuration and `include`s the mesh's half. A
controller that wrote the whole file would have to be kept in step with a Dockerfile it never sees,
and a module could not change its own image without the mesh agreeing. Asking for the user list is
not enough to receive it: the file holds every user's password hash, so the claim on `mesh-broker` is
what authorises it. And the two files share one directory of necessity — an absolute include path is
resolved relative to the including file's own directory, so a server given one from elsewhere looks
for it underneath that directory and refuses to start.
**Accounts are configuration, not API calls.** The controller composes the mesh's user list and its
permissions into a file the host declares. **How that file reaches the running server is §5's,
not this one's** — revision, first review: an earlier draft said "reloads" and cited a precedent
that does not apply to a container (see §5). No management API, no credential travelling through a
management call, and the [issue 102](../../04-ISSUES/102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md)
discipline from the first day: an address or a permission is read where it is used, never stored
with a port. Passwords are minted and sealed exactly as today; the file holds bcrypt hashes.
Alternative considered and not taken: the operator/JWT model (`nsc`), where accounts are signed
tokens resolved by the server. It is the right model for a multi-tenant NATS; the mesh is one
tenant, already has a sealing key and a controller that writes files, and would gain a second
signing hierarchy for nothing.
## 5. The broker as a module
`nats` is a catalogue module claiming the seat `mesh-broker`
([ADR 0079](../../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md): the seat
is the server, and the server changes). It declares one container (a single binary; **JetStream on a host
directory bind, not a named volume** — revision, second review:
[issue 115](../../04-ISSUES/115-a-named-docker-volume-is-invisible-and-one-flag-from-gone/00-report.md)
is resolved, and converted the store, the broker and two others away from named volumes for the
reason it names; the bus's own data is not the place to reintroduce one), its listening ports —
**the client port, which carries TLS itself** rather than standing beside a plaintext one as the
AMQP broker's 5671/5672 pair did, and the monitoring endpoint on loopback — and a configuration
file the controller composes (accounts, permissions, TLS, JetStream).
**How that file's changes reach the running server, corrected on revision.** First review: the
earlier draft named `reload-on` as the mechanism, citing the container runtime's own trust file as
precedent. `reload-on` is real, but it is a **service** field
([mesh-host declaration.go](https://git.novox.be/novox/mesh-host), `Service.ReloadOn` —
`docker.service` is reloaded via systemd, which is what the cited precedent actually does). A
**container** resource has no reload field at all — only `restart-on`, and a container's
`restart-on` is documented, exactly, to mean *recreate*. Declared as the earlier draft had it,
either the field is silently meaningless on a container resource or — if read as the nearest real
equivalent — every account, permission, or key change recreates the bus's own server: every
connection dropped, every in-flight JetStream ack lost, mid-flight the moment a module is added,
reassigned, or a person's access changes. For the one resource everything else depends on, that
is not an edge case; it is the common case.
**The fix asks nothing new of the host.** `nats-server` already reloads its own configuration
live on `SIGHUP` — accounts, permissions, everything in §4 — without dropping a connection; this
is the server's own documented capability, not something built for the mesh. So the composed
configuration file is mounted into a **directory** resource, not directly — a directory's contents
are not compared for change the way [issue 103](../../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md)'s
fix made a directly-mounted file's content, so a rewritten file inside it is not, on its own, a
reason to recreate the container. The image's own entrypoint watches that one file and sends
`nats-server` its own process `SIGHUP` when it changes — self-contained, inside the module, the
same place `modules/gitea/token.ts` keeps its own state rather than asking the host to model it.
The host's only job is what it already does for any directory resource: keep the file's content
current. Nothing is declared as `reload-on` or `restart-on` for this resource at all.
Its guard is the same rule as the deprecated broker's: the monitoring port is refused from anything but
the private network. It is raised at genesis like the store, adopted as a module in the same
phase.
**The deprecated broker is an ordinary module, not a compatibility layer.** Revision, second review
([ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md): AMQP is not a provision, and everything speaks to the `mesh-broker` seat)): earlier text here, and
ADR 0106 before it, called it `lavinmq-compat` — one purpose, the predecessor's clients, and a
retirement condition of no client connected for a period the operator sets. It is none of those.
A module may legitimately need an AMQP broker as a **backing service**, the way it needs a
database, and the provider that answers that is an ordinary module like any other: no seat, not
foundation, never raised at genesis, installed when something wants it and absent from a mesh
that does not. There is no retirement condition, because the day its last client disappears is
not a day anything is waiting for.
**Revised 2026-09-27** ([ADR 0130](../../02-DECISIONS/0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md)):
**that day is coming.** The predecessor is deprecated — some of it still running, none of it being
migrated, left to stop rather than moved — so the broker retires once nothing requires `amqp`. Still no
retirement *condition* and no end-date machinery: a provision with no consumers has its provider
unassigned, which is the ordinary mechanism and is ADR 0127 being paid off rather than revised. What
also goes with it is the tooling that reaches this installation's machines remotely, because the
predecessor's own mesh talks over that broker — so the rollout is driven from the node, or before the
broker stops.
What is deprecated is AMQP as **the mesh's transport**, which is this whole document. The rule
that remains is about direction rather than software: *inter-module communication goes over the
bus.* A module may hold a broker, a database or a cache for itself; it may not use one as a
channel to another module.
## 6. Joining: the enrolment handshake
Unchanged in shape, changed in transport. A node that has a token connects to the bus over TLS
with the **enrolment user** — a user that may publish `mesh.control.enrol` and subscribe its own
`_INBOX.enrol.<token-id>.>` and nothing else — publishes its request (the claim of the token, its
keys, its proof, its own reply subject as §2 now requires, and the found tunnel from
[ADR 0105](../../02-DECISIONS/0105-the-mesh-adopts-the-predecessors-tunnel-in-place.md)), and waits
on that inbox. The controller spends the token, records the node, composes the node's own user
into the server's configuration, and — reading the reply subject from the request's payload, never
from the transport `Reply` field the CONTROL consumer has already claimed for its own ack —
answers with the credentials sealed to the node's sealing key. The node reconnects as itself. The
enrolment user's permissions are what make a leaked token useless for anything but enrolling: it
cannot read a declaration or hear an event, and it cannot subscribe any inbox but the one its own
token derives.
## 7. A person's client
The operator asked for the mesh's tools from a workstation, and for it designed here rather than
bridged. It is three things:
1. **A person's account**: `operator issue <name>` on the controller creates a user whose
permissions are the tool subjects it may invoke — `mesh.tools.>` for an administrator, a list
for anyone else — and nothing on control, nodes or builds. It is issued, sealed to the person's
own key, and revoked, like a module's.
2. **A client that speaks the bus**: a small program on the workstation that connects as that user
over TLS, lists tools by asking the catalogue (`mesh.tools.mesh-catalog.catalog_tools`), and
turns each tool into a call — as an MCP server for an agent, and as a command line for a person.
It uses the sdk's `Broker` contract on the NATS runtime, so it is the same code path a module's
tools use, not a second protocol.
3. **Reachability**: the workstation reaches the bus over the private network once it is a node,
or over the predecessor's tunnel before that, on the bus's port; the guard and the openings
treat the bus as they do today.
Nothing is built of this before §10's bed passes; the MCP surface is a thin adapter over (2).
## 8. What a module sees, and what the wire does
**The contract a module is written against does not change.** `publish` on an envelope becomes a
publish on `mesh.events.<module>.<event>`; `subscribe` with a pattern becomes a durable JetStream
consumer on the matching subject filter; `request`/`handle` become a NATS request and a queue-group
subscription on `mesh.tools.<module>.<tool>`. The envelope's shape
([ADR 0042](../../02-DECISIONS/0042-the-shape-of-an-event-on-the-wire.md)) is unchanged; it is the
message body. A module built today runs on the new runtime without a rebuild — that is the test of
ADR 0039, and it is in §10.
**The wire underneath it changes completely, and that is a specification, not an implementation
detail.** Revision, second review ([ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md)):
an earlier draft of this section said "nothing new" and stopped there, which read as though the
change were contained inside the runtime. It is not.
[ADR 0074](../../02-DECISIONS/0074-the-wire-is-specified-not-the-types.md) settled that an SDK is an
implementation of a *specified wire*, checked by fixtures that must match byte for byte — because
two implementations that disagree about an envelope do not fail to compile, they ignore each other
while both keep running. That specification is
[design 19](19-the-module-protocol.md), and it is written in exchanges, a durable per-consumer queue
named `<node>.<module>.events`, and a shared `serve.<key>` queue. Every one of those is gone here.
So design 19 is rewritten from exchanges and queues to the subjects and streams of §2 and §3, its
fixtures are recaptured on NATS, and each SDK re-claims the capabilities it passes. That is step 3
of §9, and until it lands design 19 says in its own opening that its wire section describes the bus
being replaced. What is *not* rewritten: ADR 0074's model — protocol split per capability, an SDK
that implements the floor and events alone is legitimate, conformance executable per capability —
and ADR 0039's refusal. A new transport is when the pressure to grow the shared library is highest;
nothing is added to it here.
## 9. Moving from the bus the mesh has: five steps
Per ADR 0106 the bus moves once. Per
[ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md) the *build* is five steps,
each ending at something §10 proves, so that no part of this waits on the whole of it. Dividing the
build does not divide the bus: steps 1 to 4 leave every node on AMQP, and step 5 is still one
rollout.
**Step 1 — genesis raises the broker.** The `nats` module of §5, and a mesh raised on it from
nothing. This is built and proven although the mesh it is for will never travel this path: genesis
is where the foundation is *defined*, the one place the mesh comes from nothing, and the definition
every other path is measured against. A genesis path that exists only on paper is one nobody finds
wrong until there is a second mesh. *Ends at: the genesis bed.*
**Step 2 — adoption puts the broker in its seat.** A mesh already running does not get a foundation
module by being raised again; it adopts one in place
([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)). The
server is raised beside the deprecated broker on its own ports, carrying no mesh traffic yet, and the
`nats` module is adopted onto it. The seat it claims is **`mesh-broker`**, unchanged — the
foundation seats are named after the server's role rather than the product
([ADR 0079](../../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md)) for
exactly this case, and a seat named after the product would need renaming by every change the seat
exists to survive. *Ends at: the adoption bed.*
**Step 3 — the protocol gets its NATS binding.** §8's other half: design 19 rewritten from
exchanges and queues to subjects and streams, its conformance fixtures recaptured on NATS, each SDK
re-claiming the capabilities it passes. ADR 0074's model is unamended and ADR 0039's refusal holds —
no helper layer arrives with the new transport. A language may still arrive in pieces: connection
and events first, tools and provisioning when something needs them. *Ends at: the conformance suite,
per capability, per implementation.*
**Step 4 — the core speaks NATS.** The controller's link, the host's link, the tool runtime's
client — and with them the flows carried today by something other than the bus: a build source's
change reaching the builder, an installation, a module's own reports. Each is a conversion with a
named before and after. Observation — heartbeats, conditions, key-value state — belongs to
[research 017](../../01-RESEARCH/017-a-mesh-that-heals-itself/00-overview.md), which already
reserves it for after the move; this step does not pre-empt its design. Modules converted meanwhile
target the sdk contract and are untouched by any of it. *Ends at: each converted flow proved against
the behaviour it replaced.*
**Step 5 — the rollout.** Unchanged from ADR 0106 and previewed: the controller composes every
node's and module's account into the server standing since step 2, then rolls out the controller,
every host and every runtime built for NATS. Each node's host connects to the new bus as it comes up
and reports; the controller confirms every node heard before it stops listening on AMQP. The
predecessor's clients never notice — their broker is the compatibility module and stays. Then the
AMQP-side mesh accounts are removed from it, leaving only the predecessor's users, and it retires
when its retirement condition holds. The bus's port settings follow ADR 0100 like any port.
*Ends at: the cutover bed, then the rollout itself.*
What is not done, at any step: no dual-bus period for the mesh's own traffic, no bridge, no module
rebuilt.
**The work of these five, broken down and measured, is [design 28](28-building-the-bus.md)** —
including the two places their dependencies put a bed later than the step that names it.
## 10. How it is checked
**A bed per step, and each is green before the step after it starts** — the division in §9 is only
real if the proofs divide with it. A step that cannot name what its bed proves is not a step, and is
divided further before it is started.
**Step 1 — the genesis-broker bed**, a mesh raised from nothing, the server standing on it. The
mesh does not yet *live* on this bus — nothing speaks it until step 3's implementations exist — so
what this bed proves is the server, its configuration and the permissions, each of which the server
itself enforces and a plain client can therefore check:
- the four streams exist, asserted idempotently on a second start, with the retention of §3;
- every account and permission in the composed file is derived from the manifests' `emits` and
`consumes` and nothing else, with each user's own ack subject and its own inbox prefix;
- a user cannot publish outside its `emits` nor subscribe outside its `consumes` — refused by the
server, not by convention;
- a user cannot ack another user's delivery, and cannot subscribe another's inbox prefix — refused
by the server, not by the client's own good behaviour;
- the monitoring port is refused from anything but the private network;
- the `nats` container is not recreated when only its composed configuration file changes, and a
change to that file is live (a new user can connect, a revoked one cannot) within one
watcher-poll interval, without a restart.
**Step 2 — the adoption bed**, a mesh already running that has never had this server:
- the server is raised beside the deprecated broker on its own ports and the `nats` module is adopted
onto it in place, holding the data and the configuration it was raised with;
- the seat it claims is `mesh-broker`, and a second assignment of it anywhere in the mesh is
refused at resolution — *one per mesh*, as ADR 0079 requires;
- every node stays on AMQP throughout and nothing routes to the adopted server. This is the
check that makes steps 3 and 4 safe to run against a live mesh: adoption that quietly carried
traffic would be step 5 arriving early and unrehearsed.
**Step 3 — conformance, per capability, per implementation:**
- the fixtures of design 19, recaptured on NATS, produced and consumed byte for byte by every SDK
that claims the capability — the test ADR 0074 set, and the only one that catches two
implementations quietly ignoring each other;
- an SDK implementing connection and events alone passes those two and claims nothing more,
rather than failing as a whole;
- a module built before this design serves its tools unchanged on the new runtime — the test of
ADR 0039;
- the shared library gained nothing but the binding: its surface is the protocol and the
primitives, and a helper that arrived with the transport is a review failure, not a detail.
**Step 4 — the mesh living on it.** The implementations exist from step 3, so this is where a mesh
can first be raised on NATS and run:
- a node enrols over TLS with a claimed token, and the enrolment user cannot read a declaration;
- a push composes; the store is stopped; the push is held (nak with delay), the store returns, the
push applies, nothing was lost or duplicated;
- a node that was away gets exactly the newest declaration, and a replayed older one is refused
by sequence;
- an upgrade rolls out to two nodes;
- an enrolment request held by a `nak`-with-delay cycle still reaches the enrolling node's inbox
once the controller answers — proving the reply travels in the payload and not the transport
field a consumer's ack has already claimed;
- an event whose consumer keeps failing dead-letters after `max-deliver`;
- a module's tool is invoked from another node and from a person's client, each with an account
that can invoke it, and refused from one that cannot;
- a build source's change reaches the builder over the bus, and the build that follows is the one
the change asked for;
- an installation completes over the bus, with the same outcome the path it replaces produced;
- a node that was unreachable catches up on its reports rather than losing them.
**Step 5 — the cutover bed:** a mesh on AMQP with a predecessor stand-in on the compatibility
broker moves its bus in one rollout; every node reports on NATS afterwards; the stand-in's client
on AMQP is still connected throughout.
Unit tests hold the controller to composing accounts from `emits`/`consumes` and nothing else, to
creating the four streams and asserting them idempotently, and to spending a token exactly once;
the host to connecting as the enrolment user with nothing but enrolment permissions; the runtime
to mapping the sdk contract onto subjects exactly as §8 says.
## 11. Open, for the review
**Closed by the second revision** (2026-09-26, [ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md)):
the build was one undivided item (§9 is now five steps, each with its own bed in §10); there was no
adoption path for a mesh already running (§9 step 2); and §8 said a module sees "nothing new"
without distinguishing the sdk's contract, which does not change, from the specified wire, which
changes entirely and is design 19's (§8, §9 step 3). Two smaller corrections: the seat is
`mesh-broker` and not the product's name, and the shared library gains no conveniences with the new
transport.
**Closed by the first revision** (recorded in `MIGRATION-LOG.md`, 2026-09-24): the
`reload-on`/container mismatch (§5), the eaten reply subject on a CONTROL-stream message (§2, §6),
the missing ack permission (§4), and the un-scoped reply inbox under one account (§4). Each is
named where it was wrong, not silently fixed, so a reader comparing against the first version can
find what changed and why.
**Still open:**
- ~~Whether EVENTS should be one stream or one per emitting module.~~ **Closed**
([ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md): AMQP is not a provision, and everything speaks to the `mesh-broker` seat) (superseding [ADR 0125](../../02-DECISIONS/0125-the-bus-is-the-only-broker.md))): one stream, and not as a
preference — the streams the bus is made of are composed as configuration before any module
runs, because a provisioner is itself a module that needs a bus account to start. Bootstrapping
decides it.
- The heartbeat interval and the controller's "quiet" threshold on core NATS without persistence —
the same numbers as today are proposed.
- Whether the person's client is a catalogue module (runs on an enrolled workstation node) or a
standalone program (runs anywhere with credentials). Both, in that order, is proposed.
- Leaf nodes: a NATS leaf per machine would make every module's connection local and survive the
hub's restart. Deliberately out of scope; noted so it is not forgotten.
- **New, from this revision:** the `nats` image's own entrypoint now carries logic (watch a file,
signal a process) that no other module's container needed before. Is a one-file-watcher-and-
`SIGHUP` helper common enough across future modules with the same shape (a service that reloads
on `SIGHUP` but runs in a container) to belong in `mesh-sdk` rather than written once per module
that needs it? [ADR 0039](../../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md)'s test —
*does editing it recompile unrelated modules, and does it change often* — probably says no for
one instance; worth asking again if a second module needs the same shape.
+245
View File
@@ -0,0 +1,245 @@
---
layer: to-be
status: implemented
code:
- mesh-controller internal/catalogue/seats.go
- mesh-controller internal/catalogue/resolve.go
- mesh-controller internal/inventory/seats.go
- mesh-controller internal/inventory/migrations/0039-a-seat-is-held-by-one-assignment-on-record.sql
- mesh-controller cmd/mesh-controller/seats.go
- mesh-controller cmd/mesh-controller/source.go
- mesh-controller internal/inventory/migrations/0032-a-source-may-live-on-a-seat.sql
- mesh-catalog modules/gitea/module.json
updated: 2026-09-27
decisions:
- 02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
- 02-DECISIONS/0126-a-module-declares-its-own-seats.md
- 02-DECISIONS/0128-the-mesh-bus-is-required-not-ambient.md
- 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md
- 02-DECISIONS/0109-a-package-registry-seat-is-one-per-ecosystem.md
- 02-DECISIONS/0117-a-machines-uplink-is-a-seat.md
---
# 26 — The seats
**What a mesh can have one of, and who fills each.** A seat is a named role at a scope, held by one
module assignment. The mesh defines which seats exist. Holding one may deliver a provision, and the
list of seats with their holders is the quickest answer to "what is in this mesh".
## What a seat is
A seat has four properties, fixed by the mesh rather than by any module:
| property | is |
|---|---|
| name | what a definition names and an assignment holds, and what a person reads in the list |
| scope | node, site or mesh: where its capacity applies. Every seat in the set has a capacity of one, so one holder per scope. A bench, a seat with several holders, is a word the glossary keeps and no seat uses yet |
| delivers | the provision its holder answers for, or nothing |
| decision | the record that made it a seat |
**A definition says which seats a module can hold. An assignment says which it does hold.** The store
module can hold `mesh-store`, and it may run on every node whose capabilities match. Exactly one of
those assignments holds the seat, because that assignment says so, and a second assignment saying so
is refused. A seat makes a role singular, never a module.
**The seat points at the assignment.** Everything the mesh knows about the holder is what it knows
about that assignment: the node, the node's settings for the module, and what the module serves.
**Which assignment holds a seat is a fact on record, and changes as one act.** Revision, 2026-09-27
([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)). Until
then the holder was derived — the module that is assigned and claims the seat holds it, and a second
eligible assignment was refused. That has no way to pass a seat from one holder to the next without a
moment in which nobody holds it, and the controller finds its own bus through one of these seats: that
moment took the control plane down for an evening. So the holder is now one row the controller keeps,
written by a handover — `seat <name> --to <node>/<module>` — that names the seat and the assignment
taking it over and replaces the previous holder in the same write. Between two handovers the seat has
exactly one holder, and it is never none.
Three consequences follow. **A seat with no row is held as it always was**: the sole eligible
assignment holds it, and two eligible ones are refused — so a mesh that has never handed a seat over
behaves exactly as before, and the row appears the first time somebody does. **With a row, any other
assignment whose module could hold the seat is eligible and silent**: neither refused nor holding.
That is what lets the next holder run beside the current one until the handover, which the bus's move
needs ([28](28-building-the-bus.md), task 5.3). **And a holding is the assignment's**: unassigning the
holder takes the row with it, so a seat never points at something that is not running anywhere, and
the seat falls back to derivation rather than to nothing.
The handover refuses what would make the new holder wrong before anything is written: the seat must
exist, the assignment must exist, and the module must be able to hold the seat — claim it at its scope
and provide what it delivers, judged against the store's row and not against anything compiled into a
binary. It does not check that the module is running yet; `push` confirms that afterwards, and a
handover that could only be recorded after the new holder was up could not be the switch.
**The set is closed.** A seat the mesh does not define is refused wherever it is named, and so is one
named at the wrong scope. Adding a seat is a decision, recorded, for the reason every addition to the
host's vocabulary is one: the set is what a person reads to learn what a mesh can have, and an entry
nobody argued for is an entry nobody can explain.
## The set
**The set is derived, and only the mesh's half is written here.** Revision, 2026-09-26
([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md), superseding
[ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md)): a module
declares its own seats with their protocols, so the seats a mesh has are the mesh's own **plus
every registered module's**. The set is still closed — a seat named nowhere is refused — but it is
computed from the catalogue rather than maintained by hand, which is the property 0110 actually
needed and the table could not keep.
**And the mesh's own half is data, named for its scope.** Revision, 2026-09-27, reconciling two
records made in parallel: [ADR 0121](../../02-DECISIONS/0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md)
names a system seat for the scope it is held at — `mesh-*` for one per mesh, `node-*` for one per
machine — and [ADR 0122](../../02-DECISIONS/0122-a-seat-is-data-a-rename-is-a-database-update.md) moves
the set out of compiled code into a table the controller owns, so a rename is one write rather than a
rebuild of everything that names one.
So the set has two halves and neither is written out here: the mesh's own, which the controller holds
as rows, and every registered module's, which is computed from the catalogue. What this document keeps
is what a seat *is* — the rest would be a third copy, stale the first time somebody renamed one, which
is the fault ADR 0122 exists about.
**Every seat below is named `mesh-*`, and the prefix is the reservation rule**: a module declaring
any `mesh-*` name is refused at registration, so there is no reserved-names list to drift. Ten of
these are renamed to restore [ADR 0079](../../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md)'s
convention, which later seats departed from.
| seat | was | scope | delivers | typically held by |
|---|---|---|---|
| `mesh-controller` | — | mesh | — | the controller |
| `mesh-store` | — | mesh | — | the store the mesh's own records live in |
| `mesh-broker` | — | mesh | `mesh-bus` | the broker carrying the mesh's own bus |
| `mesh-vault` | — | mesh | `secret`, reserved | the vault |
| `mesh-artifact-store` | `the-artifact-store` | mesh | `artifact-store` | the artifact registry |
| `mesh-catalog` | `the-catalogue` | mesh | — | the catalogue |
| `mesh-npm-package-registry` | `npm-package-registry` | mesh | `npm-package-registry` | the forge |
| `mesh-git` | `git` | mesh | `git` | the forge |
| `mesh-build-machine` | `the-build-machine` | node | — | a builder |
| `mesh-dns-port` | `the-dns-port` | node | — | the local resolver |
| `mesh-intrusion-prevention` | `the-intrusion-prevention` | node | — | an intrusion-prevention service |
| `mesh-packet-filter` | `the-packet-filter` | node | — | the packet filter |
| `mesh-private-network` | `the-private-network` | node | — | the private network the mesh runs over |
| `mesh-resolver-configuration` | `the-resolver-configuration` | node | — | whichever of the alternative resolver configurations is chosen |
| `mesh-showcase` | `the-showcase` | node | — | the showcase module |
The controller holds **the mesh's own** entries in code, and a test asserts their size and that
every one names the record that made it a seat. A module's seats are not here and never will be —
they are read from the catalogue. **This table and
[ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md) govern, and code that
disagrees is what is wrong.** The implementation in progress predates several
things here: seats held by assignments rather than claimed by definitions, the `mesh-vault` seat and
its reservation, and the foundation's seats delivering nothing. It is brought to this table before it
merges.
## The foundation's seats
`mesh-controller`, `mesh-store` and `mesh-broker` name which assignment the mesh *itself* uses: the
controller, the store holding its records, the broker carrying its bus. **They route no consumer.** The
store and broker modules may run on other nodes too. A database or `amqp` consumer is served by
co-location, from whichever runs on its own node, the seat's holder included
([23 — Choosing a provider](23-choosing-a-provider.md)). A requirement cannot name one of them,
because they deliver nothing.
## A seat that delivers a provision
**A seat delivers a provision only where the mesh has one answer for everyone.** The artifact store,
the npm registry, git and the vault are each one per mesh by decision. A seat that delivers a
provision may only be held by an assignment of a module that provides it, at the seat's scope.
**A requirement may name the seat, and then its holder answers.** Naming the seat asks for *the
mesh's* one, so the holder answers **even when another provider runs on the consumer's own machine**,
and nobody is asked anything. With the seat unheld, the requirement is refused, naming the seat. A
second provider can run beside the holder and harm nothing. A forge assignment holds
`npm-package-registry`, and an npm proxy may provide the same provision on another machine. A builder
that names the seat is still served by the forge, without anybody pinning it.
**A requirement that names no seat resolves as any other**: a pin, the provider on the consumer's own
machine, the only provider. If several remain and none is local, a person chooses when the module is
assigned. The candidates are listed with the seat's holder suggested first, and the answer is recorded
as the assignment's pin ([27](27-a-module-requires-the-mesh-resolves.md)). Nothing is guessed, and
nothing changes silently because a second provider happened to appear nearby.
**Moving the role is changing which assignment holds the seat.** No definition changes and nothing is
unassigned: the forge keeps running, and keeps holding `git`, when its npm role moves. A module can
take the role only if its definition says it can hold the seat.
**The vault's provision is reserved.** Only an assignment holding `mesh-vault` may provide `secret` at
all: a definition providing it that cannot hold the seat is refused, an assignment providing it without
holding the seat is refused, and a `secret` requirement always names the seat, because there is no
other provider. A second provider of secrets would be a second place secrets live, which is what the
vault being one per mesh exists to prevent.
**What a consumer receives is what it required**, the same as for any provision: where the provider
answers, what it serves, and a credential. A consumer never reads the seat directly. The one exception
is the controller itself, which reaches the store and the broker through a narrow seat placeholder,
because it made them before any module existed and cannot be their consumer. One foundation module
also reads it today, to find its own server's port. [27](27-a-module-requires-the-mesh-resolves.md)
moves that to a host port requirement.
## A seat that delivers nothing
**A module's declared seat may promise nothing too, and that is a marker seat.** Correction of fact,
2026-09-27: [ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md)'s "a declared seat
carries a protocol" governs what a *holder* must satisfy, not that every seat offers something. A
marker seat's protocol is satisfied by holding it, which is the whole of what a marker says. Refusing
an empty one would refuse most of the node-scoped set, the showcase module's own seat included.
Nothing reaches that state by accident: an unknown manifest field is refused outright, so an empty
protocol was written as one. Checked by a registration test accepting a node seat with no protocol
and by the showcase manifest, which declares one.
Most node seats deliver nothing. They say which module is this machine's packet filter, or which of
two alternative resolver configurations it runs, and a second holder is refused. That is the whole of
their job, and it is a real one: it is the mesh saying what a machine is, in words a person can read.
## The overview
The controller lists every seat in the set with its scope, what it delivers, and its holder as a node
and a module. A seat nobody holds is listed as unheld. That is an answer, "this mesh has no forge",
and not a fault.
Holdings are derived from assignments whenever they are asked for, never stored. The list is always
what the mesh is running, because it is computed from the same thing that decides what the mesh runs.
## The git seat, and where a build comes from
A module is built from a repository, a path and a ref. The repository is one of two things, and the
mesh records which:
| form | means | recorded as |
|---|---|---|
| on the `git` seat | a repository on the forge that holds the seat | its path on the forge, and the seat |
| external | a repository anywhere else, a public forge for instance | its URL, exactly as given |
For a repository on the seat, the controller composes the clone URL at the moment of building, from
where the holder runs and the scheme and port it serves for `git`. The recorded source never contains
an address, so moving the forge changes nothing that was recorded. The build machine is not told the
difference: it receives a URL either way.
With the seat unheld, a build from the seat is refused and says why. External builds carry on.
**Not yet designed:** a credential for cloning a private repository. The mesh's own repositories are
public. The natural place for a clone credential is a `secret` from the vault, and that is a decision
still to take.
## How it is checked
The rules here are [ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md)'s —
which supersedes [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md)
and keeps every rule below except how the set is formed — and
[ADR 0111](../../02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md)'s. Each is
checked as their tables say:
| Rule | Checked by |
|---|---|
| The set is closed, and every mesh entry names its decision | 0118: a unit test on the mesh's own entries. **The refusal of an unknown seat is at registration, not in the parser** — correction of fact, 2026-09-27: a module may hold a seat *another* module declares, which is the point of naming the seat and not its provider, so whether a claimed name exists is a fact about the whole catalogue and a manifest in isolation cannot be judged on it. Registration tests cover an invented name and a name another module declares; the parser still refuses a claim on the mesh's own `mesh-*`/`node-*` namespace and a scope that disagrees with a declaration in the same manifest. |
| The set is derived, and enumerating it is a query | 0118: the overview lists the mesh's own plus every registered module's, asserted against a fixture mesh. |
| `mesh-*` is the mesh's, and a module may not declare one | 0118: a registration test refusing a manifest that declares any `mesh-*` seat, naming the prefix. |
| Two modules cannot declare the same seat | 0118: a registration test; the second is refused and the first untouched. |
| A holder satisfies the seat's protocol | 0118: a claim whose module does not serve what the seat declares is refused at assignment. |
| A seat is held by one assignment, and only by one whose module can hold it | 0118: resolution tests for a second holder and for a seat the definition does not name. 0131: `CanHold` is the one judgement, shared by registration and the handover, and its test follows the store's row. |
| A holder on record settles the seat; another eligible assignment is silent, not refused | 0131: resolution tests with a recorded holder on the same machine, on another machine, and under a seat's former name; without a record, the old rule's tests still pass unchanged. |
| A handover replaces the holder as one write, needs an assignment to point at, and goes with it | 0131: store tests — a second handover leaves one row; a handover to a module not assigned where named is refused; unassigning the holder removes the row. |
| A requirement naming a seat is answered by its holder; a foundation seat cannot be named | 0118: resolution tests with a second provider on the consumer's node, with the seat unheld, and naming `mesh-store`. |
| Several providers and none local is a person's choice | 0118: an assignment test listing candidates with the seat's holder first and recording the pin. |
| `secret` is reserved | 0118: the parser and resolution refusals for another provider and a pin. |
| Holdings are derived, and the overview lists every seat | 0118: the `seats` command test, including an unheld seat. |
| A build source on the seat records no address; an unheld seat refuses only self-hosted builds | 0111's tests. |
@@ -0,0 +1,394 @@
---
layer: to-be
status: proposed
code: []
updated: 2026-09-26
decisions:
- 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md
- 02-DECISIONS/0115-one-assignment-of-a-module-per-node.md
- 02-DECISIONS/0113-the-vault-makes-every-secret.md
- 02-DECISIONS/0114-a-shared-credential-rotates-over-two-credentials.md
- 02-DECISIONS/0126-a-module-declares-its-own-seats.md
- 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md
- 02-DECISIONS/0084-which-provider-serves-a-consumer.md
- 02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md
- 02-DECISIONS/0038-the-mesh-assigns-the-port.md
---
# 27 — A module requires, the mesh resolves
**One concept for everything a module needs.** A module definition states what it requires. Each
requirement has a contract and a kind of provider. Installing the module on a node resolves every
requirement, or refuses and says why. Nothing else reaches a module: no path it chose, no setting
beside the model, no literal it carries.
This replaces six mechanisms that grew separately: provisions read through bindings, settings,
assigned ports, machine facts, minted or accepted secrets, and literals in the definition. Each
resolved, validated and failed in its own way ([ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md),
[issue 119](../../04-ISSUES/119-a-module-definition-decides-where-its-files-live/00-report.md)).
## A requirement
A requirement has three parts:
| part | is |
|---|---|
| name | what the module calls it, unique within the module |
| contract | the fields the module may read, and what each promises: a type, whether it is secret, and anything the provider must honour |
| provider kind | which of the four kinds of provider answers it |
**A contract is shared, not per module.** A database's contract is the database's, whoever requires
it. **The controller holds every contract**, one per provision name, declared where the provision is
defined in the catalogue. Today contracts are implicit in each provider's served fields; the first
phase below makes them explicit, because nothing can be checked against a contract that is not
written down. A provider is checked against the contract it claims to answer. A module's own
specification may narrow a contract (a password of at least this length, a directory owned by this
user) and never widen it.
## The four kinds of provider
The set is closed, like the seats. A fifth kind is a decision, because each kind is a place an answer
can come from and a reviewer has to know every one.
| provider | answers | resolved by | replaces |
|---|---|---|---|
| **a module** | a database, a bucket, a vhost, a route, a secret | the rule below | provisions and bindings |
| **the node's host** | a directory, a port, a fact about the machine | always the module's own node | resource paths, assigned ports, machine placeholders, facts |
| **the mesh** | the module's identity and names, and the delivery of every answer | the controller | derived logins and generated names; the controller's delivery |
| **the operator** | a value a person chooses | the assignment, else the requirement's default | settings, carried literals |
### A module provider
Which module answers, in order:
1. **the holder of the seat the requirement names.** A requirement may name a seat instead of leaving
the provider open. It asks for *the mesh's* one, and the mesh answers with whichever assignment
holds that seat, with nothing asked of anyone. Unheld, the requirement is refused, naming the seat
([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md) (superseding [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md)),
[26 — The seats](26-the-seats.md)). Only a seat that delivers a provision can be named; naming a
foundation seat is refused, because it delivers nothing. A `secret` requirement always names
`mesh-vault`, because that provision is reserved;
2. **a pin**: the assignment names a provider, because this consumer is coupled to that provider's
contents ([ADR 0084](../../02-DECISIONS/0084-which-provider-serves-a-consumer.md));
3. **the provider on the consumer's own node**;
4. **the only provider** in the mesh;
5. otherwise **a person chooses, at assignment**. Assigning the module lists the candidates, with the
holder of a seat that delivers the provision suggested first, and the answer is recorded on the
assignment as its pin. Without an answer the module is not assigned, and the refusal names the
candidates. Nothing is ever guessed ([ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md)).
### Secrets: provisioning all the way down
**Two kinds of secret, one rule each** ([ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md)):
- **a shared secret**, a value more than one party must hold (a password, a token, an API key), is
made by the vault, and by nothing else;
- **a private key**, such as a node's sealing key, the operator's key or the mesh's certificate
authority, is made where it is used and never leaves. A private key anyone else held would no
longer be private.
**Every shared secret is a `secret` requirement, and only the vault provides `secret`.** The vault
holds the `mesh-vault` seat, and that provision is reserved to it: no other module may provide it, and
no pin can choose another provider ([26 — The seats](26-the-seats.md)).
**A provider that needs a secret for a consumer requires one, like any consumer.** A provision's
contract declares it: *for each consumer, one secret*. Resolution expands that into one requirement
per consumer, named for that consumer:
1. gitea requires `postgres-database`;
2. the database provider, to serve gitea, requires a `secret` named for gitea;
3. the vault makes it and hands it to the mesh;
4. the mesh delivers it to both of its **recipients**, each sealed to its own node: the database's
machine, which *applies* it by creating the login, and gitea's, which *presents* it;
5. the database provider creates the login, exactly as it does today, and gitea connects.
Every other shared secret takes the same path:
- a module's own secret;
- every broker account's password on the mesh's bus, where the broker's own provisioner creates the
account;
- an enrolment token, which the operator receives and the controller can only verify;
- a secret operator value, which the operator delivers to the vault;
- a secret a backend issues itself, such as a forge's API token, which the module that received it
delivers to the vault.
**The controller takes the same path, because it is a module.** Its store logins and bus accounts are
own secrets of its definition today, and become requirements of that definition. **A node's host is the
one party with no definition**, because it is what runs definitions. Its bus account is a requirement
the mesh makes for each enrolled node, answered and carried exactly as for a module.
**A provider makes resources and data.** Beyond secrets, a provider's adapter may answer with its
contract's non-secret fields: an analytics site id, a registered public name. The mesh carries them
back to the consumer as resolved values.
### The node's host
The host answers what only a machine can: where a directory is, which port is free, what the machine
is. It is always the module's own node, because none of these means anything elsewhere.
**A directory.** The contract is an owner and a mode, and the owner the image expects where it has
one. There is no persistence flag. A directory is kept while it holds anything, and data that may be
lost is a named volume, not a directory ([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md),
[ADR 0107](../../02-DECISIONS/0107-persistent-data-is-a-directory-bind-never-a-named-volume.md)).
*Where* a directory is on the machine is the assignment's:
- **a node's default layout**, a root per node with one directory per assignment beneath it, used when
the assignment says nothing;
- **a placement**, where the assignment puts one directory elsewhere: on a second disk, or where an
adopted machine's data already is ([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)).
**An operator's shared data** is an access, as before ([ADR 0051](../../02-DECISIONS/0051-shared-data-is-the-operators.md)):
never created, owned or removed by the mesh. The module requires read or read-write access. Where
the data is, is an operator value on the assignment.
**A port** is what [ADR 0038](../../02-DECISIONS/0038-the-mesh-assigns-the-port.md) already decided: the
module says which port its software uses, and the host answers with where the machine put it.
**A fact** is something the machine knows: its name on the private network, the names of the mesh's
machines. Each fact has a contract like anything else.
### The mesh
The mesh answers who the module is: its login, its broker account, the names it is known by. These
are derived by the mesh so every party agrees by construction, and no provider may make them
([ADR 0049](../../02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md)).
### The operator
A value a person chooses: a public name for an endpoint, a greeting, how many workers to run.
**It must stay cheap.** An operator requirement's contract is a type and, optionally, a default. It
needs no provider module, no grant and no credential. If asking a person for a value took more than
that, module authors would route around it, and the literals this replaces would come back.
**A secret operator value**, such as an external API key, follows the one rule for secrets: the
vault provides it. The operator hands the value to the vault, once
([ADR 0092](../../02-DECISIONS/0092-an-operator-delivers-a-pair-credential.md)), and the module requires
a `secret` like any other. The only difference is that rotation never replaces it: the vault cannot
make a new external key, so rotating one means an operator handing over a new value.
**An endpoint** is an operator value inside a route requirement: the public name is chosen on the
assignment, and the route provider answers. A public name already held by another assignment is
refused, like any other singular thing.
## How a definition reads what was resolved
**One form, naming a requirement and a field of its contract.** A definition that needs the database's
host in an environment variable, the directory's location on the host side of a mount, or the public
name in a configuration file writes the same thing: the requirement's name and the field. The
controller fills it at resolution.
This one form replaces the placeholders that exist today, one per mechanism: bound values, secrets,
ports and machine facts.
**The seat placeholder stays, for the controller alone.** The controller composes its own
declaration and reaches the store and broker it made before any module existed, so it cannot be
their consumer. One module reads the placeholder today: the store module, to find its own server's
port. That is its own port, so it becomes a host port requirement in phase 3, and after that no
module uses the seat placeholder.
**A secret field reaches a process as a file**, as [ADR 0086](../../02-DECISIONS/0086-a-secret-reaches-a-process-as-a-file.md)
decided. The one exception 0086 allows is a declared env-file with its reason; a secret field as a
value in a container's environment is refused when the definition is parsed, with no exception.
## An assignment
**A module is assigned at most once to a node**, and that pair is the assignment's identity
([ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md)). Its directories,
containers, login, broker account and settings are keyed by it, as today, and a login still fits the
tightest backend ([ADR 0049](../../02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md)).
**A module may run on many nodes, and one assignment may hold a seat**
([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md) (superseding [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md))). The definition
says which seats the module can hold; the assignment says which it does. So the store module can run
on every node, one of those assignments holds `mesh-store`, and moving that role changes an
assignment, not a definition.
What must stay singular stays so: by a seat, or by an operator value colliding, as with a public name.
## Genesis
**Genesis delivers, and the vault adopts.** The vault cannot run first: it is built on the shared
runtime base, which the installation makes only after the store, the broker and the controller exist
([21 — The installation in full](21-the-installation-in-full.md)), and it learns what to answer from
the controller over the bus. So the vault is installed **as soon as that base exists**, before any other
module built on it, and genesis generates what is needed until then:
- the store's superuser, and the broker's admin in the hashed form the broker needs;
- the bus accounts of the temporary and permanent controller (its account and the broker-management
login), the control-node's host, the builder, the broker's own provisioner and the vault;
- the controller's three store logins, and the first enrolment token.
Until the broker's provisioner runs, genesis creates those bus accounts with the broker's admin, as the
controller does today; the provisioner adopts them when it starts. Genesis seals everything to the
control-node's key, and when the vault is installed the controller **delivers the values to it,
recorded as the mesh's own**, with nobody present. That distinction keeps
them rotatable: an operator's value is never replaced, and these are, because the vault can make their
replacements.
That is the one time anything but the vault generates a shared secret, and it ends by handing them
over. It is also the answer to the objection ADR 0085 had to the vault being the only maker: the
vault cannot make what exists before it, so what exists before it is delivered to it.
**Raising the vault or the broker again is a genesis act.** Moving either seat to a new assignment, or
recovering either after it is lost, delivers the values it needs the way genesis did. It is a stated
break-glass procedure, and an ordinary assignment attempting it is refused.
## Rotation
Rotating a secret is asked of the vault, by an operator or by the vault's own policy, such as a
maximum age in the requirement's contract, and the vault makes the new value
([ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md)). An operator's external key is
not rotated by the vault, which cannot make its replacement: an operator delivers a new one.
**Each requirement says how its recipient takes a new value.** It either *applies* it, through a
provisioner (a provider setting a login's password, the broker's provisioner updating an account, a
store's provisioner changing its own superuser), or *reads it at start*. Every module in the catalogue
reads its secrets at start, and none watches them. The host already recreates a container when a file
it read at creation changes, its env-files and files mounted into it directly
([issue 103](../../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md)).
So a reader's restart is derived, and a definition declares `restart-on` only for a secret reaching a
process, or a file in a mounted directory. A secret a service takes only at first initialisation is
applied by its provisioner or marked not rotatable by the mesh, and a rotation of it is refused rather
than reported done.
**How old and new change over depends on how many parties hold the credential**
([ADR 0114](../../02-DECISIONS/0114-a-shared-credential-rotates-over-two-credentials.md), on
[research 016](../../01-RESEARCH/016-how-a-credential-can-be-rotated/00-overview.md)):
- **Two parties**, a consumer and its provider, or a module or node's host and the broker: each consumer
has two credentials, both with its rights over one resource named after the consumer. For most
providers the second is a second login derived by the mesh; where a backend's user is its resource, it
is a second token. The vault drives the rotation and records each step. It makes the new value; each
applier ensures the unused credential and confirms both authenticate; only then are readers given it
and recreated; each reader confirms by authenticating with it; and only then is the old one retired.
Nobody is left without a credential that works, a rotation can be abandoned until the old one is
retired, and `status` shows who a rotation waits on.
- **One party**, a provider's administrative credential or a module's own secret: in place, staged.
An applied one is delivered beside the current value, the provisioner changes the backend with the
current one, and only then does the new value become current. One read at start is delivered, and
the host recreates the module.
**Retiring a credential never removes what it reached.** An adapter keeps *retire a credential* and
*remove the consumer* apart, and the harness keys what it applied by consumer, so a changed login is
never a removal. The resource is removed only when the consumer no longer requires it from that
provider: unassigned, the requirement dropped, or re-resolved elsewhere. Today these are one call, and
in five providers it deletes the consumer's data, so this separation comes first. An adapter that cannot
yet ensure a second credential rotates in place, with its window stated, and is listed until it can.
## Refusing
Installation refuses when any requirement is unresolved, and **says everything at once**. For each
requirement it names what is missing and what would answer it:
- an unheld seat, and which modules could hold it;
- no provider, and which modules could provide it;
- several candidates and no choice made, and which they are;
- an operator value with no default, and that the assignment must give it;
- a provider, or the vault, that has not answered yet, and which one.
The last one is a state, not a failure. A consumer waiting for its provider or for the vault is shown
as waiting, and nothing is delivered until the answer arrives.
## What this retires
| mechanism | becomes |
|---|---|
| provisions read through bindings | a module requirement; its answer is the contract's fields |
| settings on an assignment ([ADR 0046](../../02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md)) | operator requirements on an assignment |
| a port the mesh assigns | a host requirement |
| machine facts and machine placeholders | host requirements |
| every secret the controller mints: provider credentials, own secrets, broker passwords, enrolment tokens | a secret the vault makes ([ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md)) |
| root secrets genesis mints and keeps apart | made by genesis once, then delivered to the vault, which holds and rotates them |
| a separate command issuing a broker account | a requirement resolved on assignment |
| `restart-on` naming a secret's file | a restart the host derives |
| paths in resources, mounts, bindings, secrets and received files | host directory requirements, placed by the assignment |
| literals carried in a definition | operator requirements with defaults |
Each is retired only once nothing uses it. Until then both are accepted, and a catalogue test lists
the definitions still using the old form. That list shrinks to empty, and then the old form is
removed from the parser.
## Phases
Each phase ends at a check that holds, so none of them leaves a mechanism half-replaced.
1. **Contracts, resolution and the new form.** Every provision's contract is written down and held by
the controller. The controller resolves requirements from the four providers, refuses as above,
and fills the one form. Old mechanisms keep working beside it. *Ends when* a definition written
entirely in the new form installs on a lab machine.
2. **The vault makes every shared secret, and providers answer.** The vault holds its seat and its
reserved provision; resolution expands per-consumer secret requirements; genesis delivers the
foundation's first secrets to the vault; the broker's provisioner creates every bus account; the
mesh carries providers' data back.
[Issue 103](../../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md)
is fixed in the host already. *Ends when* nothing outside the vault generates a shared secret after
genesis, a lab consumer of analytics receives its site id, and a database credential rotates over
its two credentials, the consumer recreated by derivation, never without a working login, and its data intact.
3. **Definitions move, and seats move to assignments.** Every catalogue definition is rewritten, adopted
and running assignments placed where their data already is, and each claim becomes a seat the
module can hold, held by the assignment that holds it today. *Ends when* the list of definitions
using an old form is empty, the old forms are removed, and the store module runs on two lab
machines with one holding `mesh-store`.
## How it is checked
| Rule | Checked by |
|---|---|
| Every requirement has one of the four provider kinds | The parser refuses any other. |
| A definition names no host path, node or mesh | The catalogue tests of [ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md). |
| A module provider is chosen by named seat, pin, co-location, only one, a person's choice | Resolution tests for each step: a requirement naming a seat served by its holder even with another provider on the consumer's node, and refused when the seat is unheld; several candidates and none local, where assignment lists them with the seat's holder first and records the choice as a pin, and refuses without one. |
| A provider answers within its contract | A controller test: an answer carrying a field its contract does not name, or missing one it does, is refused and not delivered. |
| A module narrows a contract and never widens it | The parser refuses a module specification that loosens a contract's field. |
| A host requirement is answered on its own node | A resolution test placing one elsewhere: refused. |
| An operator value needs no provider module | A resolution test: a requirement with a default resolves with no module assigned anywhere. |
| A secret field reaches a process as a file | The parser refuses a secret field as a container environment value, and accepts it in a file or a declared env-file with its reason. |
| A public name already held is refused | A resolution test: a second assignment asking for a public name another holds is refused, naming the holder. |
| A module is assigned at most once to a node | A resolution test: assigning a module to a node that already runs it is refused. |
| A seat is held by an assignment, not a module | A resolution test: the store module on two nodes, one holding `mesh-store`; a second assignment asking to hold it is refused. |
| A consumer waits for its provider | A resolution test with a provider that has not answered: shown as waiting, and nothing delivered. |
| Only the vault generates a shared secret after genesis | A controller test: no code path generates one. An installer test: genesis generates exactly the foundation's first secrets and delivers them to the vault. |
| Only the vault provides `secret` | The parser refuses another provider of it, and resolution refuses a pin on a `secret` requirement. |
| A provider's per-consumer secret comes from the vault | A resolution test: requiring a database expands to a secret requirement named for the consumer, answered by the vault and delivered to both recipients. |
| Restarts are derived | The tests of [ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md): an applied secret restarts nothing, and one read at start recreates its reader without a declared restart. |
| A two-party credential rotates over two credentials | The rotation tests of [ADR 0114](../../02-DECISIONS/0114-a-shared-credential-rotates-over-two-credentials.md): retiring a credential leaves the resource intact; a changed login is never a removal; readers move only after the applier confirms and confirm by authenticating; an unreachable reader keeps its old credential until it returns; rotation state survives a restart; a single-party applied secret is staged. |
| A private key is made where it is used | The per-key tests of [ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md): a node's sealing key, the operator's key and the certificate authority's key never leave where they were made. |
| The controller and a node's host take the same path | 0113's tests: the controller's definition declares requirements and no own secret; a node's bus account is made by the vault and delivered sealed to that node. |
| Moving the vault or the broker is break-glass | A resolution test: an ordinary assignment moving `mesh-vault` or `mesh-broker` is refused, naming the procedure. |
| A secret that cannot be rotated says so | A vault test: rotating a secret marked not rotatable by the mesh is refused, naming why. |
| Only the controller reads the seat placeholder | A catalogue test, from phase 3: no definition uses the seat placeholder. |
| Refusal names everything at once | A resolution test with three unresolved requirements of different kinds: one refusal naming all three. |
| The old forms retire | The catalogue test listing definitions still using one. It must be empty before a form is removed. |
## Not settled here
- The exact spelling of the one form. It must name a requirement and a field and nothing else.
- The layout a node's default root uses beneath it, beyond one directory per assignment.
- Whether a module provider's answer can change without the provider being asked, for example a
provider moving. The rule so far is that it cannot, and moving is re-resolving.
## The three gaps, answered (2026-09-26, operator)
**Where the root comes from.** A node setting, fixed at installation; a node that states none
gets the default under `/var/lib`. (The adopted-machine placement stays as proposed: a
per-assignment placement for data that must sit where it already is — mssql is novox's live
case.)
**What sits beneath it — dissolved, not decided.** The question assumed the mesh's own
writes (`/var/lib/mesh/<module>`: sealed credentials, composed bindings) need a
module-visible reservation. They do not: a module *requires* a `host-path` and receives a
location; what the mesh writes for the module is the mesh's plumbing, placed where the mesh
chooses and mounted in — never part of the module's contract. One reservation per
requirement, `<root>/<module>/<name>`.
**Resolution happens in the controller, at declaration composition.** The node receives
concrete paths exactly as today — the wire format and the host's apply do not change for
this. What changes is that no *manifest* carries a path; the controller resolves
requirement → location against the node's root setting. (The host still needs issue 126's
fix — volumes in the spec comparison — or a resolved path change cannot reach a running
container.)
**Retirement order** stays per-module data migration with verification, smallest and
empties first, the mail spool and the store last — the novox session's six-module window
(mesh-catalog #97, hq 126) is the worked example, hazard included.
+751
View File
@@ -0,0 +1,751 @@
---
layer: to-be
status: in-progress
code:
- mesh-catalog modules/nats
- mesh-controller internal/catalogue
- mesh-lab scenarios
updated: 2026-09-28
decisions:
- 02-DECISIONS/0116-the-bus-is-built-in-five-steps.md
- 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
- 02-DECISIONS/0126-a-module-declares-its-own-seats.md
- 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md
- 02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md
- 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
- 02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md
- 02-DECISIONS/0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md
---
# 28. Building the bus
**The work of [ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md)'s five steps,
in the order its dependencies allow, with what each ends at.**
[Design 25](25-the-bus-on-nats.md) is the architecture and stays the authority on *what is built*;
this document holds only the order, the sizes and the proofs, and it is wrong the moment it
disagrees with design 25 rather than the other way round.
Each step ends at something runnable. A step that cannot name what its bed proves is not a step,
and is divided further before it is started.
## How this is built, and when it is run
**Written as code with unit tests, committed per change, and taken to the lab once the pieces that
would change the outcome are in place.** The mistake this avoids is the one
[design 22](22-the-work-ahead.md) records: running a long bed against a mesh mid-transformation and
debugging paths the next step deletes. Where a fault can be reasoned out of the code path, it is —
reading, not running.
So the beds below are acceptance tests at the end of assembled work, not the tool for finding each
bug, and a step's bed is run when that step is finished rather than while it is being written.
## What the work is, measured
Counted 2026-09-26, non-test source only. The point of counting is that none of this is unknown
territory: every piece has a shape already standing beside it.
| Piece | Today | Size | Becomes |
|---|---|---|---|
| the controller's link | Go, one package | ~1 800 lines | the same package on NATS |
| the host's link | Go, one package, mirroring the contracts rather than importing them | ~1 000 lines | the same, on NATS |
| the tool runtime's client | TypeScript, one file | ~390 lines | the same, on NATS |
| the sdk's messaging surface | TypeScript: messaging, events, tools, contracts, primitives | ~360 lines across five | **unchanged**, see below |
| the broker module | the adopted AMQP broker: client, tools, provisioner, bootstrap, image, manifest | ~340 lines of module code | the `nats` module, same shape |
| the beds | 39 lab scenarios, including a broker bed, an adoption bed, a genesis bed and a store-window bed | — | four analogues and one new |
**Two measurements are worth stating on their own, because they change what the steps are.**
**The sdk speaks no AMQP, and never did.** The word appears in its source three times, in three
comments; its messaging module says in as many words that it "carries the contract, not a specific
AMQP client build." [ADR 0039](../../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md) put the
client in the runtime, and the payoff is collected here: **no module is rebuilt for this change, and
the sdk's own diff is three comments.** That is the whole reason a bus can be replaced under a live
mesh at all.
**The wire therefore has three implementations, not two, and no suite pins any of them.**
[ADR 0074](../../02-DECISIONS/0074-the-wire-is-specified-not-the-types.md) spoke of "the existing
two implementations" — Go and TypeScript. Measured, the Go side is *two separate packages* that
mirror rather than share (the host imports nothing, by
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md)), so the count is the controller's link, the
host's link, and the runtime's client. And a search for conformance fixtures finds none anywhere in
the four repositories: design 22's Phase 1.2 — the suite — has not been built.
> **This corrected the record.** ADR 0116 said step 3's fixtures were *recaptured* on NATS. There
> is nothing to recapture, so step 3 **builds** the suite, and its first job is to pin the wire the
> mesh has before changing it — a suite written only against the new bus certifies whatever the new
> bus happens to do. A fact went stale while the decision stood, which is a **progressive insight**
> ([`02-DECISIONS/README.md`](../../02-DECISIONS/README.md)): it is marked and dated in ADR 0116
> itself rather than left to be discovered here.
## The order the work actually allows
**The five steps are chunks of capability; the build order is not simply 1 to 5, and pretending
otherwise would put two beds where they cannot run.** Three edges decide it:
- **A specification precedes the implementations it governs.** ADR 0074's whole argument is that
agreement is specified and checked, not hoped for. So the wire's NATS binding is written *before*
the three implementations are, even though it is step 3 — and its conformance half can only
*finish* once two implementations exist to disagree.
- **A mesh cannot be raised on a bus nothing speaks.** A bed that raises a mesh on NATS from
genesis — enrolling a node, holding a push while the store restarts, rolling out an upgrade —
needs the controller and the host to speak NATS already. That is the implementations, and they
arrive with step 3.
- **Adoption needs the module and nothing else.** Step 2 puts a correctly configured server into a
running mesh that continues to ignore it, which depends on no link at all.
So step 1's bed proves *the server, from genesis, configured* — not a mesh living on it. The full
genesis bed is step 4's, where it can first run.
> **This corrected the record too.** ADR 0116 first attributed "a mesh raised on NATS from genesis"
> to step 1. That bed cannot run until the links exist, and a step whose proof cannot run is the
> exact failure the record was written to prevent — so step 1 now ends at the server standing,
> correctly configured and carrying nothing, and the full bed is named under step 4. The five
> steps, their names, their order and the single rollout are unchanged; only where two beds run has
> moved. Marked and dated in ADR 0116 as a progressive insight, with what the record said before.
```
step 1 module, genesis places it ──┐
step 2 adoption into a running mesh ──┤ neither needs a link
│
step 3 the wire specified ──► three implementations ──► the suite
│
step 4 the flows, and the full genesis bed
│
step 5 the rollout
```
## Step 1 — the module, and genesis raises it
> **Revised 2026-09-26** ([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md),
> [design 29](32-what-a-module-declares.md)). Tasks 1.3 and 1.4 said the controller composes every
> account and creates *the four streams* at genesis, from a fixed set. That is only the mesh's own
> half. A module declares seats with their protocols, so streams are created **at registration**
> and durable consumers **at assignment** — neither of which has happened at genesis. The fixed
> foundation set stays here; the derived machinery moves to step 3, where the declaration model it
> reads from is specified. Tasks 1.1 and 1.2, already done, are untouched by this: the module and
> its reload mechanism do not care what the configuration says.
**Why here.** Everything else needs a server to talk to, and genesis is where the foundation is
defined. The mesh this is for will never travel this path — it is already running, and takes step 2
— but genesis is the definition every other path is measured against, and one that exists only on
paper is wrong until there is a second mesh to find out.
- [x] 1.1 the `nats` module: manifest, image, one container, its client, TLS and monitoring ports,
JetStream on a named volume — the shape of design 25 §5, and the same shape the broker module
beside it already has
- [x] 1.2 the composed configuration as a **directory** resource, and the entrypoint that watches
the one file and signals the server itself — design 25 §5's correction, kept inside the module
because a container has no reload and a recreate would drop every connection the mesh has
- [x] 1.3 the controller composes that file: accounts, permissions, TLS, JetStream — a user's
permissions derived from its declaration and nothing else, over the three namespaces of
[design 29](32-what-a-module-declares.md) §2, plus its own ack subject and its own inbox
prefix (design 25 §4)
- [x] 1.4 the mesh's own streams, created at genesis and asserted idempotently on start, by the
controller as their only writer — **the mesh's own, not all of them**: a seat's streams are
created when the module declaring it is registered, and a module's durable consumers when it
is assigned, so this task is the fixed foundation set and 3.x carries the derived rest
- [x] 1.5 genesis raises it as foundation, claiming the seat **`mesh-broker`** — the seat is the
server's role, not the product. **Already true of the controller and needed no change**: it
resolves the broker by seat ("that is where the broker is, whatever else the topology says")
and names no broker module anywhere in its source. What remains is naming `nats` instead of
the deprecated broker where a genesis module set is declared, which is scenario and installer
configuration — carried with 1.6 rather than before it.
- [x] 1.7 **the composition, delivered** — the controller gathering its principals, composing the
file, and asserting the streams and consumers on start.
**In**: the user list is derived from the mesh's records and the credentials are kept.
A bus user's bcrypt hash is now recorded, keyed by the username the file needs, and the
plaintext is returned exactly once. That state is new and the reason is worth stating: on the
bus the mesh runs on today an account is a management call — mint, hand over, seal to the
holder, keep nothing — and that works because the broker remembers. Here the users are one
file rewritten whenever any of it changes, so keeping nothing would mean **the first person's
access change silently blanking every module's password**.
**Permissions are not kept, only credentials.** Authority is derived from what each module
declares every time the file is written ([ADR 0043](../../02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md));
a stored permission list would be a second account of a user's authority, able to disagree
with the records it came from while both looked internally consistent.
The derivation refuses two things where they can still be named: two users with one name — the
server reads the file as one of them and which one depends on the order — and a module
assigned but absent from the catalogue, which would compose a user with no authority and fail
on its first publish with an authorisation error that says nothing about a missing manifest.
A user the mesh has minted no password for is *named* rather than dropped or written as a user
anybody is: an ordinary situation with an obvious remedy, and the caller decides whether a
partial file is worth writing. A seat's protocol is gathered across the whole catalogue, not
from one manifest, because a seat is declared by one module and held by another.
**Out, and what each needs.**
**Delivery is in, and it settled what a module declares.** The mesh writes the *accounts* and
the module owns its *server*. The alternative was a manifest field enumerating ports, TLS
paths and a store directory so the controller could write a whole configuration — wrong,
because those are properties of the container the module raises and the controller would have
to be kept in step with a Dockerfile it never sees. So a module declares its own configuration
as a file resource and `bus-users` names where the mesh's half goes beside it; **asking is not
enough to receive it**, because that file holds every user's password hash, so the claim on
`mesh-broker` is what authorises it.
Two things a running server changed. **An absolute include path is resolved relative to the
including file's directory** — `include /etc/nats/accounts.conf` from another directory makes
the server look for it *under* that directory and refuse to start — so both files share one.
And **`verify: true` was refusing every connection in the mesh**: it makes the server demand a
*client* certificate, and nothing in the mesh presents one — a host pins this server's exact
certificate and authenticates with the password the mesh minted ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md),
design 25 §4). Every connection would have died at the TLS handshake before any password was
looked at, with an error that reads as a fault in the client. Removed; TLS is still required,
because the block is what requires it and `verify` only decides whether client certificates
are checked. **Design 25 §4 should say this**, and says nothing about it today.
Also collected here: task 1.2's payoff, end to end against the module's own image — the user
list rewritten, the module noticing and reloading the server itself with no signal from
outside, and the connection the mesh already had still working afterwards.
**Minting is in, on both halves.** A node at enrolment, a module when its credential is
issued. Three things differ from a management call and each is the point of the move: the
credential is minted into the mesh's records and becomes usable at the next composition, so no
server need be reachable for it; the password travels beside the address rather than inside it,
because a credential embedded in a URL leaks into every log line that prints a connection; and
a module's durable consumer is derived from what it declared rather than named, so it cannot ask
for delivery of something it did not say it consumes. A node reconnecting may be refused until
the composition reaches the machine running the bus — which is what the host's reconnect backoff
is for, where waiting for the push would hold an enrolment open for as long as a declaration
takes to apply.
**Which bus is one fact, and being told about both is refused at start.** Not warned about: a
mesh half on each is one where a declaration goes out on one bus and the report comes back on
the other, and every component logs success while it happens — ADR 0074's failure arriving
through configuration instead of through code. A node that came away holding a credential for
each could be half-moved, and nothing would say which half.
**The objects are asserted on every start**, not created once at genesis: a stream somebody
deleted, a mesh raised from a restored backup, or a bus whose data directory was replaced all
have records and no objects, and a node whose consumer is missing hears nothing while everything
else about it looks correct. Against a real server: every object accepted, asserting twice
changes nothing (a start that failed the second time is a controller that cannot restart), a
machine joining an already-raised bus accepted, each node's consumer bound to its own
declaration subject and no other's, and CONTROL not dead-lettering — because the store window's
bound is the controller's, and a server that gave up first would discard the push the stream
exists to protect.
**People are not in the list**, deliberately: the account model is built and `operator issue`
is not (4.4), so there is nobody to derive. Left empty rather than guessed at.
> **This corrects a tick, not a decision.** Tasks 1.3 and 1.4 are ticked and they are honest
> about what they built — the composer, the derivation, the permission model, the stream and
> consumer definitions, the asserter, all pure and held by unit tests and a golden
> composition. What nobody wrote is the *caller*. Measured on the feature branch: outside the
> package that defines them, there is **not one** use of the composer, the permission
> derivation, the stream set, the stream asserter or the principal type. Step 1's "done when"
> claims "every account and permission composed from the manifests", and a mesh raised today
> would stand up a server with no user list at all.
>
> It also needs state the mesh does not keep. Design 25 §4 says the file holds bcrypt
> hashes, and passwords are "minted and sealed exactly as today" — but today the mesh mints a
> password, hands it to the broker through a management call, seals the plaintext to the
> holder and **keeps nothing**. There is no management call here, so the hash has to survive
> for every later recomposition: the first thing a person's access change or a new module
> touches is a file that must still contain every other user's password. No bcrypt hash is
> stored anywhere in the controller today.
>
> Named as its own task rather than folded into 1.3 so the gap is visible: the parts of
> step 1 exist and the mesh does not yet do any of it.
- [ ] 1.6 the genesis-broker bed — **deferred**: beds are run once, at the end, rather than per
step (novox/hq design 22's rule, and the operator's instruction). Every claim step 1 makes
is covered by a unit test or was demonstrated against the real server; what the bed adds is
the claims that need a mesh.
> **Not done here, deliberately.** The controller builds a module's broker credential as an
> `amqps://` URL and defaults a portless genesis address to 5671. Those are correct until the
> rollout and must not move: steps 1 to 4 leave every node on AMQP
> ([ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md)), so changing the
> credential's shape now would break the running bus to serve a bus nothing speaks yet. They
> change with the links, in step 3.
**Done when.** A mesh raised from nothing has the server standing with the streams asserted and
every account and permission composed from the manifests; a user cannot publish outside its
`emits`, subscribe outside its `consumes`, ack another user's delivery or subscribe another's inbox
prefix; the monitoring port is refused from anything but the private network; a change to the
composed file is live within one watcher interval without a restart, and the container is not
recreated by it.
**The permission checks belong here rather than later** because the server enforces them itself — a
plain client proves them, no link required — and they are the whole of what ADR 0043 asks for. No
mesh traffic is on the bus yet; that is step 4's bed, not this one's.
## Step 2 — adoption puts it in the seat
**Why here.** It needs only step 1's module, it is the path the mesh that exists will actually take,
and it is what makes steps 3 and 4 safe to develop against a live mesh. A running mesh does not get
a foundation module by being raised again; it adopts one in place
([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)).
- [ ] 2.1 the server raised beside the existing broker on its own ports, carrying nothing —
installer-side, from the upstream image
- [ ] 2.2 the `nats` module assigned, which **recreates the container once, deliberately** (see
below), keeping its JetStream directory
- [x] 2.3 the seat claim, and the resolver's refusal of a second holder mesh-wide — **already
true and now proved**: the refusal is generic to any mesh-scoped seat, and three tests pin
what matters for this one — a second bus anywhere is refused naming the seat, a *different*
bus implementation is refused for the same reason (which is what lets the bus be replaced
at all), and the deprecated broker no longer contends for it, so both run on one mesh
- [ ] 2.4 the adoption bed — deferred with the other beds
> **Adoption here is not a no-op, and pretending it would be is the trap.** The host keeps an
> existing container only when its spec matches the declaration exactly
> ([`apply.go`](https://git.novox.be/novox/mesh-host): *existed && before.Spec == want && running*
> → unchanged; anything else is `rm -f` and recreate). Genesis raises the server from the
> **upstream** image, because nothing has been built yet; the module declares the **mesh-built**
> artifact, which carries the entrypoint that reloads configuration in place. Those two specs
> differ, so assigning the module recreates the container.
>
> That is correct, and it is [ADR 0067](../../02-DECISIONS/0067-genesis-is-a-pivot.md)'s pivot
> exactly: raise a temporary thing, then reinstall it as an ordinary module. It is safe **only
> because it happens while the bus carries nothing** — which is what 2.1 means by "carrying
> nothing", and why step 2 comes before anything speaks NATS rather than after. One recreate, at
> the one moment it costs nothing.
>
> **After that, never again.** The configuration is a directory mount rather than a file, so
> rewriting accounts does not change the container's spec and the entrypoint reloads the server in
> place. That is the whole point of task 1.2, and this is the moment it pays: every later account,
> permission or person's access change touches a running bus with connections on it.
**Done when.** A mesh already running has the server adopted, holding `mesh-broker`; a second
assignment anywhere is refused at resolution — *one per mesh*; and every node is still on the old
bus with nothing routed to the new one. **That last check is the point of the step**: adoption that
quietly carried traffic would be step 5 arriving early and unrehearsed.
## Step 3 — the protocol on NATS
**Why here.** The implementations cannot be written against an unwritten wire, and this is the step
that decides what "agreeing" means for everything after it. It is the largest step and the one that
pays for itself furthest away.
- [x] 3.1/3.3 **the fixtures** — one directory in the sdk, read by each implementation's own
runner rather than copied into either, because a fixture copied twice is two fixtures. The
Go emitter and the runtime's NATS client both pass the first: every required header set,
each value in the pinned shape, the subject derived the same way, and the payload the body
alone.
**The suite also had to settle what "byte-for-byte" can mean**, which ADR 0074 stated and
nothing had yet had to implement. The envelope is exact — subject, required headers, names
and formats — because that is what two implementations get wrong invisibly. The body is
not: Go sorts a map's keys and JavaScript keeps insertion order, so identical bytes would
commit every implementation to a canonical JSON encoder, to buy a property the mesh never
uses. Read strictly it would have sent somebody writing one.
Still to capture: a served tool call, a grant and its answer, and the contributions file —
the other three ADR 0074 names.
- [x] 3.2 design 19 rewritten from exchanges, queues and routing keys to the subjects and streams of
design 25 §2–§3, per capability, with ADR 0074's model untouched: floor plus capabilities, an
implementation legitimate when it claims less, identity from the sealed credential, dedup on
`x-event-id`. Claims checked against a running server are marked *verified* in the text, so
a reader can tell what was measured from what was reasoned. One limitation lifts with the
transport: a module may now call another's tool, which issue 049 recorded it could not.
- [x] 3.4 the controller's link on NATS — **both halves are through the seam, and the store
window is the server's.** `Bus` states the outbound in the mesh's words (publish an event,
declare to a node) and `Control` states the inbound (took it, dropped it, held it for the
store); each has an AMQP and a NATS implementation, and both ship, because steps 1 to 4
leave every node on AMQP and both shipping is what holds them to one envelope.
The outbound seam turned out to be eight call sites; the inbound was the larger half, and
the reason: every handler took the transport's own delivery type, so the loop could not move
without moving enrolment, reports, builds, upgrades and catch-up with it in one breath.
**The window (ADR 0083) is now what decides, once, for both.** On the bus the mesh has,
holding a message means an unacknowledged delivery kept in the controller, bounded by the
prefetch and lost if it stops. On the bus being built it is a `nak` with a delay: the
message stays the server's and the controller keeps only the moment it first could not take
it, so one that restarts mid-window has nothing to lose. Seven claims about that were asked
of a running server rather than reasoned — a report heard and gone from the work queue, one
held through a store outage and recorded when it returned, one let go once the bound passed,
a superseded one settled without being acted on, a heartbeat heard and nothing persisted,
both followed events acknowledged on a stream the controller had no ack subject for, and the
enrolment answer arriving at the address the request carried in its payload.
**Three things the wiring forced into the open.**
*Supersession is asked before the store, not after.* A report about a declaration the mesh
has moved past would otherwise wait out a restarting store to be written, and then overwrite
what the node is doing now.
*Half of a report is not about a declaration, and that half is never stale.* What the machine
**is** — the tunnel it took over, the ports its own bundle holds, what an adopted node found,
a node moving its overlay key — reaches the mesh on a report and nowhere else. A rekey set
aside as stale is a node whose overlay key never moves, and no retry is coming, because the
node said it once. So staleness is asked only of a report that is purely an apply's account.
*The controller could not have consumed a module event at all.* Its account granted no event
subject to subscribe and no ack subject on the events stream, so every announcement would
have been redelivered for ever, refused by the permission list it already had. Both are now
granted, each subject named rather than by pattern — a controller subscribing every event in
the mesh is a permission list that has stopped saying what it is for. Its consumers are
**named beside the mesh's own streams rather than derived**, because the controller files no
manifest and authority cannot come from a declaration that does not exist.
Still outstanding: a build's own shape, which travels with the builder in step 4.
- [x] 3.5 the host's link on NATS — **all three halves are through seams**, mirroring the
controller's and still importing nothing of the mesh's own (ADR 0005): the host's own
interfaces over its own libraries, agreeing with the controller only because a fixture holds
both to one envelope. A report goes through JetStream because it is the message the
store-window guarantee is about; a heartbeat stays on core, because a heartbeat in a stream is
the mesh's least valuable message competing for retention with its most valuable.
`Link` is dialling, hearing and saying in one interface, because **dialling is where the
transport is chosen** and choosing it twice is how one half of a node ends up on a different
bus from the other. `Asking` is the enrolment conversation, and it is separate for the
opposite reason: almost nothing about it is the same, and a node that fails there is not in
the mesh at all.
**What the new bus took away, and what it would not give.** A host declares nothing here: on
the old bus it declares its own queue, because a queue that is not there means a node that
hears nothing, but the object it reads through now is a durable consumer and a host's account
reaches no part of the JetStream API. So it **binds** to one the mesh made, and a missing one
is said as the mesh's to answer rather than quietly created with whatever the client defaults
to. Two things that had to be built for that: a node's declaration consumer (named after the
node, because its ack grant is derived from the node's name, so any other name is a delivery
it cannot acknowledge), and the enrolment user's **inbox** — design 25 §6 names it and the
composer granted none, so an enrolling node would have published its request and waited out
its timeout against a mesh that answered.
**The reply address travels in the payload, and that is now proved from both ends.** The
controller reads it from there (3.4) and the host writes it there and waits on it, and the
test asserts the transport's own reply field held the *consumer's ack address* by the time the
request arrived — so a future server that stopped claiming that field fails a test rather than
letting the reason quietly become folklore.
The host's **"newest wins" window narrows at the rollout rather than disappearing**, and that
is now measured rather than predicted: three declarations pushed to an absent node leave one
on the stream and it is the newest, so the catch-up half is the stream's — but three pushes to
a connected node are still three deliveries, which is the half that stays.
**The pin turned out easier here than in the tool runtime, not harder.** The Go client takes a
`*tls.Config`, so the same pinned configuration with the same verify callback does the work;
the subject-alternative-name constraint recorded under 3.6 is that client's, because it takes
PEM strings with no verify hook. A host checks the fingerprint and nothing else.
Nothing here composes an enrolment user per live token, and that is **1.7's**, not this
task's: it is one input to a composition that does not happen at all yet.
- [x] 3.6 the tool runtime's client on NATS, behind the unchanged sdk contract — round-tripped
against a real server: a tool answered across two connections, a throwing handler reaching
the caller as an error rather than a timeout, an event delivered once with its key, body,
node and event id intact. Ships beside the AMQP client and is selected at the rollout,
because steps 1 to 4 leave every node on AMQP.
**A constraint it surfaced, recorded where somebody issuing a certificate will look.** The
AMQP client pinned the exact certificate and switched hostname verification off, which is
sound because a fingerprint is stronger than a name. The NATS client exposes no equivalent
hook — its TLS options are PEM strings with no verify callback — so the pin still happens
before dialling and the library's own name check happens beside it. **The bus's certificate
must carry a subject-alternative name matching the address nodes dial it by**, or the
connection is refused by a library error rather than by anything the mesh says.
- [x] 3.7 the sdk's three stale comments, and nothing else in it — three lines, which is the
whole of the sdk's diff for the bus change, and the measurement that predicted it
- [x] 3.8 **the declaration model** of [design 29](32-what-a-module-declares.md): local names
derived to subjects, the three namespaces, permissions computed from a declaration, and a
manifest that contains no subject. Done in the controller's composer (permissions, streams,
consumers), in the runtime's client (subjects derived from the credential, never named by a
module), and as a catalogue test asserting all 72 manifests hold no subject — because the
rule held by construction, and a rule held by construction is one a later field breaks
quietly.
- [x] 3.9 **seats declared by modules** — the manifest now carries `seats` (name, scope,
accepts/emits/serves, retention) and `uses`, and registration refuses a `mesh-*` name, a
duplicate declarer, an undeclared `uses` or claim, a seat with no protocol, a scope
mismatch, and a holder that does not answer what its seat promises. **Still to do**:
creating a seat's streams at registration and its holder's work-queue consumer at
assignment, which need the JetStream client wired in.
The refusal for an unknown claim *moved* rather than disappeared — the parser cannot judge
it from one manifest any more, because another module may legitimately declare that seat,
so it is registration's. The test that encoded the old rule was rewritten rather than
deleted, and a second one pins the case the parser could not distinguish.
**Done**: a seat's work queue is derived and created, and a holder's worker with it. The
JetStream client behind them is wired and verified against a running server, which also
completes 1.4's missing half — the pure `Asserter` had no implementation until now.
- [x] 3.10 **the ten seat renames** — done in the controller's table, the ten manifests that
claim them, the controller's own shipped manifests, and every test. Not a migration after
all: a holding is derived at resolution, never stored, so nothing recorded points at an old
name (recorded as a progressive insight on ADR 0126). A **kept** rename table tells a
manifest written against an old name what it became, because a module lives in its own
repository and may be registered long after the catalogue stopped using one.
**A seat and the interface it delivers are different names.** The `git` seat became
`mesh-git` while the `git` *provision* it delivers did not change, and the same for the
package registry. A blanket replace got this wrong first and the failure read "the package
registry is served on `<nil>`", which does not say "you renamed an interface" — so a test
now pins every seat against the interface it delivers.
**Done when.** The fixtures are produced and consumed byte for byte by every implementation that
claims the capability, and a module built before any of this serves its tools unchanged on the new
runtime. **The step is not done when the code runs** — two implementations that disagree about an
envelope do not fail to compile, they ignore each other while both keep running, which is the
failure ADR 0074 exists to catch.
**And the shared library gained nothing but the binding.** A new transport is when the pressure to
add conveniences is highest, and ADR 0039's rule does not bend for it: a helper that arrives with
the bus is a review failure, not a detail. Code shared among a module's own features stays in that
module.
## Step 4 — the core speaks it
**Why here.** The links exist from step 3, so the flows that are not on the bus at all can move onto
it, and the beds that need a mesh living on NATS can finally run.
> **A blocker surfaced here that is not this step's to fix.** Every event name in the catalogue is
> still written the way a routing key on the bus the mesh has is written, so the derivation design 29
> §1 specifies turns a consumer's declaration into a subject **no emitter publishes** — thirty-seven
> manifests, and one that cannot be composed at all. Nothing fails on the bus the mesh runs on
> today, where a routing key is matched literally; it fails on the first mesh raised on the new bus
> and not before, which is why wiring the controller's own subscription is what found it. Opened as
> [issue 127](../../04-ISSUES/127-a-module-event-derives-a-subject-nothing-publishes/00-report.md).
> It holds 4.2, 4.3 and the catch-up half of 4.5; the node-facing flows — enrolment, reports,
> heartbeats, a build's outcome — are unaffected, because those subjects are the mesh's own and
> derive from nothing a module declares.
- [ ] 4.1 **the full genesis bed** — a mesh raised on NATS from nothing and living on it: a node
enrols over TLS with a claimed token and the enrolment user cannot read a declaration; a push
is held while the store restarts and applies after, nothing lost or duplicated; a node that
was away gets exactly the newest declaration and refuses a replayed older one by sequence; an
upgrade rolls out to two nodes; an event dead-letters after `max-deliver`; and an enrolment
held by a `nak`-with-delay cycle still reaches the enrolling node, proving the reply travels
in the payload and not the transport field the consumer's ack has claimed. The server-enforced
permissions were proved at step 1 and are not re-proved here
— **nothing is outstanding but the bed itself.** Both links speak NATS, the composition
happens, and every claim above has a unit test or a check against a running server behind it.
What none of them can stand in for is a mesh raising itself, which is what this bed is — so this
is where the code stops and the lab starts
- [x] 4.2 a build source's change reaches the builder over the bus, and the build that follows is
the one the change asked for — **a build is work submitted to a role now**
([ADR 0129](../../02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md)). Both sides
are behind a seam with an implementation per bus, and on the bus being built one publish does
what two did: the outcome is the role's own event, so the asker matches it by the id its
request carried, the controller records it and the catalogue places it in the graph. A build
machine needs a reply queue for nothing and a grant over nobody's inbox.
Checked against a running server: the round trip; a third party on the role's event hearing
the same outcome the asker did, which is what the decision rests on; work leaving the queue
once settled, so no second machine repeats it; work submitted with no machine holding the role
**waiting rather than failing**, and being done when one arrives; and work a machine handed
back coming round again.
The outcome carries the module name, because only the manifest says what was built and one
message now has three readers. A failed build names none: it produced no module version, and
the catalogue would otherwise place something that was never made.
- [x] 4.3 an installation completes over the bus, with the same outcome as the path it replaces —
**the installer can raise it**: a foundation template that stands up the server, writes the
server's own settings and the mesh's first user list beside them, and starts a controller
reaching the new bus. What remains is running it, which is 4.1's bed.
**The mesh composes its own user list, and at genesis there is no mesh to compose one.** So the
installer carries the first — the controller's account at a well-known bootstrap password,
exactly as the store is reached at `postgres:bootstrap` and the old bus at `guest:guest`, and
rotated with them. From the controller's first composition onward the file is the controller's.
That surfaced a gap reading would not have found: the controller's own account exists before
there is a controller to mint one, so nothing recorded a hash for it and its first composition
would have left the writer out of the file it was writing — a bus nothing can connect to,
produced by the thing connected to it. It records a hash of the credential it is using, and only
when none is recorded, so a restart cannot put the bootstrap password back over a rotated one.
**The carried list and the derived one are checked against each other**, because they are two
statements of one fact and a mesh cannot be raised twice to find out they disagreed. A template
granting less than the controller derives produces a mesh that comes up, connects, and is
refused on its first act, with an authorisation error naming a subject rather than the template
that forgot it. The check earned itself at once: the composer was granting a role's whole event
branch *and* the one event it follows, and the wider grant wins — so only the submitting half of
a role is granted now, and what comes back is named exactly.
- [x] 4.4 a person's client — **the account and the program are both in.**
**The account**: a person is not a module and holds no seat, so their authority is a list of
tools (or `*` for an administrator) and nothing else. Held to four properties, each a way of
being wrong that would not announce itself: nothing but tools, so a person cannot claim a
module said something; no ack subject, because authority over a consumer that does not exist
is authority nobody audits; no ability to answer, because a person who can answer a request is
impersonating a module on a bus where anyone may serve a tool; and two people do not share an
inbox. Issued, listed and revoked by command; stating what somebody may call replaces what was
there, because a list that could only grow is a permission nobody can take back; and forgetting
somebody takes their credential with them, or it is not a revocation.
**The program**: two surfaces over one thing — a command line and an MCP server — both adapters
over the same three calls, because a second way of reaching a tool is a second thing to keep
correct. It uses the client a module's runtime uses, so what a person may do is answered by the
same permission list that answers it for a module and an audit has nothing separate to read.
Three decisions in it worth keeping. It lists what the **catalogue** has rather than what this
credential may call: somebody seeing only their own tools cannot tell "not installed" from "not
yours", and those need different people to fix them. A failed call says which of three things
happened — nobody serves it, this credential may not, or the tool was slow — because the
remedies are in three different places and without that they are one timeout and a stack trace.
And the MCP surface decides nothing: the names are the ones a person types, the schemas are the
modules' own, an answer is passed through unshaped, and a tool that fails comes back as a tool
error rather than a protocol error, because the request was well-formed and the mesh answered it.
Both surfaces are driven against a running bus, including a host's notification being answered
with nothing and an unknown method refused.
> **Design 25 §7 says "nothing is built of this before §10's bed passes", and this was built
> before.** Recorded rather than quietly ignored: the operator asked for it, it is on the
> critical path for nothing and blocked by nothing, and the bed it waits for is 4.1's. If the
> bed changes what a person's client should be, this is what gets changed.
- [x] 4.5 reports and catch-up: a node that was unreachable catches up rather than losing them.
**The reports half is in and proved against a server** (3.4): held through the store's absence by
the server rather than by the controller, superseded ones settled by the digest they carry.
**The catch-up half needed nothing built, and that was the answer.** It existed because a queue on
the bus the mesh runs on today receives only what is published after it is bound, so everything
built before the catalogue existed was announced to nobody — and on a fresh mesh that is always
the foundation, because those are the things the catalogue needed in order to exist
([issue 050](../../04-ISSUES/050-the-catalogue-knows-nothing-built-before-it/00-report.md)). A
whole mechanism followed: the catalogue asks, the controller re-publishes.
A stream is a log and a consumer is a position in it. A consumer created later starts at the
beginning, so the builds are simply there — asked of a running server rather than assumed, since
the decision rested on it: three builds published with nothing listening, then a consumer created,
and all three waiting for it. So the question of *who replays* has no answer because nothing
replays.
> **This is the shape of the whole change, in one task.** Three ways to do the replay were weighed
> — a namespace for the mesh's own voice, the controller answering a question, a consumer reading
> from the start — and the right answer was that the bus being moved to already does it. The
> mechanism was never about builds; it was about a queue that could not remember. **A conversion
> that carried it across would have carried a workaround for a limitation that no longer exists**,
> and nothing would have looked wrong.
Retiring it is step 5's, with the rest of what only the old bus needs: the request, the
re-publishing, and the `replay` flag that told a consumer to register history without acting on it.
**Done when.** Each converted flow is proved against the behaviour it replaced, and the full genesis
bed is green. **Observation is not in this step** — heartbeats, conditions and key-value state are
[research 017](../../01-RESEARCH/017-a-mesh-that-heals-itself/00-overview.md)'s, that effort already
reserves them for after the move, and a flow built ahead of its design would be rebuilt.
## Step 5 — the rollout
**Why here.** It is the only step that moves a node's bus, and it moves every node's at once.
**The order the repositories land in is part of the rollout, not paperwork.** Derived 2026-09-27
while merging, and not obvious from any one repository, which is why it is written here rather than
left to be re-derived under time pressure:
| order | repository | why it cannot be later |
|---|---|---|
| 1 | this one | prose; nothing deploys |
| 2 | the sdk | comments only, and no module rebuilds for it |
| 3 | the client library | it is what a module calls to emit, and it is where the subject is derived. Until it lands, a locally-named event is published under the local name itself |
| 4 | the catalogue | every manifest and every module's code, renamed together. Safe only once the runtime derives |
| 5 | the controller | **it refuses an old-style event name outright**, so landing it before the catalogue makes every unconverted module unregisterable |
| 6 | the hosts | last, because nothing else waits on them |
Two properties make the sequence safe rather than merely ordered, and both are pinned by tests. A
name already in the old form passes through the derivation untouched, so a module nobody has
converted keeps working at every step. And a converted name derives to **exactly** the key the old
bus published, so steps 3 and 4 change nothing on the wire — the move to the new bus is step 5.2 and
one environment variable, not a side effect of deploying.
The failure this ordering avoids is issue 127's own: a publisher and a subscriber that disagree about
a subject produce no error anywhere. Nothing logs, nothing retries, and the mesh reports itself
healthy while reacting to nothing.
- [ ] 5.1 the cutover bed: a mesh on AMQP with a predecessor stand-in on the deprecated broker
moves its bus in one rollout, every node reporting on NATS afterwards, the stand-in's own
client still connected throughout
- [x] 5.2 the rollout: accounts composed, then the controller, every host and every runtime
together; every node confirmed heard before AMQP stops.
**Done 2026-09-28, 02:25.** Every machine reports on the new bus, the seat is held by the
module that provides it, the old broker is unassigned and forgotten, and every credential was
minted afresh at the end because two had been printed on the way. What it took, in the order
it was found, each fixed on the trunk before the next step: the control plane's `serve` and
`push` never selected the new transport (task 4.3, open until then); a machine's user was
granted neither the asking nor the delivery of its own consumer; the account had no JetStream
of its own; the control plane's client verified the bus's certificate by name instead of
pinning it; the seat table's rows carried no protocol, so no role's work queue was raised; the
build machine decided its bus from a variable its container never received; and a rotation
put new hashes on the bus before three machines had received their new memberships — which
is why there is now `rollout hand <node>` and a host adopts a delivered membership at start.
The bootstrap loop — a bus that can only be raised by a declaration that can only arrive
over that bus — was broken once, by hand: the mesh's own composed configuration started the
server, and the controller binary was run on the node directly until the managed container
could be rebuilt over the bus it was on.
- [x] 5.3 **the seat changes hands as one act.** A command takes a seat and the assignment taking it
over, and the seat is never empty in between — the emptiness is the outage of 2026-09-27, when
the control plane, which finds its own bus through this seat, lost the address and looped.
**Built 2026-09-27** (`seat_holding`, migration 0039; design 26 says how it is checked), and used
live the next night to hand `mesh-broker` from the old broker's assignment to the new one's. This
is what 5.2 uses to move `mesh-broker` from the old
broker's assignment to the new one's, and it is built first ([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)).
- [x] 5.4 **the old broker and everything that named AMQP leave the mesh** ([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md),
superseding [ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md)): the two modules that
required `amqp` are removed, the broker's module is unassigned and removed (**done 2026-09-28**; the predecessor's own tooling, which rode the same adopted broker, went dark with it, as [ADR 0130](../../02-DECISIONS/0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md) accepted), registration refuses
a manifest that provides or requires `amqp`, and a whole-catalogue check asserts none does. Not
a retirement condition — a decision, taken, with the operator's "I don't care if the predecessor
breaks" on record ([ADR 0130](../../02-DECISIONS/0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md)).
Retiring with it: the build outcome's second announcement under the module's own name, which
existed only so a catalogue deployed before the rename and one after both heard it.
> **The remote tooling goes with it too.** The predecessor's own mesh talks over that broker, so
> shutting it down ends the path that reaches this installation's machines from a workstation.
> The rollout is driven from the node, or before the broker stops — a sequencing constraint on
> 5.2, not an afterthought.
- [x] 5.5 **the AMQP transport is deleted from the control plane and the hosts**. One bus, nothing
to select ([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)).
**Done 2026-09-28.** The control plane's old consume loop, build request, tool call, management
API and account scoping went, and the host's old dialling and enrolment paths with them; a
membership or token naming any other bus is refused before anything is sent. Nothing selects a
transport any more: the variable that once did (`MESH_BUS_NATS`) now only names where the
control plane reads its own bus credential, the way any module reads a secret. **Checked by the
build**: neither repository's module file names the AMQP client library, so a line that still
used it would not compile. The store-window guarantee ([issue 083](../../04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/00-report.md))
is tested against a bus-less fake rather than the old transport's memory, which is what let
that memory go — the one thing it did that the stream does not (superseding a held report) is
the staleness check on the message itself (design 25 §3).
Found on the way: **no build had ever recorded what it stood on.** A recipe reads its base from
a build argument, so the digest was never in the file the builder derived edges from, and every
order that says *bases first* — `build --on`, `build --behind`, the merge follow-up of
[issue 131](../../04-ISSUES/131-nothing-tells-the-mesh-a-source-moved/00-report.md) — walked a
graph with no edges. The builder now reports the bases it was handed, the control plane records
them by artifact path, and the graph is read from the newest build of each module — a recorded
manifest carries no `build.on`, so the edge is derived from the build or it does not exist.
> **The old 5.4 note is history.** It recorded that a retirement *condition* was wrong from
> [ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md) onward, which framed the old broker
> as an ordinary provider with no end. [ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md) ends that
> framing in turn: the broker is not kept as a provider either, because AMQP is not a provision. Both
> readings are kept here so the two reversals can be read in order.
**Done when.** Every node reports on NATS, and nothing of the mesh's own is left connected to the
deprecated broker.
## The through-line
The order is dependency, not preference. **Steps 1 to 4 leave every node on AMQP**, so the cost of
being wrong is bounded until the last step: a step may be abandoned, or reordered after step 2,
without a rollback. The server stands before anything speaks to it; the wire is specified before it
is implemented three times; the flows move once there is something to move them onto; and the bus
itself moves once, at the end, on one day.
## What is deliberately not here
- **Observation** — research 017's, after the move, by its own design.
- **Leaf nodes** — design 25 §11 keeps this out of scope and says so; a leaf per machine is a later
question, noted so it is not forgotten.
- **The predecessor's world.** It is AMQP and it is not moving —
[ADR 0130](../../02-DECISIONS/0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md): it is
deprecated, some of it is still running, and it is being left to stop rather than migrated. Its
broker goes with it, unassigned like any provider whose provision nothing requires.
## How this list is kept true
A task is ticked when its change is committed, not when it is written. A step is done when its bed
is green, not when its tasks are ticked. If a step's tasks are all ticked and its bed has not run,
the step is **in progress** and this document says so — that gap is the thing the whole shape is
built to make visible.
@@ -0,0 +1,178 @@
---
layer: to-be
status: proposed
code: []
updated: 2026-09-27
decisions:
- 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md
- 02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md
- 02-DECISIONS/0051-shared-data-is-the-operators.md
- 02-DECISIONS/0113-the-vault-makes-every-secret.md
- 02-DECISIONS/0117-a-machines-uplink-is-a-seat.md
---
# 29 — A node has operator accounts, and the mesh owns what lives under a home
**The mesh models machines but not the people on them.** A node record holds its name, its
address, its mode — and nothing about *who a person is* on it: `jochens` on novox, `ace` on ace,
`jochen` on shanks and g14. That username is not incidental. It decides who a file under `~` is
owned by, who a user service runs as, and — the case that surfaced this — which account `ssh
<node>` logs in as. The predecessor knew it (its per-node `user:`, and the modules that wrote a
person's `~/.ssh/config`, `~/.zshrc`, `~/.config`); the mesh, taking those over, kept the machine
facts and dropped the human one.
Several things are missing, and they are one idea.
## 1. The account is a node fact
A node has one or more **operator accounts**: the human logins on it. At minimum a name; the
mesh already knows the node and its address, so `<account>@<node>` is then a complete answer to
"who am I, where." It is the mesh's to hold because everything below is derived from it, and
because it is exactly the fact that was silently lost — `ssh ace` failed to `ace` because nothing
in the mesh said ace's account is `ace`.
## 2. A resource may live under a home, owned by its account
[ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md) placed a
module's *system* data — `<root>/<module>`, owned by the module. It has no analog for the other
half of the filesystem: the things that belong under a person's home and are owned by that
person. `~/.ssh/config`, `~/.zshrc`, `~/.config/hal` — every one of these is a resource the mesh
should be able to place and own, resolved against **the account's home** rather than a system
root, and chowned to **the account** rather than to root or a module uid.
This is the same move as `${dir:…}`, one level over: a resource says `home: <account>` (or names
an account requirement), and the mesh resolves the home directory and the owning uid on the node
that account lives on. A module that writes operator config — the eventual replacements for
`hal/terminal`, `hal/claude-code`, `hal/secrets` — declares its files this way and names no
`/home/...` path, exactly as a system module now names no `/var/lib` path.
These are a **family**, not one module: an `ssh-client` module, a shell module, a `~/.config`
module, each a *universal-tier* consumer of the account fact — assigned wherever a person logs in,
which is every node, unlike the graphical stack that a capability gates.
## 3. The whole of `~/.ssh` is the mesh's — with one boundary drawn inside it
The predecessor owned a single file (`~/.ssh/config`) and left the rest alone; it drifted, because
owning one file beside foreign ones is not owning anything. The mesh should own **the directory**:
create `~/.ssh` at `0700`, chown it to the account, and own the files it places there —
- **`config`** (or the mesh's region of it): the `Host` blocks for every other node, composed
from the roster;
- **`known_hosts`**: authoritative, so the "Host key verification failed / accept-new" dance that
cost real time during enrolment simply ends;
- **`authorized_keys`**: who may log into this account, governed centrally rather than by whichever
key happened to be pasted where.
**The boundary — and it is the reason this is safe:** `~/.ssh` is the one directory where a wrong
declaration locks a person out of their own machine. So the mesh's *found-vs-owned* semantics
([ADR 0126](../../02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md),
adoption) apply *inside* the home directory. The mesh **owns** the directory and the files above; it
**holds as found — never rewrites, never removes** — the operator's own contents: their **private
keys** and their **personal drop-ins** (`config.d/personal`, the personal `Host` aliases a
workstation carries, exactly as `hosts.local` is the home the mesh never rewrites for `/etc/hosts`).
Reconcile removing an unassigned `config.d/mesh` is fine; the same logic aimed at `id_ed25519` or an
operator's own `authorized_keys` entry is a lockout. This is the login-channel cousin of the rule
[ADR 0125](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md) draws for the uplink and the sshd
module draws for the firewall: **the mesh must never be able to arrange the one failure that severs
its own way back in.** The carve-out is not a convenience; it is that rule, in `~/.ssh`.
## 4. Keys are the mesh's to generate — through a CA, and existing keys are adopted, not replaced
Key *generation* is the mesh's, not each node's improvising its own. The clean form is an **SSH
certificate authority as a seat**, the sibling of the TLS internal CA the mesh already runs:
- **Host certs.** The mesh signs each node's host key. Every node's `known_hosts` becomes one line
— `@cert-authority *.<suffix> <mesh-CA-key>` — and nothing is distributed per node; a new node is
trusted the instant its host key is signed.
- **User certs.** The mesh signs a cert naming the principals (accounts) allowed. Every node's
`authorized_keys` / sshd `TrustedUserCAKeys` becomes one trust line — no N×N key spraying — and
short-lived certs give rotation for free
([ADR 0114](../../02-DECISIONS/0114-a-shared-credential-rotates-over-two-credentials.md)).
- The **CA private key is the mesh's**, a secret the vault makes
([ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md)).
**Three kinds of key, and only one is never minted.** Host keys (server identity) and pure
machine-to-machine keys the mesh may generate end to end. The operator's **personal** private key —
possibly on a hardware token, possibly used from an off-mesh laptop — the mesh **signs into a cert
but never generates**; that, and only that, is the residue of
[ADR 0051](../../02-DECISIONS/0051-shared-data-is-the-operators.md). So "keys are mesh-owned" and
"the operator's login key is the operator's" reconcile: the mesh owns the CA and the signing; it
holds the human's private half, never mints it.
**Existing keys are not lost.** Taking ownership is *adoption*, not regeneration: a key already on a
machine is recorded and signed, not overwritten. The mesh gains authority over `~/.ssh` — it does
not clear it. An enrolling node's host key and the operator's existing key are carried forward; the
found-vs-owned boundary of §3 is exactly what guarantees nothing already there is destroyed.
## 5. How it is distributed: the controller composes, the node applies
None of this needs a node to discover the mesh, and none of it needs a control-plane module of its
own. The ssh files are **roster facts**
([ADR 0128](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md)): once the
roster view carries a node's **host key** and its **account** beside its name and address, the
`ssh-client` module ships a template for `known_hosts`, `config` and `authorized_keys`, and the
controller renders each node's copy from the full roster and pushes it. The mesh owns the data; the
module owns ssh's format; the control plane gains no ssh syntax. It is the same act as composing a
peer list or `/etc/hosts` — which is why there is **no novox-only "mesh-ssh" module**: the
centralization is the controller's composition, not a module that runs somewhere. Only non-secret
facts travel (names, addresses, accounts, host keys, the CA public key); the private key stays the
operator's, placed as an operator-owned file, referenced by path.
## 6. The two modules, and the seat between them
- **`sshd`** (server, every node) — manages sshd, owns and **reports** its host key so the roster
carries it, and trusts the user CA.
- **`ssh-client`** (client, every node) — owns `~/.ssh` per §3, consumes the roster and the CA
public key.
- **`the-ssh-ca`** (a seat, held on the control node) — signs host and user certs.
They meet at the account and the CA, not at a bespoke module. The `sshd` server side already exists;
the client/identity side and the CA are the open pieces.
## Why now, and why not yet
**Why it matters:** when HAL retires, the generators that keep `~/.ssh`, shell config and the
operator's `~/.config/hal` current retire with it. Without this, adding a node stops adding its ssh
alias and its trust, and a fresh machine has no operator dotfiles at all — the mesh would run every
service and leave the human unable to work on the box.
**Why not build it reflexively:** it is a real addition to the node model, the resource model, and
the seat set, and must be gotten right. The mechanism half is now settled — ADR 0128 is what lets
the ssh files be templates with no control-plane format — so what remains to decide here is the
model:
- **One account or several per node?** A workstation has one human; a shared box might have more.
Allow more than one without forcing the common case to name it.
- **The CA's shape.** Host-cert and user-cert principals, cert lifetime and renewal, where the CA
runs (a seat on the control node). The one thing fixed: the operator's personal key is signed,
never minted.
- **Adoption of existing keys.** How an enrolling node's host key and an operator's existing key are
recorded and signed rather than replaced — the found-vs-owned boundary, made concrete for keys.
- **The `sshd` boundary.** Server side exists; this is the client, the identity, and the CA.
- **The ssh-agent.** An agent is a *user-scoped service running as the account* — the first concrete
case of the user services §2 anticipates. It holds the operator's private key in memory; the mesh
declares the unit and sets `AddKeysToAgent`/`IdentityAgent` in `config`, and still never sees the
private half. Agent *forwarding* wants a policy, not a default: with user certs it is largely
unnecessary, and forwarding an agent into a node exposes the operator's keys to that node's root —
so prefer certificates and `ProxyJump` over forwarding.
**Not urgent, not blocking.** ssh and dotfiles work today because HAL's generators still run as the
substrate. This becomes load-bearing in the node-by-node retirement phase, not before — which is the
right time to build it, once the account and CA model are decided here.
## References
- The gap was found generating `~/.ssh/config` from the *HAL* registry (`hal/terminal`'s
postConfigure hook), which the nox mesh has no equivalent for.
- [ADR 0128](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md) — the roster
fact mechanism that renders the ssh files, format owned by the module.
- [ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md) — the
system-path placement this mirrors for home paths.
- [ADR 0051](../../02-DECISIONS/0051-shared-data-is-the-operators.md) — why the operator's personal
key is signed, never minted.
- [ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md) — the CA key is a secret the
vault makes; [ADR 0114](../../02-DECISIONS/0114-a-shared-credential-rotates-over-two-credentials.md)
— short-lived certs as rotation.
- [ADR 0125](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md),
[ADR 0126](../../02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md) —
the never-sever-the-channel rule and the found-vs-owned semantics, applied here to `~/.ssh`.
@@ -0,0 +1,137 @@
---
layer: to-be
status: proposed
code: []
updated: 2026-09-27
decisions:
- 02-DECISIONS/0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md
- 02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md
---
# 30 — The mesh updates itself on a push
**Today the mesh does not update itself; a person drives the pipeline by hand, and one class of
change freezes it.** A code change lands in `mesh-controller` or `mesh-catalog`, and getting it onto
the machines is a sequence somebody types. The predecessor's pipelines rebuilt and redeployed on a
push without anyone watching; the successor should too. This records the process as it is done by
hand now — so it can be read, and then coded — and the two things that make it more than "add a
webhook".
## The process, as done by hand
**An ordinary (non-breaking) change** — new module code, a bug fix, a manifest tweak that changes no
seat or schema:
1. `module moved <module> <commit>` — tell the mesh its source advanced (the controller repo has no
trigger, so this is manual; the catalogue's webhook does it automatically — see below).
2. `build --behind` (or `build <repo> [--ref] [--path <subdir>]`) — the build machine rebuilds and
records the new image.
3. The mesh **reconciles on its own**: the module's declaration now names the new image, the next
push/heartbeat sends it, and the host swaps the container. For the control plane this is a
self-upgrade — the running controller composes its own new image and the host replaces it. No
restart is typed.
**A breaking change** — a manifest schema the controller parses differently (a fact's shape, a
seat's name), where the new control plane cannot read the manifests the old one stored:
4. Land the code (controller + catalogue together — they are one change).
5. Rebuild + deploy the new controller (steps 1–3). **The moment it is live it refuses the
still-old-shape stored manifests, and composition freezes for every node that runs an affected
module.** Running services are untouched; only new declarations stop.
6. **Re-register each affected manifest under the new shape**, which the *new* controller accepts —
`module add <file> -source <repo> -ref <ref> -commit <commit>`. This writes the manifest to the
store without a build, so it is the fast way to lift the freeze. (The controller container is
distroless: `docker cp` the file to the container root `/x.json`; `/tmp` does not exist; the
root filesystem is writable. The file is lost when the container is recreated on the next image
swap, so copy it *after* the swap.)
7. `push --behind`, then verify `status` is clean and `seats` (or the relevant surface) shows the
new shape held by the right holders.
The freeze in a breaking change has been paid three times in one session (a fact-shape change, the
`/etc/hosts` region, a seat rename); each time it lasted seconds and no service dropped. It is
recoverable, but it is not something a push should trigger unwatched — which is the crux of what
automating this must solve.
## Why it is more than "add a webhook"
### 1. The trigger today is HAL's, not the mesh's
Build-on-push works for the catalogue because its repository has a Gitea webhook pointing at
`http://host.docker.internal:9877/webhook/gitea` — and **that receiver is `hal-gitea-tools.service`**
(`~/.hal/modules/hal/gitea/tools/server.js`), a *predecessor* component. The nox builder consumes
build work; it does not receive Git events. So the mesh's own build pipeline currently rides on a
HAL service, and:
- the `mesh-controller` repository was never wired to it, which is why the control plane is the one
thing that does **not** self-update — every controller deploy this session was `module moved` +
`build` by hand;
- when HAL is retired, build-on-push stops for the whole mesh.
**The mesh needs its own forge-webhook→build trigger**, a nox component (a module, and likely a
seat — `mesh-forge-trigger` or folded into the git seat's holder) that receives Git events and turns
them into build work over the broker, for **every** repository including `mesh-controller`. Replacing
`hal-gitea-tools` is the concrete first build. Its logic already exists to copy: match the pushed
repository (and changed paths, for a monorepo like the catalogue) against the build-context
repository of every registered module, and rebuild the matches.
### 2. The builder validates too — and a breaking change deadlocks it
The build machine embeds the same catalogue package the controller does, so **it validates a
manifest against its own compiled-in seat/schema set**. A breaking change therefore couples *four*
things, not two: the controller, the **builder**, every affected manifest, and every node's host.
This session's seat rename rebuilt the controller but not the builder, and the stale builder then
refused every manifest claiming a renamed seat.
Worse, one rename **deadlocked** the builder: the build machine's own seat was renamed
(`the-build-machine` → `mesh-build-machine`). To refresh the builder you must build it; to build it
the *running* (old) builder must accept the new builder's manifest — which claims the new name it
does not know. The old builder cannot build the new builder. Escapes:
- **Never rename a seat whose holder validates manifests** in an ordinary pass — the build machine's
seat belongs with the deferred delivering seats (ADR 0121). Reverting `mesh-build-machine` to
`the-build-machine` (deferred) lets the old builder build the new builder, which then knows the
new names.
- Or bootstrap a new builder image **out of band** (build locally, publish to the registry, register
the module at that digest), the way genesis loads the first builder — bypassing the old builder's
validation once.
Either way, self-update for breaking changes needs a **transition discipline** so a push does not
auto-freeze: the new control plane (and builder) should accept the *old and new* shape together for
one release — deprecated aliases in the seat set, a schema that reads both — then a later release
drops the old. With that, a breaking change rolls out on a push like any other: everything reads
both, the manifests migrate, the compatibility is removed. Without it, self-update would simply
automate the freeze.
## What to build
- **A nox forge-webhook trigger** (replaces `hal-gitea-tools`): receives Git events for every mesh
repository, dispatches build work to the builder over the broker, and records `module moved`
automatically. Wire `mesh-controller` to it so the control plane self-updates like everything else.
- **A transition discipline for breaking changes**: the control plane and builder accept old+new for
one release; the tooling that lands a schema/seat change emits the compatibility shim and the
follow-up that removes it. This is what makes step 4–7 above safe to trigger unwatched.
- **Config/package modules need no builder** — `module add` registers their manifest directly
(this is how the uplink managers and the re-registrations above were done). Only image-bearing
modules need the build machine, which narrows what the deadlock above can block.
## Why now, and why not yet
**Why it matters:** self-update is the difference between a mesh a person maintains by typing
pipeline steps and one that maintains itself, and it is a stated goal (parity with the predecessor's
pipelines). The HAL trigger dependency also makes it a retirement blocker: build-on-push dies with
HAL.
**Why not reflexively:** the trigger is a new component with the broker and forge in its blast
radius, and the transition discipline changes how every breaking change is written. Both should be
designed, not bolted on beside a freeze. The manual process above is the interim, and it works.
## References
- [ADR 0121](../../02-DECISIONS/0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md)
— the seat rename whose migration and builder deadlock this record is drawn from
- [ADR 0128](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md) — the
fact-shape change that first showed the breaking-change freeze
- `hal-gitea-tools.service` (`~/.hal/modules/hal/gitea/tools/server.js`) — the predecessor webhook
receiver on `:9877` the mesh currently rides on
- mesh-controller `cmd/mesh-builder` (the build machine), `internal/catalogue` (the seat/schema
validation the builder shares with the controller)
@@ -0,0 +1,65 @@
---
layer: to-be
status: proposed
code: []
updated: 2026-09-27
decisions:
- 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md
---
# 31 — A module declares its fail2ban jail, and the mesh composes them per node
**A node's intrusion filter should be composed from the modules it runs, the same way its firewall
is.** The mesh already derives a node's nftables ruleset from every assigned module's `listens` and
`guards` (the `Filtering` mechanism). fail2ban is the same shape and is not modelled: a module that
runs an authenticating service — postgres, mssql, mailu — has a jail (a filter that reads its log
and a jail stanza that bans on it), and which jails a node's fail2ban runs should be exactly the
jails of the modules assigned to that node.
The predecessor did this with per-module files: `postgres` shipped `postgres-auth.conf`, `mssql`
shipped `mssql-auth.conf`, `mailu` shipped `mailu.conf`, and the node's fail2ban read whichever were
present. When HAL retired on novox those became dangling symlinks — fail2ban ran the jails only from
memory, and a restart would have dropped them. The base was salvaged (the fail2ban module now ships
`sshd`, `recidive`, and the `ignoreip` that spares the mesh's own range), but the **service jails
are gone**, because no nox module declares one yet.
## The shape
- **A module declares its jail in its manifest**, naming no node and no path (ADR 0112): the filter
(the failregex, or a stock filter it uses) and the jail stanza (port, logpath, maxretry, bantime).
The `postgres` module says what a postgres brute-force looks like and how to ban it; it does not
say on which machine, because it does not know.
- **The mesh composes them per node.** For each node, the jails of its assigned modules are gathered
and written into the fail2ban holder's `jail.d/` (and filters into `filter.d/`), exactly as
`listens`/`guards` are gathered into the node's firewall. So a node running postgres gets the
postgres jail; a node not running it does not. The `node-intrusion-prevention` holder receives
them the way a provider receives its consumers' contributions.
- **The base stays the fail2ban module's**: `sshd`, `recidive`, and the `ignoreip` naming
`${machine:mesh-range}` so a tunnel peer is never banned.
## Why this, and not the module writing the file itself
A module could declare a `file` resource at `/etc/fail2ban/jail.d/<x>.conf` directly. Rejected: the
path is the fail2ban holder's to own (one module owns `jail.d`, as one module owns the firewall
table), the jail's logpath and defaults want the mesh's composition (the `ignoreip`, the ban action
the node uses), and two modules writing into one directory is the collision the seat/holder model
exists to prevent. The module declares *what its jail is*; the holder's composition decides *how it
lands* — the same split as `listens` (the module says the port; the mesh says the rule).
## Why now
fail2ban on novox currently runs the service jails from memory only; the next restart drops them
(the `ignoreip` is safe on disk, so the mesh-partition risk is closed, but postgres/mssql/mailu
auth-banning would be lost). This is the mechanism that restores them properly, and it is needed as
each of those modules migrates to the other nodes — ace running postgres should get the postgres
jail, composed from the postgres module's manifest, without anyone editing a node.
## References
- [ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md) — a module
names no node or path; its jail is declared the same way its `listens` are
- mesh-controller `internal/catalogue/adoption.go` (`Filtering` — the firewall composition this
mirrors), `internal/catalogue/manifest.go` (`Listens`/`Guards`, the fields a jail field sits
beside)
- mesh-catalog `modules/fail2ban` (the base: sshd, recidive, ignoreip); the service modules
(`postgres`, `mssql`, `mailu`) that will declare jails
@@ -0,0 +1,461 @@
---
layer: to-be
status: proposed
code: []
updated: 2026-09-27
decisions:
- 02-DECISIONS/0126-a-module-declares-its-own-seats.md
- 02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md
- 02-DECISIONS/0128-the-mesh-bus-is-required-not-ambient.md
- 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0041-events-are-a-relationship.md
- 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md
- 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
- 02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md
---
# 32. What a module declares, and what the bus makes of it
**A module that speaks to the mesh requires the bus, and receives what it needs to connect**
([ADR 0128](../../02-DECISIONS/0128-the-mesh-bus-is-required-not-ambient.md)). What a module
declares are *relationships*; subjects, streams, consumers and permissions are all derived from
those, and a manifest never contains one.
> **Revised 2026-09-26.** This document opened by calling the bus *ambient* — "no module requires
> it, the way no module requires a filesystem". Two counts say otherwise: of 72 modules in the
> catalogue, **49 take a broker credential and 23 do not**, so an ambient connection would mint an
> account for a third of the catalogue that never speaks; and the 49 each hand-write the path it
> lands at, which is provisioning done badly by hand. The bus is required, and a module that does
> not require it has no account at all.
**The requirement delivers the connection; the declarations shape the authority.** `requires:
mesh-bus` says *this module talks to the mesh* and grants no subject by itself. `emits`,
`consumes`, `tools`, `uses` and a declared seat say what it may say and hear. Declaring a subject
without requiring the bus is incoherent and refused at registration.
This document is the declaration model. [Design 25](25-the-bus-on-nats.md) is the bus itself —
subjects, streams, accounts, enrolment — and stays the authority on the wire.
[Design 19](19-the-module-protocol.md) is the specification an SDK implements, and is rewritten
onto this in step 3 of [ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md).
## 1. A module names locally; the mesh derives the subject
This is the load-bearing rule.
[ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md) says a module
definition names no node, mesh or path. A transport address is the same class of thing: if
manifests held literal subjects, reorganising the subject space would mean editing every module in
the catalogue, and the mesh would have hundreds of copies of a decision it made once.
| declared | derived |
|---|---|
| `emits: order.placed` | publish on `mesh.mod.<module>.event.order.placed` |
| `consumes: billing.order.placed` | durable consumer on `mesh.mod.billing.event.order.placed` |
| `tools: status` | queue-group subscription on `mesh.mod.<module>.tool.status` |
| seat `telegram-sender`, `accepts: send` | work-queue consumer on `mesh.seat.telegram-sender.accept.send` |
| `uses: telegram-sender` | publish on that seat's `accept` subjects, and nothing else |
**Wildcards, and they are the mesh's rather than a bus's.** *Added 2026-09-27, from
[issue 127](../../04-ISSUES/127-a-module-event-derives-a-subject-nothing-publishes/00-report.md).* A
consumer may write `*` for one name and `**` for the rest: `*.download.completed` is that event from
any module, and `**` on its own is every event in the mesh, which an audit logger wants and says in
one token. Spelled this way rather than the wire's, for the reason everything else here is local —
the bus the mesh runs on today spells these `*` and `#`, the one being built spells them `*` and
`>`, and a manifest naming either would stop being true when the wire changed. An emitted event
carries no wildcard: it names one event.
**A module publishes under its own name, and an event about a role belongs to the seat.** *Added
2026-09-27, same source.* The bus enforces that a namespace belongs to the module it is named for, so
an event named for somebody else cannot be published at all. Where the event is really about a role —
"the artifact store accepted an image" — the seat is the right home, because that name outlives
whoever fills it, and a consumer written against the holder's own name breaks when the holder
changes. **Not yet possible in practice**: seats carry protocol in the manifest and in the permission
model, and the shared library has no way for a module to publish on one. Until it does, such an event
lives under the emitting module's own name and the consumer carries that coupling.
**How the rule is checked, because it was not.** *Added 2026-09-27, same source.* Two checks, because
the mistake happens at two scales. Per manifest, at registration: an event is a local name, and the
old bus's form is refused with the name to write instead. Across the whole catalogue, as a test:
where a consumed event's emitter is present, it must emit that event. The second cannot demand a live
emitter for everything — a module lives in its own repository and may be installed long before the
one whose events it wants — so it says nothing about an absent emitter and everything about a present
one. **A subscription that matches nothing is not an error, it is silence**, which is why nothing
reported thirty-seven manifests being wrong the same way.
**It is `tools:`, not `serves:`.** Revision, found while implementing: the manifest already uses
`serves` for the facts a consumer needs in order to reach a provision, and two meanings under one
key in the file a module author reads most is a footgun. Worth noting that until now a module's
tools were not declared at all — they were known only at runtime, from an environment variable in
its image — so declaring them is new, and is what lets the mesh check that a module claiming a
seat answers what that seat's protocol promises.
**The `event` / `tool` / `accept` token is load-bearing, not decoration.** Revision, found while
defining the streams: a stream is defined by a subject filter, so a namespace holding both a
module's events and its tool calls cannot be filtered into an events stream without capturing
every tool invocation in the mesh — and a tool call must never be persisted
([design 25](25-the-bus-on-nats.md) §3 keeps tools on core NATS, where a lost call is a timeout the
caller already handles). The kind token is what makes `mesh.mod.*.event.>` a safe filter. The
first draft of this table had no token, which reads better and cannot be implemented.
**The test this must pass: the manifest survives the wire changing.** Reorganise the subject space
and every manifest in the catalogue is still correct. That is the property
[ADR 0039](../../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md) gave the sdk, applied to
declarations.
**A handler names an event the way its manifest does.** *Added 2026-09-27, found while fixing issue
127.* The runtime handed a handler the event name alone, so a manifest declaring
`consumes: builder.built` produced a pattern that could never match the key it was compared against,
and a module consuming one event from two emitters could tell them apart only by reading a header.
The subject already carries the emitter, so the key a module sees names it too — which makes a
disagreement between a manifest and the code a typo rather than a category error.
## 2. Three namespaces, and nothing else
**Its own** — `mesh.mod.<module>.>`. Its events and its tools. Nothing else may publish into it,
so an event's source is a fact the bus enforces rather than a claim in the body.
**Seats it holds** — `mesh.seat.<seat>.>`. Full participation: consume what the seat accepts,
publish what it emits, serve what it serves.
**Seats it uses** — publish only, and only on the `accepts` half. A sender cannot subscribe to a
seat's inbound subject and watch other modules' traffic, and cannot publish the seat's outbound
events and lie about outcomes.
A module naming anything outside these three is refused at registration. The whole permission set
is derivable from the declaration; nobody writes an access rule.
## 3. Queues are derived, never declared
A module says what it reacts to, not how delivery works. Each `consumes` becomes one durable
consumer; a seat's `accepts` becomes one work-queue consumer with a queue group named for the
seat. The module does not name them, does not know their names, and cannot misconfigure them —
and the controller stays the only writer of stream and consumer definitions
([design 25](25-the-bus-on-nats.md) §3).
**The mesh's own seats carry protocol too.** *Added 2026-09-27,
[ADR 0129](../../02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md).* A seat declared by a
module says what it accepts, emits and serves; the `mesh-*` set said only who does a job. So the mesh
had roles it could not describe — a build machine with three audiences for one outcome and no way to
derive a grant for any of them, and an event genuinely about a role with nowhere to live but the
namespace of whichever module happens to hold it. The mesh's seats now take the same three fields, and
the same machinery derives the holder's authority, its work queue and its consumers.
**So a build is work submitted to a role, like any other.** The build machine seat accepts a build and
emits an outcome, and the dedicated branch that carried builds retires: a work queue shared by several
machines is exactly what `accepts` already is, and a second mechanism for it is two places a permission
can be wrong.
**And one publish reaches three audiences without anybody's inbox being opened.** A build's outcome is
the seat's own event: whoever asked matches it by the id their request carried, the controller records
it, the catalogue places it in the graph. That is the fan-out a shared exchange gave for free, written
as a subject the mesh derived instead of a topology somebody configured — and it is why a holder needs
no permission to publish into an asker's inbox, which is the one grant design 25 §4 refuses by name.
**Retention belongs to whoever owns the namespace, not to a consumer.** A seat declares how long
its inbound backlog survives, because that is a property of the service:
```
seat: telegram-sender
scope: mesh
accepts: send retain 7d
emits: delivered, failed
serves: status
```
If each consumer could tune it, the mesh's durability would be an emergent property of whichever
manifest was edited last.
**Why a module's own events do not carry their own retention, though the same rule would allow
it.** A seat owns its namespace and gets a stream of its own, so it can say. A module's events
share one `EVENTS` stream, and three facts about JetStream decide that they must:
- **Storage is not a property of a subject.** A subject is only an address; a *stream* is a
separate object that captures subjects matching a filter. So "this topic is durable" is always
really "some stream covers it", and something has to create that stream.
- **Overlapping streams are refused, not merged.** Verified against the server: a per-module
stream beside a shared `mesh.mod.*.event.>` is rejected with *subjects overlap with an existing
stream*. So "one stream by default, its own for a module that wants different retention" is not
available — it is all of one or all of the other, and a filter cannot express an exception
either.
- **A stream per module breaks cross-module consumption.** An audit logger consuming every
module's events is one consumer on one stream today; with a stream each it becomes one consumer
per module, created and destroyed as modules come and go.
So: one stream, and **per-subject caps** for the fairness that actually matters — a noisy emitter
cannot evict a quiet one, which is verified (a cap of three, ten messages on one subject and one
on another, leaves four). What is genuinely unavailable is a different *age* per module, because
JetStream ages per stream and not per subject. A module that truly needs its own retention has a
way to say so: declare a seat, which owns its namespace and gets its own stream.
**Scope gives per-node workers without a new concept.** A module running on three nodes that each
need their own queue declares a node-scoped seat: one holder per node, three queues, same
machinery. Mesh-scoped and node-scoped seats already exist; here they do the work of "one shared
service" versus "one worker per machine".
## 4. Five relationships
| | provision | event | job | state | tool |
|---|---|---|---|---|---|
| shape | 1:1 resource | 1:many | N:1 | 1:1 | 1:1 |
| addressed to | a provider | the emitter's own namespace | a **seat** | one node | a module or seat |
| who must act | the provider | nobody | exactly one holder | that node | the server |
| credential | sealed, per consumer | none | none | none | none |
| reply | — | none | none, or an event later | a report | awaited |
| retention | — | age and size | work queue, explicit ack | **last per subject** | none |
| declared | `provides`/`requires` | `emits`/`consumes` | seat `accepts` / `uses` | the mesh's own | `serves` |
**Job** is the one [ADR 0041](../../02-DECISIONS/0041-events-are-a-relationship.md) had no room
for. Its table has events at 1:many and provisions at 1:1; a module submitting work to a service
is neither. It is not an event, because an event is a broadcast nobody is obliged to act on and a
second holder would do the work twice. It is not a provision, because there is no resource and no
credential. What makes it safe is not cleverness in the subscribe call but the seat: exactly one
holder, so exactly one worker, by construction.
**State** is the shape the deploy path needs and nothing else uses. A declaration is not an event
— replaying yesterday's is actively harmful — and not a job. Only the newest matters, which is
last-per-subject retention, and a node that has seen sequence *n* refuses *n−1* by construction.
That is the wire-level answer to
[issue 107](../../04-ISSUES/107-a-declaration-carries-no-order/00-report.md).
## 5. Seats
A module declares a seat with its protocol, and the mesh enforces one holder at its scope
([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md)). A caller declares that
it uses the *seat*, never the module, so the implementation can be replaced under it.
- The set of seats is **derived** — the mesh's own, plus every registered module's — so it is both
closed and extensible, and enumerating it is a query rather than an inventory.
- `mesh-*` is **reserved**: the prefix is the reservation rule, and a module declaring one is
refused at registration.
- Two modules declaring the same name: the second is refused.
- A module may not claim a seat whose protocol it does not implement.
- **Nobody holding a seat is not an error.** The stream exists from registration, so work queues
until a holder appears. Install the telegram module a week later and the backlog flushes.
## 6. The lifecycle: build, publish, deploy
Every shape above appears once, in order, and no step knows where the next one runs.
**A change lands.** The module holding `mesh-git` emits `pushed` — repository, ref, commit. An
event, because it is a fact about git and git's identity is the meaning.
**The change becomes work.** The controller consumes `pushed`, asks the catalogue which modules
are built from that repository and path
([ADR 0069](../../02-DECISIONS/0069-a-module-is-a-repository-and-a-path.md)), and submits one
**job** per affected module to the `mesh-build-machine` seat. The builder stays simple: it builds
what it is handed, and never resolves anything. A builder that dies mid-build has its job
redelivered, because a work queue with explicit ack is what that means.
**The artifact is published.** The builder pushes to the registry seats and emits `built` —
module, version, digest. An event again: a fact about the builder.
**The build cascade is that event fanning out through a graph the mesh already has.** A module
whose image is built *on* another's artifact declares that in `build.on`. So `built` reaches the
controller, which walks the declared graph and submits rebuild jobs for everything downstream. A
dependency cascade is not special machinery — it is one event, one derived graph, and the same job
queue.
**Deployment is state, not a message.** The controller composes each affected node's declaration
and publishes it last-per-subject. A node that was away gets exactly the current one, never a
queue of superseded ones, and a replayed older one is refused by sequence.
**Applying is reported to a role.** The host applies and reports to the `mesh-controller` seat —
not to an address it was given at genesis. Held and retried while the store restarts
([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)).
What disappears across that chain is every address. No webhook URL, no registered callback, no
"which node is the builder on", no controller endpoint baked into a joining node. That is the
class of bug
[issue 102](../../04-ISSUES/102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md)
names, dissolved rather than fixed.
## 7. Modules depending on each other
Three kinds, and conflating them is how deployment ordering goes wrong.
**Build-time** — A's image is built on B's artifact. Resolved by the cascade above; nothing at
runtime cares.
**Provision** — A requires a database from B. A **hard** dependency: the credential must exist
before A can start, so resolution gates delivery and A is shown as waiting until B has answered
([design 27](27-a-module-requires-the-mesh-resolves.md)).
**Seat** — A uses B's seat. A **soft** dependency, and this is the one the bus changes. A starts
whether or not anyone holds the seat, because the stream absorbs the gap. Deployment order stops
mattering for everything expressed this way, and a service being restarted, moved or upgraded is
not an outage for its callers — it is latency.
That difference is worth choosing on purpose. A dependency expressed as a provision must be
ordered; the same dependency expressed as a seat need not be.
## 8. Versioning a protocol
A seat's protocol is a compatibility surface between modules that do not know each other and are
deployed at different times. Four ways it can change, and they are not equally dangerous:
| change | example | detectable |
|---|---|---|
| **additive** | a new `accepts` subject, a new optional field | nothing breaks |
| **removal or rename** | `send` becomes `deliver` | yes, mechanically |
| **shape** | an optional field becomes required | yes, if shapes are specified |
| **semantic** | `send` starts meaning *queue for tomorrow* | **no** |
**Additive is free.** A seat may grow without a version, without re-registering a caller, and
without ceremony. Most change is this.
**A breaking change is refused while anyone is bound.** Registration computes a compatibility
fingerprint over the seat's protocol — its subjects and the shapes they carry. A registration that
alters the fingerprint while callers are bound is refused, **and the refusal names them**. The
mesh already holds the `uses` graph, so this is derived rather than declared, and it turns a
runtime breakage into a registration-time conversation.
**When a break is genuinely needed, the version goes in the subject, not the name.** The seat stays
one thing; `mesh.seat.<seat>.v2.<verb>` runs beside v1 and the holder serves both. A caller moves
when it is ready. Versioning the *seat name* was considered and rejected: it forks the role, so
"one holder" stops meaning one provider of the capability, and every document naming the seat has
to be found and changed.
**Binding is recorded, not inferred.** A caller declares `uses: telegram-sender` with no version,
and resolution binds it to the current one and records that — the same **pin** machinery
[design 27](27-a-module-requires-the-mesh-resolves.md) already uses when resolution had a choice
to make. Moving to v2 is a deliberate re-pin, so nothing drifts onto a new protocol because it
happened to be newest.
**Retirement is reported, never automatic.** When the `uses` graph shows nothing bound to v1, the
overview says it is retirable. The mesh does not remove it.
**And none of this catches a semantic change.** Same subject, same shape, new meaning: no
fingerprint sees it, and no check proposed here would. The defences are review, and pushing
meaning into shape wherever it can go — a required `channel` field is caught, a changed
interpretation of an existing one is not. Saying so is better than implying the fingerprint is
complete, because a team that believes it is complete stops reviewing for the case it misses.
## 9. Provisioning over the bus
Provisioning rides the bus, and the provider stops having an address.
| part of a provision | shape |
|---|---|
| the requirement resolving to a provider | the controller's, not the bus's |
| the grant reaching the provisioner | request/reply to a **role** |
| `holds` — the reconcile question, every minute | the same call, on a timer |
| `provisioned` / `deprovisioned` | events |
| the credential reaching the consumer | §10 — fetched, never carried |
A provisioner's interface is already three calls — create, remove, holds — which is exactly a
`serves` protocol. So **a provision interface is a seat whose protocol is those three**, which is
why [design 26](26-the-seats.md) already allows a seat to deliver a provision. The two concepts
were converging before this document; here they meet.
What stays different, and must not be unified away: a provision has a **per-consumer resource and
a sealed credential**, created and destroyed per consumer. A seat protocol has neither — it is a
role you send to. Collapsing them would mean pretending a database is a subject.
### `mesh-bus` and `nats` are two interfaces, never one name
The mesh's own bus is **`mesh-bus`**, delivered by the `mesh-broker` seat and answered by the
controller — because the bus's accounts are configuration rather than something a provisioner
creates, so there is no provisioner process in the path and nothing waiting on a bus account in
order to make bus accounts. A module that runs a NATS server of its own and offers it as a
backing service provides **`nats`**, exactly as the deprecated broker provides `amqp`
([ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md): AMQP is not a provision, and everything speaks to the `mesh-broker` seat)).
They are never the same name. A manifest saying `nats` could otherwise mean either the mesh's
nervous system or a private queue, and the difference between those is the whole architecture.
0119's rule decides which is legitimate: a private bus is a backing service, never a channel to
another module.
### Where addresses survive
"Where is it?" is two different problems, and the bus solves one of them completely and the other
not at all. Keeping them apart matters, because a reader who thinks the mesh no longer has
addresses will believe a class of bug is fixed when it is untouched.
**Something that is on the bus: the address disappears.** A host reporting in used to need the
controller's address — recorded somewhere, at some moment, and wrong as soon as anything moved.
Now it publishes to `mesh.seat.mesh-controller.report` and the bus routes it to whoever holds the
seat. Nothing anywhere records where the controller is, so nothing can record it *wrongly*. The
same is true of the builder, the catalogue, the telegram sender. This class is not mitigated; it
is gone, because the information is no longer stored.
**Something that is not on the bus: the address stays, exactly as before.** A module that requires
a database does not reach postgres over NATS — it opens a postgres connection, because postgres
speaks postgres and is not listening on any subject. Its credential contains a host and a port,
and no amount of subject addressing changes that.
So the fix for that second class is unchanged and is not this document's:
[ADR 0098](../../02-DECISIONS/0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md) —
fetch the fact where it is used rather than storing a copy — which is what
[issue 102](../../04-ISSUES/102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md)
is actually about. The bus makes that class *smaller* by removing every mesh-internal address from
it. It does not make it empty.
**And one address is irreducible: the bus's own.** A node has to know where the broker is before
it can use subjects for anything, so that one cannot be a subject. It is the addressing equivalent
of §10's bootstrap — the first thing cannot be found by the mechanism that finds everything else.
## 10. Secrets, and why they never enter a stream
Everything the vault does is request/reply to the `mesh-vault` role: mint, fetch, rotate. In that
sense it is as much on the bus as anything else.
**The bus is not trusted with a secret, and does not need to be.** A secret is sealed to its
recipient, so what crosses the bus is ciphertext only that recipient can open. The broker sees
that a secret moved, and to whom — metadata, which is acceptable — and never a plaintext.
**But sealed is not enough on its own, because a stream persists.** A sealed secret written into
a JetStream stream is a durable ciphertext sitting in the mesh's own storage, and the day a
sealing key leaks, that stream is an archive rather than a moment. So:
- **A secret travels on core request/reply, never through a stream.** No persistence, no replay,
nothing to exfiltrate later.
- **A declaration names a secret; it does not carry one.** Declarations are the state shape, which
*is* a stream — so the host fetches the secret from the vault at apply time, over the core path.
That is [ADR 0098](../../02-DECISIONS/0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md)'s
existing discipline — *fetched from it, not carried* — applied to the one payload where carrying
it is worst.
**The bootstrap, which is circular and has a precedent.** The vault makes every secret
([ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md)), including the bus's own
passwords. The vault is a module, and a module needs a bus account, whose password the vault
makes. Nothing can go first.
This is the shape [ADR 0067](../../02-DECISIONS/0067-genesis-is-a-pivot.md) already resolves for
the control plane: **genesis is a pivot.** The controller mints the handful of foundation
credentials itself, raises the store, the broker and the vault, and then the vault takes over and
mints everything from there — the same move as raising a temporary control plane and reinstalling
it as an ordinary module once the registry exists.
So there are exactly two things the normal path cannot make, both at genesis, both ending the
moment the mesh can mint for itself: **the bus's own accounts** (§the bootstrap argument in
[ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md): AMQP is not a provision, and everything speaks to the `mesh-broker` seat) (superseding [ADR 0125](../../02-DECISIONS/0125-the-bus-is-the-only-broker.md)) — a provisioner is a module and
needs an account before it can run) and **the vault's own credential**. Any third exception is a
design failure, and naming these two is what makes a third one visible.
## 11. Open
**Semantic change has no mechanical defence** (§8). Recorded as open rather than solved, because
it is the residue of a question the rest of §8 answers and the part a fingerprint cannot reach.
**Whether a module may declare a seat it does not itself claim** — the contract as one thing, the
implementation as another, which is how two competing implementations would ever exist.
**Whether `consumes` naming another module couples too tightly.** It is kept here deliberately —
an event's provenance is its meaning — but a consumer of `billing.order.placed` does depend on
billing existing under that name.
## 12. How it is checked
- **A manifest holds no subject.** A catalogue test: no manifest contains a string matching the
subject grammar. The rule is worthless if it is followed by convention.
- **Permissions are exactly the three namespaces.** A composition test per module: the derived
permission set equals what its declaration implies, and a hand-written addition to it fails.
- **A sender cannot read the queue it writes to.** A bed: a module declaring `uses` is refused
subscribe on that seat's inbound subject.
- **One holder, one delivery.** A bed: a seat's job delivered once with the holder running, and
a second claim of the seat refused.
- **A queued job survives no holder.** A bed: submit with the seat unheld, assign the holder,
the job is delivered.
- **The cascade rebuilds exactly the dependents.** A bed: publish an artifact two modules build
on, and exactly those two are rebuilt.
- **A stale declaration is refused.** A bed: replay sequence *n−1* after *n*, and the node refuses
it rather than applying it.
+7
View File
@@ -34,6 +34,13 @@ document is written and this one's status becomes `implemented`.
| [`22-the-work-ahead.md`](22-the-work-ahead.md) | Everything decided and not yet built, in dependency order, each phase ending at a run | [ADR 0074](../../02-DECISIONS/0074-the-wire-is-specified-not-the-types.md), [ADR 0075](../../02-DECISIONS/0075-two-stores-and-which-provides-what.md), [ADR 0014](../../02-DECISIONS/0014-no-npm-workspace.md) |
| [`23-choosing-a-provider.md`](23-choosing-a-provider.md) | Which of several providers of a kind serves a consumer, and when a module carries its own instead | [ADR 0084](../../02-DECISIONS/0084-which-provider-serves-a-consumer.md), [ADR 0027](../../02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md) |
| [`24-the-secrets-vault.md`](24-the-secrets-vault.md) | The module that owns a secret — a `secret` provision, and the boundary of what it owns | [ADR 0085](../../02-DECISIONS/0085-a-secret-is-a-provision.md), [ADR 0031](../../02-DECISIONS/0031-the-control-plane-authenticates-nobody.md), [ADR 0048](../../02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md) |
| [`25-the-bus-on-nats.md`](25-the-bus-on-nats.md) | **Proposed.** The architecture of the mesh's bus on NATS: what rides which subject under which guarantee and whose account, how a node joins, how a person reaches a tool, and how the mesh moves from the bus it has | [ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md), [ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md), [ADR 0043](../../02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md), [ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md) |
| [`26-the-seats.md`](26-the-seats.md) | **Proposed.** What a mesh can have one of, who fills each, and a seat's holder answering for the provision it delivers — including the `git` seat a build's source can live on | [ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md) (superseding [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md)), [ADR 0111](../../02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md), [ADR 0109](../../02-DECISIONS/0109-a-package-registry-seat-is-one-per-ecosystem.md) |
| [`27-a-module-requires-the-mesh-resolves.md`](27-a-module-requires-the-mesh-resolves.md) | **Proposed.** One concept for everything a module needs: a requirement with a contract, answered by one of four kinds of provider, resolved at assignment or refused. Retires settings, placeholders, facts and paths in definitions | [ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md), [ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md), [ADR 0114](../../02-DECISIONS/0114-a-shared-credential-rotates-over-two-credentials.md), [ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md) (superseding [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md)) |
| [`28-building-the-bus.md`](28-building-the-bus.md) | **Proposed.** The five steps of the bus work in the order their dependencies allow, each ending at a bed — with the surface measured, so no step's size is a guess | [ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md), [ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md), [ADR 0074](../../02-DECISIONS/0074-the-wire-is-specified-not-the-types.md) |
| [`29-a-node-has-operator-accounts.md`](29-a-node-has-operator-accounts.md) | **Proposed.** The mesh models machines but not the humans on them: a node gains operator accounts, and a resource may live under a home owned by its account — what would own ~/.ssh, dotfiles and ~/.config when HAL retires | [ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md), [ADR 0051](../../02-DECISIONS/0051-shared-data-is-the-operators.md) |
| [`32-what-a-module-declares.md`](32-what-a-module-declares.md) | **Proposed.** What a module declares and what the bus derives from it: three namespaces, subjects from local names, queues never declared, the five relationships, and the build-publish-deploy lifecycle on one bus | [ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md), [ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md), superseded by [ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md) (superseding [ADR 0125](../../02-DECISIONS/0125-the-bus-is-the-only-broker.md)), [ADR 0041](../../02-DECISIONS/0041-events-are-a-relationship.md) |
## Not yet written
@@ -1,8 +1,8 @@
---
status: open
status: resolved
opened: 2026-09-22
located-in: []
fixed-by:
located-in: [mesh-catalog]
fixed-by: mesh-catalog — the 11 not-defensible modules (de-spiegel, gitea, hello-web, mailu, mssql, n8n, novox.be, only-office, photos, photos-eef, photos-filip) had their hardcoded machine-side port mapping stripped to the bare software port, letting ADR 0038's existing assignment machinery (internal/inventory/ports.go's PortFor, declaration.go's publishedOn) assign it as designed. postgres, lavinmq, distribution (foundation, genesis-rewritten) and unifi (protocol-fixed) were left as-is — the two defensible kinds this report names.
amended-design:
---
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-23
located-in: [mesh-host internal/apply]
fixed-by:
fixed-by: mesh-host PR #22 — a container records the digest of every file it reads at creation, its env-files and files mounted into it directly, and is recreated when one changes; a pre-upgrade label is accepted once, and the plan names the file. A mounted directory still needs restart-on.
amended-design:
---
@@ -8,6 +8,35 @@ amended-design:
# 108 — The registry has no garbage collection, and two doors make it harder to add
## 2026-09-26 — the second door was attached to the wrong thing
*Left in the title and in the text below rather than rewritten, because the reasoning that assumed
two doors is what a reader needs to see retracted.*
Two separate things were treated as one. **The mesh's own artifact store holds the store seat**: it
is internal, reached by name over the overlay, with no accounts, because being on that network is the
permission ([ADR 0082](../../02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)).
**Serving a registry publicly is a service the mesh can host** — a module with its own name, its own
accounts and its own storage, like any other thing it runs for somebody. The conversion this report
was written beside gave the seat-holding store a second, public, authenticated door over the same
filesystem, which is neither of those: it is the internal store with an external face.
That second door is not being built. A publicly served registry, if one is wanted, is a module beside
the store rather than another way in to it, and it brings its own storage with it.
What that leaves here:
- **The complication in the title is gone.** One door means one registry process on the store's
filesystem, so the shared blob-descriptor cache and the deletion-cached-by-the-other-door problem
this report worried about do not arise at all.
- **The original issue is untouched, and is the whole of it.** The mesh's registry has no garbage
collection, never had, and the settings a collection routine depends on are not enabled.
- **One thing is sharpened rather than removed.** Enabling deletion on the only door enables it on a
door with no accounts, reachable by everything on the overlay. The predecessor kept deletion behind
its authenticated door — which it could, having one. So adding garbage collection now includes
deciding whether deletion is exposed on that door at all, or only ever performed by a routine the
mesh runs against its own store.
## What was observed
Reviewing the conversion that gives the mesh's image registry its public, authenticated name,
@@ -1,8 +1,8 @@
---
status: open
status: resolved
opened: 2026-09-24
located-in: []
fixed-by:
located-in: [mesh-controller internal/inventory, mesh-controller internal/catalogue, mesh-controller cmd/mesh-controller, mesh-catalog modules/dnsmasq]
fixed-by: [mesh-controller#73 overlay-name + namesInTheMesh, mesh-catalog#106 dnsmasq daemon.json merge]
amended-design:
---
@@ -0,0 +1,78 @@
# Diagnosis
## 2026-09-24 — where the mesh's own record ends, and what stands beside it
Checked the mechanism first. The resolver module's config (`mesh-catalog modules/dnsmasq`) is
generated whole from `nodes.conf`, the `dnsmasq.fact-node-zones` fact `mesh-controller` computes:
one `address=/<node>.<suffix>/<address>` wildcard per node the mesh has a name for, plus
`local=/<suffix>/` — which tells dnsmasq it is *authoritative* for the whole suffix, so anything
under it that is not one of those wildcards is refused, not forwarded. That is the exact mechanism
the report describes: **the mesh does not lack an upstream to ask, it has told itself there is
nothing to ask.**
Checked whether "forward to the resolver it replaced" (the third open question, ADR 0104's shape)
is literal here the way it was for the proxy. It is not, and the difference matters: the proxy's
predecessor kept running throughout its migration, a separate process on ports the mesh's proxy did
not yet hold. DNS has one process on one port. `mesh-catalog`'s dnsmasq module, on assignment,
replaces `/etc/dnsmasq.conf` whole and restarts the one `dnsmasq.service` unit — the same unit the
predecessor's own resolver is running as, right now, on this control-node:
```
$ systemctl status dnsmasq
● dnsmasq.service ... Active: active (running) since Fri 2026-09-18 ...
```
There is no predecessor process left standing after that assignment for a forward rule to reach.
An adapter in ADR 0104's literal shape — a second thing running, written into by a first — does not
fit; whatever answers this has to be data carried across the cutover, not a live thing forwarded to
across it.
## What is actually missing, and where it already exists
The report's first open question — *should a carried peer be nameable, the operator saying which
machine a carried address is, before it enrols* — turns out not to need a guess. The predecessor's
own resolver config, still live on this control-node, already states it:
```
# /etc/dnsmasq.d/hal-dns.conf — Generated by HAL dnsmasq-app
address=/novox.internal/10.10.0.1
address=/ace.internal/10.10.0.2
address=/g14.internal/10.10.0.4
address=/shanks.internal/10.10.0.3
```
Checked against what the mesh itself recorded when it took the tunnel over (`overlay show`):
```
a peer of the tunnel (zTKYP6Sz…) 10.10.0.2 not yet enrolled
a peer of the tunnel (XT1l25Z0…) 10.10.0.3 not yet enrolled
a peer of the tunnel (mq76YPW0…) 10.10.0.4 not yet enrolled
```
`10.10.0.2`/`.3`/`.4` match `ace`/`shanks`/`g14` exactly, one for one. This is not the operator
inventing a fact under time pressure — it is the predecessor's own record, current, and it has been
correctly serving these three names for six days without correction. Recording it is transcription,
not assertion.
**located-in, tentatively:** `mesh-controller internal/inventory` (where a carried peer's record
lives today — `CarriedPeer`/`TunnelPeer` in `tunnel.go` carry a public key and an address but no
name field) and `cmd/mesh-controller` (a command to set it — nothing today lets an operator attach
a name to a carried-tunnel-peer record; `node add` is for enrolling nodes, not naming peers, and
`node public-domain <name> <d>` is the closest existing shape to model a new verb on). Separately,
`mesh-controller internal/catalogue` (`facts.go`'s `nodeZones`) would then need to emit an
`address=` wildcard for a *named carried peer* the same way it does for a node — it renders only
from the mesh's `addresses` map today, keyed by node name, and does not distinguish a named carried
peer from an unnamed one. `mesh-catalog modules/dnsmasq` itself needs no change: it already
restarts on the `node-zones` fact and would pick up the new wildcard the moment `internal/catalogue`
emits it.
## What is still open
- The mechanism for *recording* the name is not designed — a new field on the carried-peer record,
a new CLI verb, and what happens if the name later disagrees with what the peer states on
enrolling (it should win; nothing says so yet).
- Whether this generalises: the predecessor's static hosts file was the answer *here* because it
happened to be readable and correct. A migration without one still has the report's harder second
option (refuse the resolver until every peer enrols) as its fallback.
- Not implemented. This diagnosis is what the fix would touch and why the ADR 0104 framing in the
report's third option does not transfer directly — not the fix itself.
@@ -0,0 +1,151 @@
---
status: resolved
opened: 2026-09-24
located-in: [mesh-catalog modules/minio]
fixed-by: mesh-catalog — the object-store module repinned to a maintained fork of the withdrawn server image, its runtime sidecar built from source rather than pulled, and its data moved off the predecessor's live directory. The standing condition this report names is not closed by it — see What was done.
amended-design:
---
# 113 — The object store's images were withdrawn upstream, and only a node that already holds them can still run it
## What was observed
On 2026-09-24, during a service-by-service cutover, the object-store module could not be built on
a node that did not already hold its images. Both images the module needs answer an anonymous
pull with `401 UNAUTHORIZED`:
```
<registry>/minio/minio 401 <registry>/minio/operator 200
<registry>/minio/mc 401 <registry>/minio/console 200
```
The module pins the server image **by digest**, in its `module.json`:
"image": "<registry>/minio/minio@sha256:14cea493…"
*Corrected 2026-09-24. This first said the module pinned a **tag** whose default was four and a
half years old. That described the **predecessor** mesh's object-store module — a different file
in a different repository — not the module being cut over to. The conflation, and what it cost,
is retracted in full in [the diagnosis](01-diagnosis.md).*
The cause is upstream and outside the mesh: the vendor **deleted** the community server and client
repositories. It is not an access policy that a credential could answer, and nothing about the
mesh's own registry configuration, resolver or trust settings is involved.
- The vendor removed both repositories from the main public registry on **2026-09-11**. Its API
answers `404` for the server repository while a sibling in the same namespace answers `200`.
- The secondary registry that the wider ecosystem repointed to as a stopgap **no longer lists
them either**. Sixty-eight repositories in that namespace are still public and pull normally;
the server and the client are simply absent, and the namespace is now dominated by the vendor's
commercially licensed line.
- The open-source repository was archived in **February 2026**, and the community edition has been
source-only since **October 2025**. No new images are published anywhere.
## What did not happen, and why it is recorded
The first reading of this was that the registry had *disabled anonymous pulls for the whole
vendor namespace*. That was wrong in a way worth keeping, because the evidence looked conclusive:
- The anonymous token carries `"actions": []` for the affected repositories and `['pull']` for
working ones — a real signal, but it is **also exactly what a repository that does not exist
returns**. A deliberately invented repository name in the same namespace produced a
byte-identical response. The signal cannot distinguish *revoked* from *absent*.
- The token also carries `"$disabled"`, which was read as confirmation. It appears on **every**
repository on that registry, including the ones pulling successfully. It describes image
**signing**, not access.
Sibling repositories in the same namespace pulling normally is what rules out a namespace-wide
policy, and the registries' own APIs — `404` against `200` — are what establish deletion.
## The predecessor mesh was not blocked, which is the other half
*Scope, corrected 2026-09-24: everything in this section describes the **predecessor's** delivery
machinery and its object-store module. It is what made the instance harmless, and it is why
dropping the module from the queue was unnecessary. It says nothing about how the mesh being built
resolves images, which is a different mechanism.*
A node that already holds the images runs the module normally. The node carrying the cutover holds
the pinned server image, the client, and the load-balancer image the module composes with, all
pulled years ago. Its resolved version variable matches the cached tag exactly.
This is by design and not by luck. The deploy stage pulls **best-effort** and then asserts only
that every image the composition declares **resolves locally**, precisely so that an image which
exists on the node but can no longer be fetched does not fail a deploy. The code comment naming
the precedent describes this case exactly — *"an old tag pulled years ago and since removed
upstream"* — and records that failing on the pull instead had previously made a module
undeployable while all of its images sat on the node.
So the deploy logs a warning and succeeds. Dropping the module from the cutover queue was not
necessary.
## Why it matters beyond this instance
The instance is harmless; the standing condition is not.
1. **No new node can ever provision this module.** Every node that does not already hold the
images is permanently unable to obtain them, and the same will be true of any module whose
upstream withdraws an image.
This is the failure mode of a deliberate design choice, which is why it is worth recording
rather than patching. The foundation design chooses **references over payload** — *"the bundle
names images by digest and the host fetches them"* — on the stated grounds that
*"reproducibility comes from pinning the identity of a thing rather than carrying its bytes"*
([to-be 07](../../03-DESIGN/01-to-be/07-the-foundation.md)). That reasoning is sound. It holds
only while a pinned identity stays **resolvable**, and nothing in the mesh's control guarantees
that for an image in somebody else's registry. The passage is about
the foundation bundle, and this module is not in it; but the pattern — pin the identity, fetch
the bytes on demand — is how every module gets its third-party images, so the exposure is
general even though the sentence is local.
2. **The pinned release is permanently unpatched.** It is four and a half years old, upstream is
archived, and no security fix will ever reach it.
3. **Nothing detects this class of failure.** The condition is invisible until a node without the
image tries to deploy. The very guard that correctly stops this from breaking existing nodes —
assert local resolvability, not the pull — also means a warning is the only trace, and a
warning is not a rule. A mesh cannot state that its modules are installable while the only
evidence is that they are already installed.
The third point is the general one, and it is not specific to this vendor: an image pinned against
a registry the mesh does not control — **by tag or by digest, it makes no difference** — is a
dependency with no guarantee behind it, and the mesh currently learns it has lost one only by
trying to use it.
[Issue 064](../064-a-mesh-build-cannot-fetch-a-modules-external-dependencies/00-report.md) is the
nearest precedent, and it does not cover this. That issue asked whether the mesh's build
environment can **reach** a declared vendor image — a network-policy question, answered by
requiring the image be declared as a build input — and it assumed that an image, once declared,
stays fetchable. Withdrawal is the case the assumption does not cover: no network policy and no
declaration makes a deleted repository resolvable, so a module can satisfy 064 in full and still
be unbuildable on a node that holds nothing.
## What was done
The module was repinned to a maintained fork of the server image, published to a registry that
still serves it; its runtime sidecar is now built from source rather than pulled; and its data was
moved off the predecessor's live directory. The object store runs on the control-node from that
pin, and a node holding nothing can obtain it again.
That answers the instance and none of the three points above. The mesh still cannot say which of
its other pinned third-party images are still obtainable, and it would still learn of a withdrawal
only when a node without the image tried to deploy. The replacement question — S3 the protocol
rather than this product — is carried by
[research 015](../../01-RESEARCH/015-the-object-store-after-minio/00-overview.md); the detection
question is carried by nothing, and is the first of the open questions below.
## Open questions
- Should the mesh **hold** the images it depends on — mirroring third-party images into its own
registry at adoption, so a module's installability does not depend on an upstream's continued
goodwill? That is the fix that generalises. It costs storage and a policy about what to mirror,
and it is a deliberate move **away** from references-over-payload for third-party images
specifically — so it should be decided as such, not smuggled in as a fix.
- ~~Should a module's images be pinned **by digest** rather than by tag?~~ **Answered, and the
premise was wrong.** This module already pins by digest, and it made no difference: the
repository was deleted, so the digest resolves to nothing. A digest buys an exact, auditable
artifact; it buys no protection whatever against withdrawal. Struck rather than deleted, because
the question was asked from a mistaken reading of the manifest and that is worth seeing.
- What **checks** that every module in the catalogue is still obtainable from a node that holds
nothing? Nothing does today. A periodic cold-pull of the catalogue would have caught this on
2026-09-11 rather than thirteen days later, mid-cutover.
- For this module specifically: replace the product. The design already says the dependency is on
**S3 the protocol, not the product** — see
[research 015](../../01-RESEARCH/015-the-object-store-after-minio/00-overview.md).
@@ -0,0 +1,136 @@
# 113 — Diagnosis
## 2026-09-24 — the trail
The symptom arrived already carrying a diagnosis: *the registry has disabled anonymous pulls for
the entire vendor namespace.* Everything below was an attempt to confirm that, and it did not
survive.
### Step 1 — the token is not the test
The reported evidence was the anonymous pull token's contents: `"actions": []` and `"$disabled"`.
A token is an intermediate artifact. The test is whether a manifest can actually be fetched with
it, so the first step was to request one:
```
GET /v2/minio/minio/manifests/<pinned tag> -> 401 UNAUTHORIZED
```
That confirmed the failure but said nothing about its scope or cause.
### Step 2 — a control ruled out the stated cause
The same request flow, in the same minute, against other repositories:
| repository | token actions | manifest |
|---|---|---|
| the vendor's server | `[]` | **401** |
| the vendor's client | `[]` | **401** |
| the vendor's operator | `['pull']` | 200 |
| the vendor's console | `['pull']` | 200 |
| the vendor's sidecar proxy | `['pull']` | 200 |
| an unrelated public project | `['pull']` | 200 |
**A namespace-wide policy is ruled out.** Two repositories fail; their siblings in the same
namespace pull normally.
### Step 3 — both pieces of the original evidence were red herrings
- `"$disabled"` is present on **every** repository on that registry, including all of the
successful ones above. It belongs to the signing context, not to authorisation. It carries no
information about this failure at all.
- `"actions": []` with a `401` is **indistinguishable from a repository that does not exist.** A
deliberately invented repository name in the vendor's namespace returned a byte-identical
response — empty actions, `401`. The signal cannot separate *access revoked* from *not there*,
so it cannot support the conclusion it was used for.
That second point turned the question from *who revoked access* to *is it still there*.
### Step 4 — the registries' own APIs establish deletion
Asked directly, rather than through the pull path:
- **Primary registry:** its repository API answers **404** for the server repository, and **200**
for a sibling in the same namespace. The repository is gone, not private.
- **Secondary registry:** a listing of the vendor's namespace returns **68 public repositories**.
The server and the client are **absent from the list**. Present are the operator, console,
sidecar proxy, benchmarking and key-management images — and a large, newer set belonging to the
vendor's commercially licensed line.
### Step 5 — upstream confirms, and dates it
The vendor deleted the community server and client from the primary registry on **2026-09-11**.
This was the last step of a staged withdrawal: free image publishing stopped in **October 2025**,
the community console UI was removed mid-2025, and the open-source repository was archived in
**February 2026**. The secondary registry was where the ecosystem repointed as a stopgap; it has
since lost the two repositories as well.
**Conclusion: the images were withdrawn, not restricted.** No credential can answer this, because
there is nothing left to authenticate against. A different image source is the only remedy.
## Why the mesh kept working, checked rather than assumed
The claim that the module "genuinely isn't ready for cutover" was tested and is false for the node
in question.
- The node **holds** the pinned server image, the client, and the load-balancer image the module
composes with — all pulled years before the withdrawal.
- The module's resolved version variable on that node matches the cached tag **exactly**, so the
composition references an image that is present.
- The composition declares **no pull policy**, so a present tag is used as-is.
- The deploy stage pulls best-effort, then asserts only that every declared image resolves
locally. A pull error becomes a logged warning when all images are present.
- The start path restarts the unit and does not pull. The one `--pull always` in the tree sits
inside a generated guide describing the **superseded** approach, not in the code that writes
units.
So a deploy of this module on that node succeeds today.
## RETRACTED — the four claims this diagnosis called unsubstantiated
*Added 2026-09-24, the same day, after the error was pointed out.*
This diagnosis originally carried a table headed *"Claims in the original report that could not be
substantiated"*, asserting that a `module.json` did not exist, that no digest pin existed, that an
all-zeros runtime digest appeared nowhere, and that a readiness document was not on disk. **The
table was wrong and it is withdrawn in full.** The original report was accurate.
| Claim, as reported | Actual finding |
|---|---|
| The manifest is a `module.json` pinning the server image by digest at line 83 | **True, and exactly.** `modules/minio/module.json`, digest pin, line 83. |
| The module's own runtime artifact carries an all-zeros placeholder digest | **True.** A second container resource pins a runtime sidecar at an all-zeros digest, meaning nothing was ever published for it. |
| A dated readiness document records several modules with placeholder digests | **Unverified, not disproven.** It is not on the machine searched. The migration record is a separate private repository that is not checked out there, so its absence locally is not evidence. |
| The cached server image is an older release than the module pins | **Unverified.** What was checked was the *predecessor's* tag pin against the cache, which did match. Whether the cached image is the digest this module pins was never checked. |
### Why it went wrong, stated plainly
**One repository was searched, and absence in it was reported as absence.** The catalogue of the
mesh being built is a **separate repository**, not checked out on the machine where the search ran.
Every one of the four claims was about that repository. The searches were real and their output was
reported honestly; the inference drawn from them was not warranted.
Compounding it, the predecessor's object-store module and the one being cut over to were treated as
the same thing. They are different files, in different repositories, with different shapes: the
predecessor's is a compose file pinning a **tag** with a version variable, and it declares no
sidecar; the one being cut over to is a JSON manifest pinning a **digest**, and it declares two
container resources. Findings about the first were written up as findings about the second.
**The lesson worth keeping, because it is not specific to this issue:** *"zero occurrences anywhere
in the tree"* is only ever as strong as the tree that was searched, and a diagnosis must name which
tree that was. This one did not, which is what let a one-repository search read as a mesh-wide fact.
A confident rebuttal of a correct report is worse than no diagnosis, because it sends the next
person looking in the wrong place with the authority of a written record behind them.
What none of this changes: the images are gone upstream, and that finding stands on the registries'
own APIs.
## What is located, and what is not
**Located:** the object-store module in the catalogue of the mesh being built — it pins, by digest,
a server image that no longer exists anywhere public, and a runtime sidecar that was never
published.
**Not located, and deliberately left open:** the general condition. The mesh has no mirror of the
third-party images its modules depend on and no check that a module is obtainable by a node
holding nothing. That is a design gap rather than a defect in this module, and it is stated as an
open question on the report rather than answered here.
@@ -6,7 +6,7 @@ fixed-by:
amended-design:
---
# 113 — Should the controller run as a container, or as a process the host supervises directly?
# 114 — Should the controller run as a container, or as a process the host supervises directly?
## What was observed
@@ -71,3 +71,16 @@ checked.
a controller outage tolerable regardless of which resource type it is?
- If the answer is "keep it a container," what does that answer, precisely, that this issue asked —
so the next person who notices the same asymmetry finds it answered rather than open again?
## The general case
[Issue 117](../117-a-modules-own-code-is-a-container-and-a-process/00-report.md) is the same
question asked of every module rather than of the controller: a module's own code is a `container`
in [ADR 0047](../../02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)
and a `process` in the to-be design, and no record moves it. Its
[diagnosis](../117-a-modules-own-code-is-a-container-and-a-process/01-diagnosis.md) answers the
first open question above: the host's `process` shape is built, applied and tested, including the
restart and run-to-completion semantics — so this would not need host-side work first.
The two do not collapse into one. The controller is not a code-carrying sidecar, and `network: host`
is what makes the asymmetry visible here and nowhere else.
@@ -0,0 +1,75 @@
---
status: resolved
opened: 2026-09-24
located-in: [mesh-catalog]
fixed-by: mesh-catalog PR #54 — distribution, lavinmq, postgres and searxng converted to host directory binds; data copied and verified (mesh-store stopped cleanly first for a crash-consistent copy), old named volumes kept as the rollback path
amended-design:
---
# 115 — A named Docker volume is invisible to the operator, and one flag from gone
## What was observed
On the control-node, 2026-09-24, mid-migration, checking every module in the catalogue for how it
mounts its data. Five container mounts across four modules use a **named Docker volume** rather
than a host directory:
```
distribution mesh-registry → mesh-registry-data:/var/lib/registry
lavinmq mesh-broker → mesh-broker-data:/var/lib/lavinmq (+ mesh-broker-tls)
postgres mesh-store → mesh-store-data:/var/lib/postgresql/data
searxng valkey → searxng-valkey-data:/data
```
Every other module in the catalogue — more than forty of them — mounts a host directory,
`/var/lib/<module>/...`, matching what `DATA-CUTOVER.md` and every rehearsed recipe tonight
assumes. These four are the exception, not a second convention.
`mesh-store` is the one that matters most: it holds every database migrated tonight, including a
live `keycloak` restore verified minutes before this was written.
## Why it matters beyond this instance
[`03-DESIGN/01-to-be/22-the-work-ahead.md`](../../03-DESIGN/01-to-be/22-the-work-ahead.md) shows
this was a deliberate choice, not an oversight — *"Data survives on the named volumes"* — but that
sentence answers a narrower question than the one this issue raises. It says a named volume
survives **ordinary container recreation** (a rebuild, a `take`, a routine `docker rm -f` and
push), which is true and which every module already gets from either a named volume or a host
directory equally.
What it does not address: a named volume is **one flag away from deleted**, in a way a host
directory structurally cannot be.
- `docker rm -f` alone does not remove a named volume — it persists, unreferenced, until something
targets it by name.
- `docker rm -fv`, `docker volume rm`, and `docker system prune --volumes` all do target it, and
the difference from the command this migration already uses routinely (`docker rm -f` — see
`HANDOFF.md`'s own "a restart does not re-read anything" rule) is one character.
- A named volume is invisible to an operator working the way this migration has worked all
night: `ls`, `find`, `grep` across `/services/*` and `/var/lib/*`. Finding it requires knowing to
ask Docker (`docker volume inspect`), and its actual bytes sit under
`/var/lib/docker/volumes/<name>/_data`, a path nothing points at.
- Nothing external can back it up, snapshot it, or notice it growing without going through
Docker's own volume machinery — a host directory is a directory; a filesystem-level backup job
already reaches it for free.
`mesh-store` carries the sharpest version of this: every module's database, the mesh's own
inventory, identity and licence stores — the single foundation piece the rest of the mesh depends
on — sits somewhere the operator's ordinary tools do not look.
## Decided
**Persistent data is a host directory bind, never a named volume. A named volume may hold only
data that is disposable if lost.** All four hold real state and all four are in scope — including
`mesh-registry` and `searxng`'s cache, not only `mesh-store`. Decided 2026-09-24; the record is
[ADR 0107](../../02-DECISIONS/0107-persistent-data-is-a-directory-bind-never-a-named-volume.md).
## Open questions
- The safe migration path for each, in order of stakes — `mesh-store` live and holding every
database migrated tonight, `mesh-broker`, `mesh-registry`, then `searxng`'s cache, which is
genuinely disposable and may not need migrating at all if it is rebuilt rather than moved.
- Should the catalogue refuse a module declaring a named volume for anything but disposable data
at `module add`, the way [issue 091](../091-a-module-definition-carries-a-machine-port/00-report.md)
asks the same of a hardcoded machine port — so a convention violation is caught at registration
rather than found by reading the whole catalogue?
@@ -0,0 +1,152 @@
---
status: resolved
opened: 2026-09-25
located-in: [mesh-controller examples/route-proxy]
fixed-by: mesh-controller PR #58 — the four capabilities implemented in the reference proxy; the table re-keyed by host and path with a total ordering; a declaration carrying a credential refused rather than served, and an unreadable secret failing closed
amended-design: 03-DESIGN/01-to-be/08-connectivity.md
---
# 116 — The mesh's proxy applies no policy to a request: no authentication, no source restriction, no path scoping, no redirects
## What was observed
On 2026-09-25, deciding whether the mesh's own reverse proxy
([to-be 08](../../03-DESIGN/01-to-be/08-connectivity.md)) can replace the ingress the mesh adopted
from the predecessor, as the permanent public entry point. The standing requirement is that the mesh
does **at minimum** what the system it replaces already does, so the comparison is against the
predecessor's real, live configuration and its module catalogue — not against a feature list.
**The proxy's entire request path is a host lookup and a forward.** Its handler takes the request's
host, finds a target in a table, and either answers a named 404 or hands the request to the reverse
proxy. There is no authentication check, no source-address check, no redirect handling and no
middleware chain anywhere in the program.
The table is the reason this is structural rather than a missing feature: it maps **host → one
target**, and the lookup strips the port and lowercases the host. **A host cannot be routed two ways.**
The design is deliberate as far as it goes — *"a route hands back a name, not a credential"* — but
that governs the **route grant**. It says nothing about what a request arriving at that name is
allowed to do, which is what the live configuration relies on.
## What the predecessor actually relies on, counted
Two sources, because neither alone is complete: the **module catalogue**, and the **ingress's own
dynamic configuration** on the node. The catalogue misses what was hand-written on the node; the
node's directory misses what modules declare as container labels. An earlier version of this report
read only the latter and undercounted as a result.
### 1. Basic authentication — three dependents in the catalogue
| Module | What it is the only gate on |
|---|---|
| the key-value store | its browser UI, which has no login of its own |
| the relational store | a **database web UI** |
| the ingress itself | its own dashboard — twice, counting a desktop flavour |
Every one is a credential-less admin surface whose sole protection is a middleware the mesh's proxy
does not have. The database web UI is the worst of the three, and the ingress dashboard means the
ingress is currently protecting itself with a mechanism its replacement lacks.
### 2. Outright refusal on a path — an incident mitigation
One hand-written rule on the node blocks external access to an internal API path on the forge. Its
own header records it as **incident response to a compromise**, closing the write primitive that was
abused. It is expressed as an allow-list containing a single documentation-range address — that is,
a deny-everyone — and the middleware is named accordingly.
It is **not** an address-scoped allow-list in any useful sense, and reading it as one points at the
wrong fix. What it needs is the ability to refuse a request outright, scoped to a path.
The same header already records where this belongs: *"not mesh-managed. Durable home is the
route-proxy module; re-home when convenient."*
### 3. Path-scoped routing with priority — and this one gates the others
That rule matches a **path prefix** on a host that is **already routed elsewhere**, and carries an
explicit high priority so it shadows the ordinary route. The mail module needs the same shape for a
different reason: it routes a certificate-challenge path on a host that otherwise goes to the mail
front end.
Because the table maps a host to exactly one target, **neither is expressible today, and adding
authentication and a source filter would not make them so.** Path scoping with priority is a
prerequisite for the refusal rule, not a feature beside it.
### 4. Redirect rules — live, and they fail quietly
Two routes canonicalise a `www` name onto its apex with a rewriting redirect. They appear in the
node's configuration and **not** in the catalogue, so a catalogue-only survey misses them. They are
the easiest of the four to lose, because losing them produces no error — just two public names that
quietly stop redirecting.
## What compared cleanly, and is not part of this issue
- **Large and streaming request bodies.** Four modules raise or remove the body cap — the object
store, the file-sync application, the image registry. The standard library's reverse proxy streams
with no default cap, so this needs nothing added.
- **Connection upgrades**, used by at least one console route: native to the standard library's
reverse proxy. Not yet verified live against this proxy, but not absent by design the way the four
gaps above are.
- **Certificate issuance.** The proxy refuses to certify any name the mesh did not route, and
defaults to a staging issuer until a node opts in to production
([issue 004](../004-certificate-issuance-targets-production/00-report.md)) — stricter than the
hand-maintained configuration it would replace.
- **Several public names for one module.** Already solved by the `contributes` many-shape; needs
nothing from the proxy.
## Why it matters
Under the "at minimum" rule, **nothing can be called a replacement for the predecessor's ingress
while any of the four is missing** — and one of them is a live mitigation for an exploited
vulnerability. The two outcomes if it is left unfixed are both bad:
- the affected routes stay on the adopted ingress indefinitely, leaving the mesh running two
reverse proxies side by side with no principled division between them; or
- they are migrated anyway, and three credential-less admin surfaces become reachable by anyone who
can resolve a name, while a known-exploited path loses the block that was put in front of it
during an incident.
## Open questions — answered
These were design questions, not implementation details, and the fix was not written before they
were answered. All three were settled in
[ADR 0108](../../02-DECISIONS/0108-a-route-carries-the-policy-applied-to-a-request.md): policy goes
**on the route**, the set is **closed at the four**, and a declaration **names a secret and never
carries one**. Kept as asked, because what was rejected and why is the half worth having.
- **Where does request-level policy come from?** Today a contribution carries a name and a port.
Extending it to carry policy keeps the mesh as the source of truth, consistent with everything
else a route already does. A separate proxy-side settings layer keyed by route name decouples
policy from the grant but adds a second place to look. The module's own README notes *"the contract
is the file, not this program"*, so this is a contract decision and not a property of one
reference implementation.
- **May a declaration carry a credential?** A password hash in a contribution puts a secret in a
declaration. That cuts across how the mesh mints and holds secrets, and it should be settled
deliberately rather than as a side effect of whichever option is less code.
- **Is the right general shape "these four", or something narrower?** Authentication, refusal, path
scoping and redirects are what the predecessor uses *today*. Whether route-level policy should be
an open middleware surface, or exactly these four and no more, is worth deciding before any of it
is written — an open surface is far harder to withdraw than to add.
## What the fix covers, and what it does not yet allow
*2026-09-25, on resolution.* The gap this issue reports is closed: the proxy applies policy, the
four capabilities exist, the table is keyed by host and path with a total ordering, and the two
failure modes that would rot quietly are held by tests — a declaration carrying a credential is
refused rather than served, and an unreadable secret makes the route refuse rather than open.
**It does not yet let an operator move the affected routes.** That needs the mesh side: a manifest
able to declare these values, and the controller minting the secret that `auth` names. Until both
exist the capability is reachable only by writing the routes file by hand, so the routes held back
on the adopted ingress stay there.
Resolved rather than left open because the issue reports a gap **in the proxy**, and that gap is
gone. The remaining work is not this fault persisting; it is the ordinary build-out of a contract
this record's decision created, and it belongs to
[to-be 08](../../03-DESIGN/01-to-be/08-connectivity.md) rather than here.
**One thing found while fixing it, worth keeping.** Priority was first read with the reader for
ports, which caps at 65535 — and the one real rule this has to reproduce is declared at 100000. It
parsed to zero, so refusal and path scoping would both have shipped looking complete, passing their
own tests, and doing nothing on the only case that motivated them. *A validator borrowed from a
neighbouring field is a silent default*, and the test that now guards it goes through the proxy,
because at the parser the value looked fine.
@@ -0,0 +1,101 @@
---
status: located
opened: 2026-09-25
located-in: [hq, mesh-catalog modules/showcase, mesh-sdk src/tools/index.ts, mesh-tools]
fixed-by:
amended-design:
---
# 117 — A module's own code is a container in one record and a process in another
## What was observed
Asked what the "sidecar" is — the second container a code-carrying module runs beside its
service — and whether a supervised process would do instead. Reading the records to answer it,
the repository answers both ways, and nothing reconciles them.
| record | status | what runs a module's own code |
|---|---|---|
| [ADR 0047](../../02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) | **accepted**, 2026-09-04 | "a **container**, the tool runtime carrying that module's compiled code" — one module, one process, one account; events and tools in that same process, "not a second one to scope and seal" |
| [`01-to-be/18-building-a-module.md`](../../03-DESIGN/01-to-be/18-building-a-module.md) | proposed, 2026-09-21 | a resource type table in which `container` is "an image" and **`process`** is "**its own code**, in three modes", whose default mode is "a unit restarted when it exits", supervised by the machine |
| [`01-to-be/20-writing-a-module.md`](../../03-DESIGN/01-to-be/20-writing-a-module.md) | proposed, 2026-09-21 | one module declaring **four** `process` resources — events, tools, provisioner, a scheduled ingest — each with its own `run` argv, and the sentence "it is why these are `process` rather than four containers" |
Three disagreements, not one:
1. **Container or unit.** ADR 0047 chose a container and said why: a node-wide runtime loading
every module's code could not hold a per-module account, so the runtime is per-module. The
design docs choose a supervised unit running an argv and give no reason, because they do not
record that they are choosing.
2. **One process or several.** ADR 0047's "one module, one process, one account" is the whole
content of its second and third sections. The worked guide declares four for one module and
presents four as the point.
3. **Whether the record was consulted at all.** Neither design doc names ADR 0047 in
`decisions:`. No record supersedes or extends it on this. **The string `process` as a resource
type appears in no decision record** — the shape exists only in two `proposed` design docs.
Meanwhile the thing as built is the container. [ADR 0029](../../02-DECISIONS/0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md)
records that "anything that is a service plus a sidecar currently has to publish a port to talk
to itself," which is one of the things the host's `network` shape was added for.
[Issue 113's diagnosis](../113-the-object-stores-images-were-withdrawn-upstream/01-diagnosis.md)
found a catalogue module declaring "two container resources," the second a runtime sidecar
"pinned at an all-zeros digest, meaning nothing was ever published for it."
[Issue 095](../095-a-module-assigned-after-genesis-has-no-broker-account/00-report.md) is a
sidecar crash-looping on a credential while its service served correctly.
[ADR 0093](../../02-DECISIONS/0093-a-fixture-that-runs-a-modules-runtime-carries-its-name.md)
records that a bed wanting "a sidecar without its server raises the server."
### And the word is in no glossary
"Sidecar" appears sixteen times across five records — two decisions and three issues. It is
absent from [`00-META/glossary.md`](../../00-META/glossary.md), and absent from every document
under [`03-DESIGN/`](../../03-DESIGN/), in both layers. ADR 0047, which creates the thing, never
uses the word; it says "runtime process" and "runtime container". The glossary's own rule is that
"a new name for an existing thing lands here first, in the same change that introduces it in
code," and the page exists because "the terms kept drifting in conversation." A reader asking
what the sidecar is has nowhere in the design layer to look, which is how this was found.
## Why it matters beyond this instance
- **A module author reading the current guide writes a `process`; the catalogue as built declares
a `container`.** [`20-writing-a-module.md`](../../03-DESIGN/01-to-be/20-writing-a-module.md) is
a worked guide with a manifest in it. Whichever of the two is wrong, somebody follows it.
- **The cost of the container shape is paid in four places and totalled in none.** A published
image per code-carrying module, a network so a module can reach itself, a bed that cannot run a
runtime without raising the server it manages, and a credential failure that presents as the
module's own bug. Each record argues its own piece is worth paying. No record puts them beside
the alternative.
- **Both shapes carry a cost the other does not, and neither is written down.** A container
carries its own interpreter; a `process` declaring `run: ["node", "index.js"]` needs an
interpreter present on the machine, which is the machine dependency the statically linked host
([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)) exists to avoid. And `run` is an argv,
where [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) refuses `action` because the link may
not carry a command — a refusal [`18-building-a-module.md`](../../03-DESIGN/01-to-be/18-building-a-module.md)
restates on the same page that it introduces `process`.
- **This is the repository's own named failure mode, in its own records.** `cycle.py` enforces
that a to-be doc names *at least one* decision. Both docs do, so both pass, while introducing a
resource type no decision records and contradicting an accepted one. The rule is "no design
without a decision"; the check is "no design without *a* decision." An unenforced rule is
indistinguishable from a wrong one, and these two documents are what that gap looks like when
something walks through it.
## Open questions
- Which is the decision — container or supervised unit? If the design docs are right, ADR 0047
needs superseding rather than quietly outliving. If ADR 0047 is right, two proposed documents
and a worked manifest describe a resource type that does not exist.
- Is one account per module satisfied by a per-module *unit* as well as a per-module *container*?
ADR 0047's argument rules out a node-wide runtime sharing one account. It does not appear to
rule out a unit holding one scoped credential, and nothing has said so either way.
- If several processes for one module are right, what holds the accounts? ADR 0047 refused "a
second one to scope and seal" for events beside tools. Four processes are four somethings.
- How does a `process` get its interpreter, and does declaring one reintroduce the machine
dependency the host is built to avoid?
- Is `run` an argv the link may carry, given `action` is refused for being one? If the answer is
that a `process` reconciles and an `action` does not, that distinction is not written down.
- What is the thing called, and where does the design layer describe it? Whichever shape wins, no
document in either layer currently says a code-carrying module runs a second thing beside its
service.
- **How would this have been caught?** A decision and a design doc disagreeing on a resource type
is mechanically checkable: the resource types a design doc names are a closed set, and every
member of it either appears in a decision or does not. Whether that check is worth writing is
part of this issue, not settled by it.
@@ -0,0 +1,201 @@
# Diagnosis — 117
## Which trees were searched, 2026-09-25
Named first, because [issue 113](../113-the-object-stores-images-were-withdrawn-upstream/01-diagnosis.md)
is the record of reporting absence in one repository as absence in the mesh.
| Searched | At |
|---|---|
| `mesh-host`, `mesh-catalog`, `mesh-tools`, `mesh-sdk`, `mesh-controller` | `main`, fresh shallow clones |
| `hq` | `main`, and the two branches named under finding 7 |
**Not searched:** the private migration repository; the open pull requests on the catalogue and
the controller; any branch of a code repository other than `main`. A statement below about "the
catalogue" is a statement about its `main`.
## The report's central question is answered: the shape exists
`mesh-host` `internal/declaration/declaration.go` defines `TypeProcess Type = "process"`.
`internal/apply/process.go` applies it — it writes the unit, writes the timer for a scheduled one,
and gates what follows a run-once one. It has tests of its own in both packages. The resource
carries a bundle `source` with a `digest`, a `run` argv, `env` and `env-file`, a `user`,
`restart-on`, and the `run-once` and `schedule` modifiers.
So the report's alternative — "if ADR 0047 is right, two proposed documents and a worked manifest
describe a resource type that does not exist" — is **disproven**. It exists, it is implemented, it
is tested, and the host's vocabulary is now **twelve** shapes rather than the nine
[ADR 0029](../../02-DECISIONS/0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md)
counted.
## The enforcement ADR 0029 asked for is intact, and it recorded this gap rather than closing it
ADR 0029 said "the vocabulary is nine, and the count moves with a record. The test that asserts it
names this one." That test exists — `internal/declaration/declaration_test.go` asserts the count is
twelve and fails with the reason rather than a number. Above the assertion, a paragraph per
addition names what made it one:
| shape | the test names |
|---|---|
| `network`, ninth | ADR 0029 |
| `access`, tenth | ADR 0051 |
| **the eleventh** | **`03-DESIGN/01-to-be/18-building-a-module.md`** — a design document, `status: proposed` |
| `opening`, twelfth | ADR 0100 |
The eleventh is this one. The test still calls it `daemon`, the code calls it `TypeProcess`, and
its paragraph is the only one that names a design document where the others name a decision.
Independently: in `declaration.go`, `TypeProcess` is the **only** shape in the vocabulary whose doc
comment cites no ADR — `network` cites 0029, `access` 0051, `opening` 0100, `user` and the refusal
of `action` cite 0005.
**So ADR 0029's mechanism worked exactly as designed and was not enough.** It requires every
addition to name something. It does not require that something to be a decision, and the one
addition that named a proposed design document instead is the one this issue is about.
### A correction to this trail, recorded because it was one grep from being a finding
The first search here was for `len(Vocabulary())` and found nothing, and the working conclusion for
two steps was that no count assertion existed any more — which would have been written up as "the
mechanism ADR 0029 relied on is gone." It is not gone. The test binds the slice to a local variable
first, so the assertion reads `len(speaks) != 12`. The claim was wrong, it was caught by reading the
file rather than by grepping it, and the shape of the error is the same one issue 113 recorded: a
negative search result read as a fact about the world.
## The argument the report asked for already exists, in a test comment
The report asked why a container rather than a supervised process, and said the reasoning was not
written down. It is — in `declaration_test.go`, as the eleventh shape's paragraph:
> Running code of one's own meant a `container` and therefore an image; running a script meant a
> `service` and a unit somebody else had to install. One intent — run this and keep it running —
> expressed two unrelated ways, with the hosting chosen before anything could be declared. […] It
> is a full-host shape rather than a portable one: it needs a process supervisor to install into.
> It does NOT need a container runtime, which is the point — only software that genuinely needs
> isolation asks for a container.
That is a decision's Context and Consequences, in a Go comment, in another repository. Nothing in
`02-DECISIONS/` contains it. `TypeProcess`'s own doc comment adds the rest — that three modes beat
three kinds, and that a first draft added a `daemon` for the long-running case alone.
## The catalogue is containers, and the one exception is the reference module
71 modules on `main`. Counting the `type` of every declared resource:
| `container` | `process` |
|---|---|
| 115 | **3** |
All three `process` resources are in **one** module: `showcase` — the module
[`20-writing-a-module.md`](../../03-DESIGN/01-to-be/20-writing-a-module.md) is a worked guide for.
### And in that module, the tools do not run
`showcase` declares its migrate, server and reporting steps as `process`. Its fourth resource, the
one for tools, is a **`container`** — and its image is the module's `helper` artifact, which the
same manifest declares as `kind: upstream` from a bare distribution base. Its command is
`sleep infinity`. It mounts the broker credential and sets the variable naming it, and runs nothing.
Meanwhile the module's `code` bundle declares six entrypoints. Three are run by the three `process`
resources. The tools entrypoint and the provisioner entrypoint are **run by no resource in the
manifest.**
Two consequences worth stating separately:
- **The worked guide does not match the module it documents.** The guide shows four `process`
resources, the fourth being `{"id": "tools", "type": "process"}`. The module has three and a
container.
- **This is the condition ADR 0047 was written to end, in a new shape.** That record's Context says
the conversion "produced tools and events that, as it stands, never execute," and its first
Consequence is that they become runnable. In the reference module they do not execute again —
not for want of a runtime this time, but because nothing declares one that runs them.
## The harness has no per-module boundary, and nothing refuses a second module
This is where ADR 0047's isolation argument is load-bearing, so it was checked rather than assumed.
- `mesh-sdk` `src/tools/index.ts`: `serveTools` iterates `collectTools()` over a module-level
registration array and serves **every registered module's** tools over the **one** `broker` it
was handed.
- `mesh-tools` `src/main.ts`: the modules to load come from one variable as a **comma-separated
list**, and the runtime sets its module and node identity from the **single** credential.
- `mesh-sdk` `src/events/index.ts`: an emitted event's `x-source` is stamped from that single
module identity.
Put together: load two modules into one runtime and everything the second emits is attributed to
the first, because there is one credential and the identity comes from it. That is precisely the
failure ADR 0047 predicted — "able to emit as any of them" — reached by a different route, since
the credential is correct and there is only one of it for two modules. **Nothing in either
repository refuses the second module**, and no test asserts that a runtime serves one.
### Ruled out, in fairness to the implementation
- **The serving key conforms.** ADR 0047 replaced a single `tools.invoke` dispatch with a per-tool
key, and the SDK does that: a tool is served on `<module>.<tool>` with the account scoped
`serve.<module>.*`. The superseded `tools.invoke` survives only in **prose** — the doc comment
directly above the conforming code, and the `mesh-tools` README, which also describes the runtime
as per-node. The code is ahead of its own documentation.
- **The credential shape conforms.** The sealed per-module credential file is preferred in code, and
the plain URL is documented as the bootstrap case before a module has an account — not the
ordinary path.
So the account is the right shape and the key is the right shape. It is the **process boundary**
that is declared nowhere and enforced by nothing.
## An unmerged report already asks the narrow version of this
Branch `issue/113-controller-container-or-process`, one commit, 2026-09-24, adds a report titled
**"Should the controller run as a container, or as a process the host supervises directly?"** with
`located-in: [mesh-controller module.json, mesh-host internal/apply]`. Its observation is that the
controller is declared a `container` with `network: host` — so container network isolation, the
property that resource type usually buys, is not in use — and it asks what `type: container` buys
that `type: process` would not.
It was unmerged and numbered 113, which is taken. A sibling branch,
`issue/113-record-the-repin-and-fold-114`, is why `114` was free.
**That report and this one are the instance and the general condition**, and they do not conflict:
it asks about one module that is not a code-carrying sidecar at all, and reaches the same question
from the opposite end. So it lands in this change as
[issue 114](../114-should-the-controller-be-a-container-or-a-process/00-report.md), its commit and
authorship intact, with a section pointing here — rather than being folded in and losing the
`network: host` observation, which is its own and is not reproduced above.
This diagnosis answers its first open question. The host's `process` shape does support what
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) describes for the host's own launcher — the
unit, the timer, restart, and run-to-completion gating are implemented and tested — so that report
does not need host-side work before it can be decided.
## What is located, and what is not
**Located — and it is not a code defect.** The implementation and the design layer agree with each
other; the **decision record is what is missing**, and the accepted record that occupies its place
says the other thing. ADR 0047 is `accepted`, cited by the module protocol, and unsuperseded, while
the host it describes has had a purpose-built shape for a module's own code since the eleventh
vocabulary entry.
| Owner | What is theirs |
|---|---|
| `hq` | the missing record for the `process` shape; ADR 0047 left standing; the worked guide that does not match the module |
| `mesh-catalog modules/showcase` | tools and provisioner entrypoints that no resource runs; a tools container that sleeps |
| `mesh-sdk src/tools/index.ts` | several modules served over one credential, unrefused and untested; a doc comment describing a superseded dispatch |
| `mesh-tools` | a README describing a per-node multi-module runtime the code no longer prefers |
**Not located, and deliberately open:** whether `process` or `container` is *right* for a module's
own code. This diagnosis establishes that the question was answered in practice and never recorded
— not which answer is correct. The arguments on both sides now exist in writing; they exist in a
test comment and a proposed design document, and one of them contradicts an accepted decision.
## What would close it
1. A decision record for the `process` shape, carrying the argument currently in
`declaration_test.go`, and saying what becomes of ADR 0047 — superseded in whole, or in the part
that names a container.
2. `18-building-a-module.md` and `20-writing-a-module.md` naming that record in `decisions:`, and
the worked manifest agreeing with the module.
3. The eleventh shape's paragraph in the vocabulary test naming a decision, like the other three.
4. **How the rule is checked, since a rule states how it is checked:** every shape in the host's
vocabulary names a decision, asserted where the count is already asserted — which turns "no
design without a decision" into something stronger than "no design without *a* decision" for
the one vocabulary where each entry is a security decision.
5. Whether a runtime may serve more than one module answered either way, and asserted — a refusal
if not, a test that two modules' events keep their own source if so.
@@ -0,0 +1,49 @@
---
status: located
opened: 2026-09-25
located-in: [mesh-catalog modules/umami]
fixed-by:
amended-design:
---
# 118 — umami's store answers the dial and times out the query
## What was observed
The analytics service's public name has answered `502` through the whole of 2026-09-25's migration session
(first noted mid-afternoon, still true at night). The container restart-loops on a
timescale of about a minute. Its own log, every cycle:
```
✓ DATABASE_URL is defined.
✓ Database connection successful.
Invalid `prisma.$queryRaw()` invocation:
Raw query failed. Code: `N/A`. Message: `Operation has timed out`
```
The connection is established — the dial succeeds — and the first raw query then times
out. This is not a credentials fault and not an unreachable store.
## What it is not
- Not the routing layer: the `502` is Traefik faithfully reporting a backend that is
restart-looping. The stale duplicate Traefik router for this name (a HAL-era
hand-authored file beside the mesh-written one) was removed the same night and changed
nothing, as expected.
- Not the mesh's grant machinery: the binding and sealed secret compose, and the store
accepts the login — a wrong credential refuses the dial, and this dial succeeds.
## Where to look
A dial that succeeds and a query that times out, from a container on one network to a
store on another, has the shape of a path-MTU or conntrack fault (large response packets
dropped after the small handshake ones pass), or of the store accepting the TCP
connection while the backend it proxies for is wedged. Neither is proven. What is known
to differ for umami against every working consumer of the same store tonight is nothing
yet — that comparison is the first move.
## Why it is filed rather than chased
The 2026-09-25 session's scope was routing and the build chain; this fault predates the
night's changes, survived them unchanged, and needs its own sitting with the store's own
logs beside the consumer's.
@@ -0,0 +1,135 @@
---
status: located
opened: 2026-09-25
located-in: [mesh-catalog modules, mesh-controller internal/catalogue]
fixed-by:
amended-design:
---
# 119 — A module definition decides where its files live on the machine
## What was observed
A review of where module code reads its files turned up a cross-cutting pattern. Every module
definition in the catalogue chooses, in its own manifest, where on the machine its files live.
Counted on the catalogue's `main`, 2026-09-25:
| where in the definition | host-path strings |
|---|---|
| directory and file resources | 257 |
| container mounts, host side | 230 |
| own secrets | 78 |
| bindings | 53 |
| env-files | 50 |
| secrets | 35 |
| container environment | 28 |
| accesses | 21 |
| receives, grants | 24 |
| everything else | 13 |
**789 host-path strings in 70 of the 71 definitions.** Mounts are checked: a container may not
mount a path its module never declared ([ADR 0091](../../02-DECISIONS/0091-a-mount-is-declared-three-ways.md)).
Nothing checks the same path where it is retyped as a value: an environment variable, an env-file
line, a literal in module code.
## 2026-09-26 — counted by what the path *is*, and what a node's layout would replace
The count above treats every host path the same. They are not the same, and only one of the
categories is this issue's subject. Recounted by role across the 71 definitions:
| What the path is | Count | Does a definition naming it name *this* machine? |
|---|---|---|
| A module's own data, `/var/lib/<module>/…` | 289 | **Yes** — this issue |
| What the mesh writes for a module, `/var/lib/mesh/<module>/…` | 225 | **Yes** — this issue |
| The operator's shared data, `/services/…` | 166 | Separately answered: an `access`, never created or owned by the mesh ([ADR 0051](../../02-DECISIONS/0051-shared-data-is-the-operators.md)); *where* it is belongs on the assignment |
| A system file the mesh owns, `/etc/…` and a few under `/var/` | 18 | **No.** The path is the fact — that file is at that path on every machine of the kind. Naming it says nothing about this installation |
| A path inside a container | 236 | **No.** The software's own contract, true in any mesh that runs the image |
So the category a node's default layout would replace is **514**, not 789 and not 698: a module's own
data, and what the mesh writes for that module. The rest is either already answered or was never the
problem.
## What the design already says, and what it does not
[ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md) (proposed)
makes a directory a host provision whose contract is the owner and mode the module needs, and says
where it lands on the machine is the assignment's.
[To-be 27](../../03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md) (proposed) resolves
that two ways: **a node's default layout — a root per node with one directory per assignment beneath
it** — used when the assignment says nothing, and **a placement** when the assignment puts one
directory elsewhere, such as where an adopted machine's data already is.
That is the reservation model: assigning a module to a node gives it a directory, and the definition
never names one. Three things are not written anywhere:
1. **Where the root comes from.** Nothing says it is a node setting, fixed when the node is installed,
and nothing says what a node that states none falls back to.
2. **What sits beneath it.** To-be 27's own open list ends at exactly this: the layout beneath the
root, beyond one directory per assignment. The catalogue currently keeps two subtrees per module —
its data, and what the mesh writes for it — so this decides whether an assignment's directory holds
both, or the mesh keeps its own.
3. **The order the existing 514 are retired in**, which 0112 also leaves open.
## Why this is not a manifest edit
Every one of those 514 paths names a directory that holds data a service is using. A definition that
changes where it looks, without the data moving with it, does not fail: the mesh creates the new
directory with the right owner and mode, the container starts, and the service comes up **empty**.
The loudest case in this mesh is an object store consumer whose bucket holds 174.9 GiB; a mail spool
and the mesh's own store are the same shape.
So the retirement of these paths is a **data migration with a verification step**, module by module,
not a change to a definition. That is the reason this section documents and stops: the count and the
categories are what someone needs to plan it, and the planning belongs with whoever can see the
machine.
### Where that has already gone wrong
- **A provider that would provision nobody, silently.** One DNS provider mounts its grants
directory at a short path inside its container, then tells its provisioner to read the
contributions file at the host path, which does not exist in there. Nothing requires the
provision today, so it has not failed yet. When a consumer arrives, it will get no record, and
nobody will be told.
- **The mesh's own wire carries host paths into containers.** Each contribution names its
consumer's credential as "the file on this machine holding that consumer's credential", a host
path computed from the provider's grants directory. So every provider has to mount that directory
at the *identical* path, or it cannot read what it was given. Ten of the eleven providers with a
grants directory do. It is a convention nothing states or checks, and the eleventh is the
provider above.
- **The warning that would have caught it is lost in the SDK.** The controller always writes the
contributions file, even when empty, so a provider can tell "nothing asked" from "never written".
The SDK's reconcile loop treats an unreadable file as empty, and logs nothing.
- **Code carries copies with nothing checking them.** Several modules default a path in code when an
environment variable is unset. Five of those defaults disagree with the value their own manifest
sets. One of them is a host path used inside a container that does not mount it.
### And a module cannot be assigned to one node twice
Everything that identifies a running module is keyed by the module's name: its directories, its
container names, the login it presents to a provider, its broker account. Two assignments of one
module to one node would share every one of them. Assigning the same application twice is an
ordinary need: production beside staging, one site per customer, two instances of one service
configured differently, two stores of one engine.
[ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md), which answers
this report, declines that need rather than meeting it. A module is assigned at most once to a node,
because every identity in the mesh is already a module on a node. The cases above become different
modules, or the same module on different machines.
## Why it matters beyond this instance
A definition that names machine paths is not portable between nodes. It cannot follow data onto a
second disk, or onto a machine being adopted with its data already in place, without editing the
module. It cannot run twice on one node. It keeps every path in two or three places with nothing
checking that they agree. The defects above are what that allows, and each was found by reading,
not by any check.
## Open questions
- Should a definition name any host path at all, or should every location come from the
assignment and the mesh?
- If a directory is something a module *requires* rather than *declares*, what is its contract:
ownership, mode, whether it is kept when the module goes?
- What identifies an assignment, if a module may be assigned to one node more than once?
- What would the contributions file carry instead of host paths, so a provider needs no
identical-path mount?
@@ -0,0 +1,62 @@
---
status: located
opened: 2026-09-26
located-in: [mesh-sdk src/provisioner, mesh-catalog modules/redis]
fixed-by:
amended-design:
---
# 120 — A provisioner remembers what it did, not what is there
## What was observed
The provisioner harness every provider is built on keeps, in memory, a hash of what it last applied
for each consumer: the login, the password and the values. On each pass it skips a consumer whose
hash has not changed. It never asks the backend whether what it made is still there.
The cache module shows what that allows. Its server is configured with a password and a data
directory, and **no ACL file**. So the per-consumer ACL users its provisioner creates exist only in
the server's memory. The server and the provisioner run in separate containers:
1. the provisioner creates an ACL user for each consumer, and records it as applied;
2. the server restarts, for an upgrade or a crash, and comes back with no consumer users;
3. the provisioner, still running, sees nothing changed in what it receives, and does nothing;
4. every consumer of the cache fails to authenticate, and **nothing reports it**. The provisioner's
log is quiet, and the mesh's status is green.
The consumers recover only when the provisioner itself restarts, because its memory is then empty.
Rotating the cache's administrative password happens to cover it, because that file is mounted into
the provisioner too and recreates it. Nothing else does.
Evidence, from the catalogue's and the SDK's main branches: the harness's reconcile loop (`applied`,
keyed by login, compared by hash before `create`), and the cache module's rendered configuration,
which names no ACL file. Found during research 016, how a credential can be rotated, proposed
alongside to-be 27.
## Why it matters beyond this instance
The cache is the case where the backend forgets on its own. The same gap opens whenever a backend
loses what was provisioned while the provisioner keeps running: a store restored from a backup taken
before a consumer was added, a login removed by hand, a server recreated on an empty data directory.
In each of them the provisioner reports that everything is applied, because it compares against its
own memory and not against the backend.
The harness's other half has the same shape. A consumer's contribution that disappears while the
provisioner is down is never removed, because only logins the running process applied are candidates
for removal. What the mesh wants and what the backend holds can drift in both directions, and the
harness sees neither.
This is the design permitting a silent failure. *A provider makes what its consumers require true*
is stated in [to-be 13](../../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md), and nothing
checks it after the first pass.
## Open questions
- Should the harness check each consumer's credential against the backend on every pass, or
periodically, instead of trusting its memory? Most adapters' `create` is already idempotent, so the
cheapest fix may be to drop the hash short-cut and apply every pass. What does that cost for a
provider that recreates an access key on every create, as the object store does?
- Should the cache keep its users in an ACL file, so a restart does not lose them? That fixes this
instance and leaves the gap for the others.
- Where does the record of what was applied live, if not in memory? ADR 0114, still
proposed, puts rotation state with the vault. The same place may answer this.
@@ -0,0 +1,83 @@
---
status: located
opened: 2026-09-25
located-in: [mesh-catalog modules/builder, mesh-catalog modules/gitea, mesh-controller internal/catalogue/resolve.go]
fixed-by:
amended-design:
---
# 121 — `builder` requiring a real `package-registry` grant deadlocks a genesis bootstrap
## What was observed
On the control-node, 2026-09-25, fixing `builder`'s hand-faked `package-registry` binding (a
hardcoded JSON fragment standing in for a real grant, found live-blocking nothing but carrying no
correctness at all — see the migration log's account of the same night). The fix declared
`requires: ["artifact-store", "package-registry"]` on `builder` and let the mesh mint the grant
properly, mirroring how `artifact-store` was already required.
Running `mesh-controller`'s own test suite against the catalogue as changed surfaced three tests
that pinned a different, deliberate design:
```
--- FAIL: TestTheBuildersCarriedPackageBindingTakesThePortFromTheNode
--- FAIL: TestTheBuildersCarriedBindingStartsWhereTheForgeServes
--- FAIL: TestTheBuildersPackageBindingKeepsItsIdentity
```
Read together with their own comment in `foundation_manifests_test.go`:
> The builder carries a binding because at genesis nothing provides `package-registry` to resolve
> one from; the day the forge is a module, the same consumer is told what the forge serves.
The tests expect `builder` to carry its own `package-binding` resource — a `merge: json` file with
`protected: [provision, from, at, as]`, settable only on `port` via the ordinary settings layers —
matching defaults (`port: 3000`, `scheme: http`, `npm-path: /api/packages/<owner>/npm/`, `as:
mesh-builder`, `from: gitea`) that agree with what `gitea`'s own `serves.package-registry` declares.
Tonight's fix removed that resource entirely.
**The dependency this protects against is real, not hypothetical.** `gitea`'s own `module.json`
carries a `build` section — its runtime image is compiled from a Dockerfile via `mesh-tools`'s
build/runtime bases, the same as every other built module. `builder` is what runs that build. So:
- `builder` now `requires: package-registry`, satisfiable only by `gitea`.
- `gitea` cannot run — cannot exist as a container at all — until `builder` has built its image.
On the control-node tonight this is invisible: the mesh is already running, `gitea` is already built and
assigned, and the grant resolves immediately. A genesis bootstrap — a new mesh raised from nothing,
or that node fully re-raised for disaster recovery — hits the order the tests describe: `builder` is
needed to build `gitea`'s image before `gitea` can be assigned, so `builder` cannot yet hold a real
`package-registry` grant, so (as tonight's fix has it) `builder` cannot start.
## Why it matters beyond this instance
This is the same shape ADR 0075 already named for the mesh as a whole ("the bootstrap still has no
package registry... that is the same pivot as everything else and it is not solved here") — but it
now has a concrete second collision: the one module capable of building a package-registry
provider's own image would refuse to run before that provider exists, precisely because it was
made to depend on it correctly.
It is also a near miss worth writing down for its own reason: the fix that broke this was reviewed
against `go build`, the full test suite (which flagged it immediately — three tests, not zero), and
matched an explicitly-stated target design in `22-the-work-ahead.md` written a week earlier. Nothing
about the fix was careless; the deadlock was invisible because the node it ran against had already
crossed the point where it would bite.
## What is not decided here
- Whether `builder` should carry both — its own default/fallback binding for when no real grant can
be resolved yet, and a real `requires: package-registry` for when one can — and if so, what
decides which one is live at a given moment. Nothing in this mesh's `requires`/`provides`
resolution currently expresses "this, or a carried fallback if it cannot resolve" for any
provision; this would be the first.
- Whether the fallback belongs on `builder` specifically (matching what the removed tests already
encoded) or is a shape the mesh should offer any module bootstrapping a provider it also depends
on — the "builder builds the thing it needs to ask a favour of" pattern is not unique to
`package-registry` and could recur.
- Whether genesis itself should sequence around this instead (build gitea before granting anything,
with `builder` never resolving `package-registry` during that specific window) rather than the
manifest carrying two shapes of one credential.
Tonight's `builder` fix (`mesh-catalog modules/builder/module.json`) is left in place — correct for
the steady state, live and working on the control-node — with the three tests above left failing rather than
reverted or hacked to pass, so the gap stays visible rather than quietly patched over.
@@ -0,0 +1,94 @@
# 121 — Diagnosis
## 2026-09-26 — the seats work renamed the requirement; it did not reorder the bootstrap
The seats change landed the same day this was filed (ADRs 0109–0111 accepted, and the code on both
main branches), and it touches exactly the two manifests this report names. That made it worth
asking first whether the deadlock had been fixed in passing. It has not.
### Step 1 — what the seats work changed
`builder`'s requirement was renamed from `package-registry` to `npm-package-registry`, which is now
a **seat** the forge holds; the forge claims `npm-package-registry` and `git` at mesh scope and
serves both. The provision is the same relationship under a name the mesh defines rather than one a
manifest invented.
### Step 2 — the seat's holder cannot answer when there is no holder yet
The seat is consulted in exactly one branch of resolution: where **several** nodes provide the
wanted provision and no pin says which. That is the ambiguity case, and it is not the bootstrap
case. With nothing providing it at all, resolution takes the branch for no providers and **refuses**,
naming the remedy:
> nothing in this mesh provides "npm-package-registry", wanted by builder — assign gitea to a node
A refusal, not a deferral. So a seat delivering a provision does not make a requirement optional
before the seat is filled, and nothing about the seats work moves the order.
### Step 3 — ruled out: genesis has no escape hatch
Resolution has an `Unchecked` mode that "takes brokered requirements on trust instead of refusing
when nothing answers them", which would be exactly the exemption a bootstrap needs. It is not one:
it exists for the first of two passes — working out what each node offers — and its own comment
states that **a declaration is never built from an unchecked resolution**. The second pass, the one
that produces the declaration, refuses.
### Step 4 — the other arm of the cycle is intact
The forge's manifest still carries a `build` section, its runtime image compiled from the
toolchain's build and runtime bases like every other built module. `builder` is what runs that
build. So both arms stand: the builder cannot be assigned before the forge provides the registry,
and the forge cannot exist as a container before the builder has built its image.
## 2026-09-26, later — Step 4 is retracted: the forge's service is not built
Step 4 said the forge's runtime image is compiled by the builder, so the forge cannot exist as a
container before the builder has built it. **That is not what the forge's definition says**, and it
took one command to disprove.
The forge's **server** is an upstream public image pinned by digest. What the mesh builds is the
forge's **runtime sidecar** — its own agent beside the service, a separate container. The store and
the broker have exactly the same shape: an upstream server image, a built sidecar. And that shape is
the point, because [ADR 0078](../../02-DECISIONS/0078-the-store-and-broker-are-modules.md) raises the
store and the broker at genesis as plumbing and **adopts them in place** as ordinary modules
afterwards. Nothing stops the forge being raised the same way: its service can be serving git,
packages and OCI before anything at all has been built.
So the second arm of the cycle, as this diagnosis stated it, does not exist.
**What survives is narrower and is still worth answering.** A grant — the credential the builder
would use against the package registry — is minted and delivered by the **provider's runtime
sidecar**, and that sidecar is built. So the question is whether a consumer can be granted a
provision whose provider's sidecar is not yet running, and whether the ordering
*raise the service → grant → build the sidecar* simply works. That is sequencing inside the mesh's
own machinery, not a chicken-and-egg about images, and it may not be a deadlock at all.
The report's open questions should be read against this. In particular, a carried fallback binding
may be answering a problem that adopt-in-place already solves.
**Why the error is recorded rather than edited away:** the wrong version was reached by taking a
record's bootstrap argument at face value instead of checking it against how the store and the broker
are actually raised — which is the one comparison the manifests make available in a single command.
## What did change, and why it is the part worth recording
The three tests that pinned the carried-fallback design — the builder's own `package-binding`
resource, its port taken from the node, its identity kept — were **deleted** by the controller's
seats PR and replaced by seat-based tests that pass.
The report deliberately left those three failing so the gap would stay visible. It is no longer
visible: the suite is green, and the only remaining trace is this issue. That is the same way the
deadlock was invisible in the first place — a running mesh resolves the grant instantly, because the
forge is already built and assigned. Twice now, the evidence has been removed by something correct.
## What would actually answer it
- **A provision that resolves to a carried fallback when nothing provides it yet.** Nothing in the
model expresses "this, or a carried default until something can answer", for any provision. This
would be the first, which is why it belongs in a record rather than in a manifest.
- **Genesis sequencing instead** — build the forge before granting anything, with the builder never
resolving the registry during that window.
- **The forge's runtime image not being the mesh's to build** — an external image breaks the cycle
by removing the dependency rather than by expressing it.
None is chosen here. The report's open questions stand as written.
@@ -0,0 +1,116 @@
---
status: located
opened: 2026-09-26
located-in: [mesh-controller internal/catalogue/declaration.go, mesh-catalog modules/keycloak, mesh-catalog modules/minio, mesh-catalog modules/nextcloud]
fixed-by:
amended-design:
---
# 122 — A module cannot ask for its own public name, so three manifests wrote this mesh's names into the catalogue
## What was observed
Reviewing the open module changes on 2026-09-26, three of them independently put a name belonging to
**one particular mesh** into a module definition, each for a different piece of software and each as
the only way to make the software work:
- The identity provider's manifest gains `KC_HOSTNAME: https://<label>.<public-domain>` as a literal,
because the software generates absolute URLs and, behind a proxy, cannot derive them.
- The object store's manifest gains `MINIO_BROWSER_REDIRECT_URL: https://<label>.<public-domain>` as
a literal, for the same reason — its console redirects to an absolute URL.
- The file-sync module's S3 bucket is renamed from `nextcloud` to a name carrying this mesh's own
name, so the module matches a bucket that already exists here.
No reviewer introduced these carelessly: each is the value the software needs, and there is nowhere
else to put it.
**And they are not the first.** Asked whether merging them would set a precedent, the catalogue
answered no: **five modules already on the main branch carry one**. The clearest is the workflow
automation module, whose environment file states its host, its protocol and a full absolute webhook
URL as literals; two more — the proxy and the builder — carry a complete clone URL for a repository
on this mesh's own forge, scheme, host and port included.
**Counted properly, 2026-09-26.** The five were what a first look found. A sweep of all 71
definitions, for values that could only be true of one installation:
| What is named | Occurrences | Modules |
|---|---|---|
| A public domain, or a name under one | 49 | 26 |
| The node's own name | 58 | 19 |
| A routable IP address | 12 | 1 |
| A host path for a module's own files — the node's to decide | 514 | 70 |
| An email address | 0 | 0 |
The path row is [issue 119](../119-a-module-definition-decides-where-its-files-live/00-report.md),
counted again — and counted differently, which is the correction worth keeping. **A path inside a
container is not a fact about the machine.** `/run/secrets/…` and the directory a server keeps its
data in are the software's own contract, true in any mesh that runs it; only the host side of a mount
names where it landed. The first sweep matched path-shaped strings and so counted both halves of every
mount and every in-container location a value mentioned. Counted by role instead — directory and file
resources, the host side of mounts, accesses, and the targets of binds, grants, receives and secrets —
it is 698 across 70 definitions — of which **514** are the category at issue. The rest are a system
file the mesh owns (`/etc/…`, where the path is the fact and is the same on every machine) and the
operator's shared data, which ADR 0051 already answers as an access. Issue 119 carries that breakdown
and what a node's default layout would replace. Issue 119's own figure is role-counted already and close to this; it
additionally counts paths written into environment values, a few of which are container-side.
The rest of the table is this issue. **Thirty of the seventy-one definitions name this
installation** in one of the first three ways, and the worst single case is not a domain at all: a
mail module states the node's own public IPv4 as the address it trusts a real-IP header from, so a
node that moves, or gains a second address, silently stops attributing mail to the right sender. A
private subnet, the mail domain, the site name and the administrator's domain sit in the same block
of values.
Two upstream resolvers are excluded from the count deliberately: naming a public DNS service is a
policy default, true of any mesh that wants it, not a fact about this one.
So the three under review are the visible edge of a pattern the catalogue already follows, which
changes what this issue is for. It is not a matter of refusing three changes; there is no version of
this catalogue today that does not name the mesh it was written in, and a mechanism is the only
thing that removes any of them.
## What the mesh offers instead, and why none of it answers
A manifest may interpolate `${secret:…}`, `${seat:…}`, `${bound:…}`, `${port:…}` and
`${machine:…}`. The machine form resolves `at`, `name` and `address` — the machine's identity on the
**private** network. None of them yields a public name.
The mesh does compose public names: a route contribution's label is joined to the node's public
domain, `<label>.<public-domain>`, and the mesh interprets neither half. That happens inside the
controller when a declaration is built, and the result reaches the proxy that serves the route. It
does not reach the module that asked for the route, so a module that must **tell its own software**
what it will be reached at cannot read what the mesh already worked out.
So the workaround is the only expressible option: write the answer down, in the definition, for the
one mesh it is true of.
## Why it matters beyond these three
A module definition is meant to hold what is true of the module everywhere, with everything
particular to an installation resolved at assignment — the argument
[ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md) (proposed)
makes in full, and what [issue 119](../119-a-module-definition-decides-where-its-files-live/00-report.md)
found for paths. A hostname is the same kind of fact as a path, and arriving by the same route: not
because anyone disagreed with the principle, but because the mechanism that would honour it does not
exist yet.
The cost is concrete. A second mesh installing the identity provider from this catalogue gets the
first mesh's hostname, and its login flow breaks in a way that looks like a proxy fault. Nothing
checks for it: a literal hostname in an env value is a well-formed string, and no rule distinguishes
it from a version number.
The bucket case is the same shape with a different subject — an adopted resource's real name is also
particular to one installation, and also has nowhere to live but the definition.
## Open questions
- Should a module be able to name what it will be reached at — a `route`-scoped interpolation
resolving to the composed public name of a route it contributes, so the answer the mesh already
computes is the one the software is told? That is the smallest fix and it stays within the existing
vocabulary.
- Does the scheme belong to it? Every instance here wrote `https://`, which is true of a route the
proxy terminates and not of the module's own port.
- For an adopted resource such as an existing bucket: is that a setting on the assignment (where
ADR 0112 would put it), and if so what reads it — the provisioner, or the module's own values?
- What check would notice the next one? A definition naming a public domain is detectable in the
shape of the value, which is more than nothing, and less than a rule.
@@ -0,0 +1,74 @@
---
status: located
opened: 2026-09-26
located-in: [hq 00-META/glossary.md, hq 02-DECISIONS/0075-two-stores-and-which-provides-what.md, mesh-controller internal/catalogue/seats.go, mesh-catalog modules/distribution]
fixed-by:
amended-design:
---
# 123 — The image registry is named after a role, and *artifact* is defined as one format
## What was observed
Reading the registry work back to the operator on 2026-09-26, three wordings turned out to disagree
with each other, and the disagreement was doing real damage: a reader — including whoever writes the
next module — cannot tell what the mesh's image registry is for from what it is called.
**One: the word for a delivered thing is defined as one format.** The glossary says *artifact* is
"what the mesh delivers to a machine to install and run: **an OCI image, by digest**", served by the
`artifact-store`. But a module definition's own build vocabulary already names four kinds of artifact,
all in use in the catalogue today — `image` (48 of them), `upstream` (3), `bundle` (1), `archive` (1) —
and the controller's manifest code documents the field as "image or archive". So the catalogue builds
artifacts that are not images, while the word for them means image, and the store named after the word
serves images only.
**Two: the seat is named after a role, and every other one is not.** [ADR 0079](../../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md)
names the foundation seats after their servers — the store, the broker. [ADR 0109](../../02-DECISIONS/0109-a-package-registry-seat-is-one-per-ecosystem.md)
settled that a package registry seat is one **per ecosystem**, not one for all of them, which is why
there is now a seat for one language's packages and another for git. The image registry's seat is
`the-artifact-store`: named after neither its server nor what it serves. It is the last registry named
after the job it happens to be doing.
**Three: a module is not an image.** An image is one thing a module may install or build; a module may
also ship archives, files, directories, databases and a service it did not build at all. Prose that
calls a built image "the module's image" reads as though a module *were* an image, and the manifest's
own vocabulary does not: it says artifacts, each with a kind.
## Why it matters
The cost is confusion, in the place where the mesh explains itself. Two people reading
`artifact-store` will reasonably take it for two different things — "where OCI images live" and "where
everything a module delivers lives" — and those stop being the same thing the moment a module delivers
an archive, which one already does.
It also hides a question that ought to be asked plainly. Because the store is named after a role, a
reader assumes the role needs its own implementation. What the records actually say is narrower:
[ADR 0075](../../02-DECISIONS/0075-two-stores-and-which-provides-what.md) keeps two provisions because
packages and images are two protocols — a language's client cannot install from an OCI registry — and
it **already allows the forge to provide `artifact-store` too**, saying a mesh may choose it. The
reason the OCI registry is a second implementation is a bootstrap argument: something must serve images
before the mesh can build anything.
**That bootstrap argument is weaker than it reads.** The forge's server is an upstream public image;
only its runtime sidecar is built ([issue 121](../121-builders-real-package-registry-grant-deadlocks-genesis/01-diagnosis.md),
where the opposite was claimed and retracted). The store and the broker have the same shape, and
[ADR 0078](../../02-DECISIONS/0078-the-store-and-broker-are-modules.md) raises both at genesis as
plumbing and **adopts them in place** as ordinary modules. The forge can be raised that way too, which
would leave the question of a second registry implementation resting on the window before it is
adopted, rather than on protocols.
## Open questions
- Should the seat and the provision be named after what they serve — an OCI registry, images by digest
— leaving *artifact* free to mean what the build vocabulary already means by it: anything a module
produces, of any kind?
- Does the mesh need two registry implementations at all, or can the forge hold every registry seat
once it is adopted in place? The protocol argument keeps two **provisions**; it does not by itself
require two **servers**.
- If a window before adoption is the only reason for the second one, is that window a **seat**, or a
step in genesis that ends?
- Renaming touches the glossary, two records, the seat set, the builder's requirement and every image
reference in the catalogue. Which of those can be checked mechanically, so the rename is not a prose
exercise that leaves the code disagreeing?
- What check would keep the glossary honest — a definition tested against the kinds a definition may
actually declare, rather than restated by hand?
@@ -0,0 +1,65 @@
---
status: located
opened: 2026-09-26
located-in: [mesh-controller internal/catalogue/declaration.go, mesh-sdk src/provisioner, mesh-catalog modules/minio]
fixed-by:
amended-design:
---
# 124 — A consumer cannot be told a value its provider derived for it, so it transcribes one
## What was observed
The object-store interface gives each consumer a bucket. The provider **derives** which one: it
takes the login the mesh minted for that consumer and normalises it into a bucket name, then creates,
checks and removes exactly that. Its own comment states the reasoning — derived from the login, so
teardown recomputes it with nothing to persist — and it never reads the bucket a consumer's
definition named.
The consumer still has to tell its own software which bucket to use. It has no way to be told, so
all three consumers wrote the answer down by hand. Two transcribed it correctly. The third named the
predecessor's bucket, while the key the mesh scopes for it admits only the derived one: it would have
authenticated successfully and been refused on every object, which reads like a credential fault and
is not one. That was found by reading, not by running, and it was filed as part of the fix.
## Why the value cannot arrive
A definition may interpolate five things: a secret, a seat, a port, a machine fact, and a **bound**
value — what a provider said a consumer must know. Everything about the arrangement is already
delivered that way: the endpoint, the port, the region, the login itself.
The bound values come from what the provider `serves`, and **`serves` is a literal block in the
provider's definition**: the same values for every consumer. A value derived *per consumer* has no
way into it. Nothing else can carry it either — a provisioner's contract takes the provision and
returns nothing, so what it derived stays inside it.
So the mesh has a value it computed, a consumer that needs exactly that value, and no channel
between them. The gap is not that the answer is unknown; it is that the answer cannot be passed.
## Why it matters beyond this instance
Every interface where the provider names the resource has this shape. A database provisioner that
prefixed database names, a queue provider that scoped vhosts, an object store that derives buckets —
each would force the consumer to reproduce the provider's rule in its own definition. That is a copy
of somebody else's logic, kept in agreement by hand, which is the failure
[issue 119](../119-a-module-definition-decides-where-its-files-live/00-report.md) records for paths
and this one records for names.
It also shifts a rule out of the place that can enforce it. The provider knows its own naming rule
and can refuse a bad one; a transcription in a consumer's definition is a string, and no check
compares it to what the provider will actually create. The one wrong instance was well-formed.
## Open questions
- Should a provider be able to return consumer-visible values from provisioning — the natural
channel, since the provider is what derived them, but it means a grant carries data the provider
wrote rather than only data the mesh minted.
- Or should `serves` be able to say a value is derived from the consumer's identity, with the
mesh performing the derivation — which keeps the provider declarative, and requires the mesh to
know normalisation rules belonging to somebody else's protocol. It already knows one: the identity
limit is 20 characters *because* of what an object store's access key accepts.
- Either way: should a consumer that names a resource its provider will not use be **refused** rather
than ignored? The contributed bucket was read by nothing, and looked authoritative for months.
- What would have caught the wrong instance? A test that resolves a consumer's grant and compares the
bucket in its own configuration against the one the provider would create is a check that could
exist today, for any interface, without the mechanism above.
@@ -0,0 +1,59 @@
---
status: open
opened: 2026-09-26
located-in: [mesh-host internal/apply, mesh-controller]
---
# 125 — a hold is not a line in the apply report, and an operator flew blind into an outage
## What was observed
During the route-proxy edge cutover on novox (2026-09-26): the module was assigned, the
push reported success, `status` said the node was doing everything it was told — and the
module's three containers did not exist. The operator stopped the predecessor's proxy on
the strength of those reports, and every public name on the node went dark until rollback.
The cause was correct behaviour, invisibly reported. The first (rolled-back) route-proxy
attempt had left `/var/lib/route-proxy/*` on disk; on re-assign, the adopted node *found*
those directories, held them (ADR 0100, exactly as designed), and held every container
that mounts them — `"would mount /var/lib/route-proxy/ca, found on this adopted node;
not run until route-proxy is taken"`. All of that lived only in `state.json`. What the
operator saw:
- the push: `sent novox 346 resource(s)` — the controller's count of what it sent;
- the node's journal: `applied 330 resource(s)` — sixteen fewer, with no line saying
which sixteen or why;
- `status`: green — a held resource is not "wrong", so nothing was flagged;
- `node show novox`: the holds list did NOT include route-proxy's (it showed only holds
the *controller* knew about from take-time listings, not what the node decided at
apply-time).
Four surfaces, none carrying the one sentence that mattered: *route-proxy is assigned
but not taken, and its containers will not run until it is.*
## Why this is a real fault and not operator error alone
The operator error (an edge-flip runbook that omitted `take`) was only possible because
every surface reported success. A system whose correct refusals are indistinguishable
from completed work will keep converting small procedural gaps into outages. The
`sent 346 / applied 330` discrepancy was the single visible symptom, and interpreting it
required reading `state.json` by hand.
## What would have prevented it
Any one of:
1. **The apply report says what it held.** `applied 330 resource(s), 16 held for
untaken modules (route-proxy: 13, …)` — one line in the journal.
2. **`status` counts holds against untaken-but-assigned modules.** A module assigned,
pushed, and running zero of its containers is at minimum worth a "waiting on take"
line — it is never converged in any useful sense.
3. **`node show <node>` shows the node's own held list**, not only what take-time
computed — the node already records it in `state.json` with reasons.
## Precedent
The photos cutover hit the same semantics benignly the same week (assign → held
containers in `Created` state → take), and the mailu cutover documented "take is the
verb, and ADR 0100 meant it". The semantics are consistent and right; the reporting is
what let them be forgotten at the worst moment.
@@ -0,0 +1,47 @@
---
status: open
opened: 2026-09-26
located-in: [mesh-host internal/apply]
---
# 126 — a volume path is not in the spec comparison, and a roll-out raced a data move
## What was observed
Landing the "module data lives in /var/lib" change on novox (mesh-catalog #97), two
distinct faults surfaced in one hour:
1. **Building a module with a roll-out upgrade policy IS deploying it.** gitea's policy
was roll-out; the `build` that registered its repathed manifest sent it to the node
immediately, which recreated the container mounting the *not-yet-renamed* (empty)
`/var/lib/gitea/data`. The forge came back as its own install page, fresh host keys
and all, and every subsequent pipeline build died on `repository not found` — which
also blocked the fix, since re-registering the other modules needed the forge. The
operator narrative "build, then move data, then push" is only safe under the record
policy; nothing warned that one module in the batch would skip the pause.
2. **Changing a container's volume paths does not recreate the container.** After the
final push, five of the six repathed modules kept their old containers running
("Up 13–26 hours") — the new declaration's volume paths differ from the running
containers' mounts, and the apply judged them current. Same class as mesh-host #27
(`dns`/`ip` absent from the comparison): a field the comparison does not read is a
field that can never change a running container. Benign here only because a rename
on one filesystem preserves the mounted inode — the running containers keep serving
the same bytes the new path names, and the next natural recreation converges. A
cross-filesystem move, or a path change to *different* data, would have silently
split the module between two worlds.
## What would have prevented it
- `build` printing the module's upgrade policy when that policy will act on the result
("gitea rolls out on build — the node will receive this immediately"), or a
`--register-only` flag for exactly this choreography.
- Volumes (and every other container field) in the spec comparison, or the honest
refusal: "this field changed and I cannot apply it without recreation."
## Recovery that worked
Instant renames both ways broke the circular dependency (forge needed for builds,
builds needed for the push, push needed for the forge): data back to the old path,
old-spec forge started, artifacts rebuilt, data renamed forward, push. Nothing lost;
the install-page junk was discarded twice.
@@ -0,0 +1,39 @@
---
status: located
opened: 2026-09-27
located-in: [mesh-controller cmd/mesh-controller/push.go]
---
# A declaration that shrinks to empty is skipped, so the node keeps what it should drop
## What was observed
Fixing the broker-opening leak (the foundation port scoped to the broker's host) made ace's
declaration compose to **zero resources** — ace is adopted with nothing assigned, and the
stray opening was its only resource. `push ace` then printed `ace is assigned nothing —
skipped` and sent nothing. ace goes on holding `adoption.opening-tcp-5671-incoming` in its
ufw, because it was never told the resource is gone.
`composeEach` (push.go) skips any node whose composed declaration has no resources. That is
right for a node that never had anything. It is wrong for a node that **had** resources and
now composes to none: the empty declaration is the correction, and skipping it leaves the last
non-empty one in force forever.
## Why it matters
Any adopted node whose openings (or other baseline resources) are all removed keeps the stale
ones until something else pushes a non-empty declaration to it. Converge is unaffected — it
composes the full ruleset fresh — so this is an incremental-push gap, not a firewall-safety
one. But "the mesh cannot tell a node to drop its last resource" is a real hole in reconcile.
## The fix, roughly
Send the empty declaration when the node's last-sent declaration was non-empty — i.e. skip
only when empty-and-was-already-empty. Requires push to know (or the host to be told) that the
node held something. Simplest: always send to a placed, enrolled node; let an empty declaration
mean "own nothing", which the host already applies correctly when it receives one.
## Workaround used
On ace, one command drops it permanently (the corrected controller never re-composes it):
`sudo ufw delete allow 5671`. At ace's converge it would clear on its own.
@@ -0,0 +1,93 @@
---
status: resolved
opened: 2026-09-27
located-in: [mesh-catalog modules, mesh-control internal/catalogue, mesh-tools src]
fixed-by: mesh-catalog 7b06a7a, mesh-tools fbeb373, mesh-control 05ff606
amended-design: 03-DESIGN/01-to-be/32-what-a-module-declares.md
---
# 127 — A module's event derives a subject nothing publishes
## What was observed
[Design 29](../../03-DESIGN/01-to-be/32-what-a-module-declares.md) §1 says a module names an event
locally and the mesh derives the subject: `emits: order.placed` becomes
`mesh.mod.<module>.event.order.placed`, and a consumer declaring `consumes: shop.order.placed`
subscribes the emitter's own subject. That derivation is built and tested.
**Every event name in the catalogue is still written the way a routing key on the bus the mesh has
is written** — `module.<module>.<verb>` — and the derivation reads it as `<emitter>.<event>`. Asked
of the composer directly, with the module names and declarations the catalogue holds today:
| declared | derived |
|---|---|
| `builder` emits `module.builder.built` | publish `mesh.mod.builder.event.module.builder.built` |
| the catalogue consumes `module.builder.built` | subscribe `mesh.mod.module.event.builder.built` |
| a media module emits `module.<itself>.download.completed` | publish `mesh.mod.<itself>.event.module.<itself>.download.completed` |
| a player consumes `module.*.download.completed` | subscribe `mesh.mod.module.event.*.download.completed` |
The consumer's subject names a module called `module`. **No cross-module subscription in the
catalogue matches what any emitter publishes.** Thirty-seven manifests declare events; every one of
their consume declarations derives this way.
Two further consequences of the same cause, found in the same check:
- One module declares `consumes: "#"` — the wildcard of the bus the mesh has, which is not a
subject at all. The composer **refuses it outright**, so that module's account cannot be composed
and the module cannot be assigned.
- One module emits under a name that is not its own — it declares `module.<other>.image.pushed`
while being a differently named module — which the derivation puts inside *its* namespace. Whether
that is legitimate is a design question: design 32 §2 makes an event's source a fact the server
enforces, and this is a module claiming another's name in its own event.
None of it fails on the bus the mesh runs on today, where a routing key is matched literally and
nothing derives anything. It fails only once the subject is derived — which is to say it fails on
the first mesh raised on the new bus, and not before.
Evidence: run against the controller's own `PermissionsFor` on the current feature branch, with the
declarations read from the catalogue's manifests. Found while wiring the controller's consume side
(design 28 step 3.4), when the controller's own subscription had to be written and the subject it
would have to name turned out not to be the one design 32 specifies.
## Why it matters beyond this instance
**This is the failure [ADR 0074](../../02-DECISIONS/0074-the-wire-is-specified-not-the-types.md)
exists to catch, arriving by a route the conformance suite does not cover.** Two implementations
that disagree about an envelope do not fail to compile — they ignore each other while both keep
running. Here it is not two implementations disagreeing but a *declaration* and a *derivation*
disagreeing, and the symptom is identical: every service starts, every log is quiet, and nothing
reacts to anything.
The fixtures cannot catch it. They pin one emitter's envelope against one subject, and both halves
of that pair are correct. What is wrong is only visible when an emitter's derived subject is set
beside a consumer's derived subject — a check nothing performs, because until the subject was
derived there was nothing to compare.
It also means the rule 3.8 established is weaker than it reads. That task asserted **no manifest
contains a subject**, which holds: a manifest contains a local name. What nothing asserts is that a
local name derives to a subject some emitter actually publishes, and the rule as stated is satisfied
by thirty-seven manifests whose names derive to nothing.
And it blocks work already scheduled. Design 28's step 4.2 (a build source's change reaching the
builder over the bus) and 4.3 (an installation completing over the bus) are both event flows through
exactly these pairs, and the catch-up flow the controller answers is a third — the controller
currently replays a build announcement under its *own* name rather than the builder's, which a
consumer filtering the builder's subject will not hear either.
## Open questions
- Is a local name converted per manifest (`emits: built`), or does the derivation keep accepting the
old form and strip a redundant prefix? The first is thirty-seven manifests and a rule that can be
checked; the second is a rule that cannot, because `module.foo.bar` is also a legitimate three-part
local name.
- What checks the pair? An emitter's derived subject against every consumer's derived subject is a
whole-catalogue check, not a per-manifest one — and a module lives in its own repository and may
be registered long after the catalogue was checked.
- What are `#` and `*` in a consumed name? The bus the mesh has and the bus being built spell
wildcards differently, and a `consumes` pattern is the one place a module writes one.
- May a module emit an event named after another module, and if not, what does the module that does
it today declare instead?
- Who replays? A catch-up answer published by the controller under a builder's subject is the
controller signing an event as another module, which is the thing the derived namespace prevents.
If it must not, then a replay is a different message from an announcement, and the consumer needs
to be told so.
@@ -0,0 +1,109 @@
# Diagnosis — 2026-09-27
## Where it lives
Three places, and only one of them is a bug in code.
**The manifests, in the module catalogue.** Thirty-seven declare events, and every one of them
spells an event the way a routing key on the bus the mesh runs on today is spelled —
`module.<module>.<verb>`. [Design 29](../../03-DESIGN/01-to-be/32-what-a-module-declares.md) §1 says
a module names an event **locally and bare** (`emits: order.placed`) and a consumer names
`<emitter>.<event>` (`consumes: billing.order.placed`). So the manifests are stale against a rule
that was already decided, not wrong against an undecided one. **This is the whole of the reported
symptom.**
**The manifest's own documentation, in the parser.** The comments on `emits` and `consumes` still
describe the old convention and give the old examples — "dotted topic keys, e.g.
`module.umami.site.created`", and `"#"` named as the audit logger's pattern. A module author reading
the file they read most is being told to write the thing that does not work. That is why the drift
was uniform across thirty-seven manifests rather than scattered: nobody was mistaken, everyone
followed the documentation.
**Nothing checks either one.** `ParseManifest` validates the module name, the slug, what it provides
and what it requires. It says nothing about an event name. So a local name that derives to a
namespace belonging to a module called `module` is accepted by every check the mesh has, and the
first thing that notices is a subscription that never fires.
## What was ruled out
**The derivation is not wrong.** Asked directly, with the module names and declarations the
catalogue holds, `PermissionsFor` produces exactly what design 32 §1 specifies for the input it is
given: it reads a consumer's `<emitter>.<event>` and builds the emitter's subject. Given
`module.builder.built` it reads the emitter as `module`, which is a correct reading of an incorrect
declaration.
**The conformance fixtures are not at fault and could not have caught it.** They pin one emitter's
envelope against one subject, and both halves of that pair are correct. What is wrong is only
visible when an emitter's derived subject is set beside a *consumer's* derived subject — a
comparison nothing performs, because until the subject was derived there was nothing to compare.
**Task 3.8's rule is not broken, it is weaker than it reads.** That task asserted **no manifest
contains a subject**, which holds: a manifest contains a local name. Nothing asserts that a local
name derives to a subject some emitter actually publishes.
## What is still a decision and not a conversion
Converting the manifests is implementing design 29, not deciding anything. Three of the report's
open questions are not:
- **Wildcards.** Design 29's table has no wildcard row, and two manifests need one: a module
consuming every download completion across several media modules, and an audit logger consuming
everything. The two buses spell wildcards differently, and a `consumes` pattern is the one place
a module writes one.
- **A module emitting under another module's name.** One manifest declares an event named for a
*provision* rather than for itself. Design 29 §2 makes an event's source a fact the bus enforces,
so this cannot survive as written — and the remedy is probably not a rename but a **seat**, which
is what a name stable across whoever implements it already is.
- **Who replays a build announcement.** The controller answers a catalogue's catch-up by
re-publishing builds under its *own* name, which no consumer of the builder's subject hears, and
for which it holds no grant. Publishing them under the builder's subject would be the controller
signing an event as another module — the exact thing the derived namespace prevents. So the
catch-up is either a different message or a different mechanism, and that is a decision.
## Owners
`located-in` names the manifests and the parser. The replay question reaches the controller and the
catalogue module together and is recorded above rather than in that field, because it is not where
this symptom lives.
# Fixed — 2026-09-27
Converted, and the rule now has checks. What it took was larger than the report said, in two
directions nobody had looked.
**The module code, not just the manifests.** Forty-three files pass an event name to `emit()` at
runtime, and the runtime builds the subject from what it is handed. A converted manifest with
unconverted code would have had the permission and the subject disagree — the same silence, one layer
down.
**Both clients had to learn the mapping.** Each passed the name straight through, which was right on
the bus the mesh runs on today only because modules were writing routing keys. So the old bus's client
now turns a local name into `module.<emitter>.<event>` on the way out and back on the way in.
**Without that, converting the modules would have broken the mesh that is actually running** — which
is the opposite of what fixing this was for.
**The declaration and the handler spoke different vocabularies.** The key a module's handler saw was
the event name alone, while its manifest names `<emitter>.<event>`. So a correct manifest produced a
pattern that could never match. The subject already carries the emitter; the key names it now.
## The three open questions, answered
- **Wildcards**: `*` is one name, `**` is the rest, spelled the mesh's way and derived to each bus's
own. `**` alone is every event, which is what the audit logger wanted and now says in one token.
- **A module emitting under another's name**: not allowed, and the remedy is the seat rather than a
rename — a role's name outlives whoever fills it. **Deferred in practice**: seats carry protocol in
the manifest and in the permission model, and the shared library cannot publish on one, so the
module that did this emits under its own name and its consumers carry that coupling. Worth a task
when a seat's holder needs to emit.
- **Who replays a build announcement**: still open, and narrowed. It cannot become a reply to the
catalogue's inbox: answering a module's inbox needs `_INBOX.>`, which is the blanket grant
[design 25](../../03-DESIGN/01-to-be/25-the-bus-on-nats.md) §4 refuses. So the remaining options are
a subject the controller may publish and the catalogue may subscribe, or a durable consumer that
starts at the beginning of the stream and removes the need to ask at all. Recorded on the work
breakdown as the catch-up half of task 4.5 rather than here, because it is no longer this symptom.
## What it found while running
Two dangling subscriptions that predated this and nothing had reported: a module emitting an event its
manifest never declared, which the new bus refuses outright, and a module waiting for an event nothing
emits — a demo that could never be triggered, because only that module may publish under its own name.
@@ -0,0 +1,73 @@
---
status: located
opened: 2026-09-26
located-in: [mesh-controller internal/catalogue/facts.go, mesh-host internal/apply]
---
# 128 — the machine's hosts file is written whole, and on a workstation it is shared
## What was observed
The private network asks for the `node-names` fact, and the mesh delivers it as
`/etc/hosts`. `nodeNames` composes a **complete** file — its own header, `localhost`, the
machine's own name, and every name in the mesh — and the host writes it over whatever is there.
On an adopted workstation the file the mesh holds contains, besides the predecessor's block of
mesh names:
- the distribution's own lines (`localhost`, the machine's `.localdomain` name);
- two marked blocks (`# BEGIN … # END …`) maintained by a local-development tool, pointing a
dozen development hostnames at `127.0.0.1` — rewritten by that tool whenever its project
list changes;
- hand-added entries of the operator's.
Today the file is only **held** ([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)):
the private network was assigned, not yet taken, so nothing was lost. Taking it — or
converging the node, which takes everything — replaces the file. The development tool's
entries disappear, its projects stop resolving, and every later write it makes is overwritten
at the next change to the mesh's names (a machine joins, a route is contributed), silently and
without a failure anywhere: the development tool thinks it wrote its block, and the mesh thinks
it owns the file.
This is [ADR 0102](../../02-DECISIONS/0102-the-mesh-writes-into-a-shared-file-never-over-it.md)'s
failure exactly — a file the mesh shares with software it did not install, written over — in a
file 0102 did not name, because its merge verb is structured (`into: json`) and a hosts file is
not JSON.
A second, smaller finding from the same reading: the fact's contents depend on which machines
hold the private network. A machine that is enrolled but not yet assigned the private network
is in neither `node-names` nor `node-zones`; its name resolves on the others only for as long
as a predecessor's hosts block survives. Taking the hosts file before every machine is on the
private network loses that name too.
## What would have prevented it
- A **marked-region** merge in the host's vocabulary: `into: "block"` (or similar) — the host
owns only the lines between its own begin and end markers, keeps everything outside them
byte for byte, records what the region held before, and on undeclare removes the region and
nothing else. The shape local tools already use for this very file.
- The `node-names` fact written as that region — no header of its own, no `localhost`, no
machine name — so the distribution's lines and every other tool's stay where they are.
- A converge preview that names a held file the take would replace *whole*, with its line
count before and after, so a person sees "hosts: 31 lines → 12" before the flip.
## The fix, as built (in review)
- **Host:** a file resource may say `"into": "block"`. The host owns only the lines between
`# BEGIN mesh <id>` and `# END mesh <id>` and keeps everything outside them byte for byte. A
new region goes at the `end` by default, or at the `start` (`"at": "start"`) for files where a
line's meaning depends on what stands above it; a region already present is never moved.
Undeclared, what the region held before is put back, or the region is removed and nothing
else. Replacing nothing, it is written on an adopted node without being held — so a machine
gets the mesh's names before its private network is taken.
- **Controller:** `node-names` is a fact written into a shared file, emitted as that region: the
mesh's names only, no header, no `localhost`, no `127.0.1.1` line.
- **Order:** a host older than the block mode refuses the whole declaration on an unknown
`into`, so hosts are upgraded before the controller that emits it.
## Evidence to carry into diagnosis
- `internal/catalogue/facts.go`, `nodeNames`: the complete file is built here.
- The host's file resource supports `into: "json"` only; anything else is a whole write.
- `node show <node>` on the adopted workstation: `holds file /etc/hosts
mesh-wireguard.fact-node-names`, original kept.
@@ -0,0 +1,62 @@
---
status: open
opened: 2026-09-26
located-in: [mesh-controller, mesh-catalog step-ca]
---
# 129 — nothing makes a machine trust the mesh's own certificate authority
## What was observed
On an enrolled, adopted workstation — on the private network, resolving the mesh's names
through the mesh's resolver — every HTTPS name the mesh serves internally fails verification:
```
curl https://<a name the mesh routes internally>/
curl: (60) SSL certificate OpenSSL verify result: unable to get local issuer certificate (20)
```
The route proxy presents a certificate issued by the mesh's internal authority (step-ca, the
`internal-acme-ca` provision). The machine's trust store holds the **predecessor's** authority
and a developer tool's local root, and nothing of the mesh's. No module installs the mesh's
root, and no fact carries it: step-ca's only consumers are proxies, which obtain certificates
over ACME and never need the root on the machine they run on.
[Issue 048](../048-nothing-makes-a-machine-trust-the-mesh-registry/00-report.md) found the same
shape for the mesh's registry and resolved it by treating the private network as the transport
security ([ADR 0082](../../02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)):
the runtime pulls in the clear, over the tunnel. That answer does not carry over. A browser, git
over HTTPS, a package manager and every TLS client a person or a module uses verify the
certificate chain, and there is no "insecure registries" for them — nor should there be.
Consequences today, all silent until someone tries:
- a person on a workstation cannot open any internal HTTPS name without a warning;
- git over HTTPS to the mesh's forge fails, so the working clone URL is ssh-only;
- a module on a non-hub machine that calls another module's internal HTTPS name fails
verification unless its image happens to carry the root;
- the predecessor's authority cannot be retired from any machine while anything there still
speaks TLS to a mesh name, because it is the only authority those machines trust.
## What would have prevented it
- A **mesh fact carrying the internal authority's root** (public material; the controller or
the step-ca module is its source), written onto every machine on the private network — the
same reasoning that has the private network write the registry trust and the names: being on
the network is what makes a machine one that speaks to the mesh's names.
- A resource that puts it where the machine's TLS clients look — on Arch,
`/etc/ca-certificates/trust-source/anchors/` — and **refreshes the extracted bundles**
(`update-ca-trust`). The refresh is the open design question: it is a command, and the link
may not carry an action ([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)). A declared
one-shot unit, or a host primitive for "trust this anchor", are the obvious candidates.
- Removal symmetric to arrival: undeclared, the anchor goes and the bundles are refreshed again,
so a machine leaving the mesh stops trusting it.
## Evidence to carry into diagnosis
- `step-ca` module: provides `acme-ca` / `internal-acme-ca`, listens on 9000 for proxies; no
resource writes its root anywhere but its own state directory.
- The private network's generator writes `/etc/hosts` and the registry trust, and nothing
about certificates.
- On the workstation, the trust anchors present are the predecessor's authority and a local
development root; `trust list` shows no entry for the mesh.
@@ -0,0 +1,66 @@
---
status: located
opened: 2026-09-27
located-in: [mesh-host internal/apply/apply.go (remove)]
amended-design: 02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md
---
# 130 — undeclaring a service stops it, even one the mesh only reloads or only keeps running
## What was observed
Reviewing the uplink modules ([ADR 0125](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md))
found that the host's `remove` path stops every `service` resource that is no longer declared:
`SetServiceState(..., "stopped")`, reported as "stopped; the unit file is not the host's to
delete". `store.Orphans` matches by id alone. So any of these stops the unit:
- the module is unassigned — by mistake, or to switch it for another;
- the node is sent a deliberately-empty declaration ([issue 127](../127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md));
- a later catalogue version renames the resource's `id`.
That is right for a service the mesh brought into being. It is wrong for a unit the mesh
declares only to act on — and the catalogue already has two:
- **The private network declares `docker.service`** (`registry-trust-reload`, state `running`)
so that a change to the registry trust reloads the runtime ([ADR 0102](../../02-DECISIONS/0102-the-mesh-writes-into-a-shared-file-never-over-it.md)).
Unassigning the private network stops the container runtime, and every container on the
machine with it — including ones the mesh does not manage.
- **The sshd module declares `sshd.service`.** Unassigning it stops the machine's ssh daemon:
the lockout the same module's `listens` rule says a firewall must never arrange.
The uplink modules would have added a third and a fourth: unassigning the network manager's
module would have stopped the network manager, taking the machine off the only link the mesh
reaches it by.
## What would have prevented it
- A service resource that says the unit's **lifecycle is the machine's**: declared with no
`state`, the mesh never starts, stops, enables or disables it; it only reloads or restarts a
*running* unit when a trigger changes; undeclared, it is left exactly as it is. (Being built
on mesh-host `feat/a-file-written-into-a-marked-block` for the uplink modules.)
- Then: `registry-trust-reload` declared that way (the runtime is the machine's), and the sshd
module's service too — a machine's ssh daemon outlives any module that configures it.
- A plan or unassign preview that names every unit an undeclare will stop, so the consequence
is read before it happens.
## Resolution
[ADR 0126](../../02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md):
undeclaring removes what the mesh made and gives back what it changed. The host records the state
it first found a unit in, and undeclaring returns the unit to it — a unit found running (the
container runtime, sshd, a network manager) is left running; one the mesh started (the packet
filter a converge loaded) is stopped again; nothing is started on the way out; a record from
before the host kept what it found leaves the unit alone. That covers the runtime, sshd and the
uplink modules at once, without each module opting out; the private network and the sshd module
need no change.
A first draft — never stop a unit the mesh did not create — was rejected while implementing it:
returning a converged node to adopted unloads the mesh's filter by exactly this path.
Found on the way: an undeclared `process` failed every apply on its node (`remove` had no case
for it). Now removed with its unit, timer and bundle — the mesh's own code. `user` and `archive`
have the same gap and are left for their own decisions: removing a login or unpacked files is not
something to settle in passing.
The unassign preview is partly answered — the host's plan names each unit it will stop — and the
controller's side is left open.
@@ -0,0 +1,107 @@
---
status: resolved
opened: 2026-09-27
located-in: [mesh-catalog modules/gitea, mesh-controller cmd/mesh-controller, mesh-controller internal/builder]
fixed-by: mesh-catalog #124 — the forge module watches for merged pull requests and emits `pull.merged` with the merge commit and the clone address; mesh-controller #110/#111 — the control plane follows that event on the bus, marks every module built from that repository and branch as moved, and builds them bases first, stopping when a base fails; mesh-controller #113/#114 — a build records the bases it was handed and the graph is read from builds, without which "bases first" had no edges to order by.
amended-design: 03-DESIGN/01-to-be/28-building-the-bus.md
---
# 131 — Nothing tells the mesh a source moved, and it reports itself current anyway
## What was observed
Six changes were merged to the trunk of six repositories in one sitting. The build machine built
nothing. Its last build, minutes before the first merge, was still the one it reported; no build was
requested, refused or failed, because none was ever asked for.
Asked afterwards what was wrong, the mesh said:
> 4 machine(s), all doing what they were told, all heard from, running what the mesh would send them,
> and every module current with its source
Every one of the six had moved. The last clause was false for all of them, and it is the clause a
person reads to decide whether there is anything to do.
## Why it matters beyond this instance
**The mesh learns a source moved by being told, and there is no longer anything to tell it.** The
command exists — a person names the module and the commit — and so does the question the overview
answers. What is missing is whatever used to connect the two. One repository still carries a forge
webhook aimed at a port; the rest carry none, and the port belongs to a different service than the
one the arrangement implies. So the state is not "the trigger is broken" but "there is no trigger,
and nothing says so".
**A wrong answer is worse here than no answer.** "Every module current with its source" is
indistinguishable, to a reader, from a mesh that has genuinely caught up. The overview is built to be
the thing you check instead of checking by hand, so a confident false negative removes the habit that
would otherwise have caught it. Nothing in the mesh is at fault for being out of date — it is at
fault for saying it is not.
**It is also why "current with its source" cannot be a stored fact.** The mesh compares what it built
against what it was last told the source was, and calls that agreement. Two facts agreeing tells you
nothing when both come from the same place.
## The intended shape, which is decided and not built
The forge emits what happened to it — a pull request merged — and the build machine reacts by
building what that commit affects. That keeps the forge ignorant of the build system and the build
machine ignorant of the forge's internals, which is the same argument
[ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md) makes for addressing an event
to its emitter: the merge is a fact about the forge, and what should be rebuilt because of it is not
the forge's business to know.
The forge's module already declares the event. The build machine declares that it consumes nothing.
## What the trigger cannot be
**Not one build per changed module.** The modules form a graph: several are built from one
repository, and some are the base another is compiled on — a runtime image, a compiler base, a
repository whose context a second module builds from. Firing a build for each changed module
independently would start work that cannot succeed yet and produce a failure per dependent, for one
cause.
Observed while catching the mesh up by hand on 2026-09-27: a compiler base had to move before
anything compiled against it could build, and when it failed, the right behaviour was for its
dependents to wait rather than each fail the same way. Fifteen modules shared the cause. A trigger
that reports it fifteen times has buried it.
So whatever reacts to the forge's event resolves what changed into an order, builds the bases first,
and holds a dependent while its base is unbuilt or failed. That is a larger thing than "rebuild what
the commit touched", and knowing it now is cheaper than discovering it from fifteen identical
failures.
## Open questions
- Is "the source moved" still a thing a person can assert by hand once the event path exists, or does
the hand-operated form become the thing that made this failure possible?
- Which commit does the build machine act on — the merge, or each commit it brought — and what does
it do when several arrive for one module at once?
- How does the overview stop being able to lie? Comparing what was built against what was recorded
will always agree. Whether the trunk has moved is a question only the forge can answer, so either
the overview asks it, or it stops claiming to know.
- Does this want to be the same mechanism as the build request on the bus
([ADR 0129](../../02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md)), or does it sit in
front of it?
## What was done (2026-09-28)
The shape above was built as described: the forge's module emits the merge, the control plane
consumes it, and nothing on either side knows the other's internals. The build is asked for the
merge commit, not each commit the merge brought — the trunk moved once, to one place. Several merges
for one module arriving in a row are followed in turn, each moving the recorded source to its own
commit, so the last one to arrive is the one the mesh ends up built from.
The hand-operated form stays. `module moved` is how a source is recorded without a forge — a module
built from a repository elsewhere, or a mesh whose forge module is down — and it is the same act the
event performs, so the two cannot disagree about what "moved" means.
**Bases first needed edges, and there were none.** The order this report asked for was written and
walked a graph that no build had ever recorded: a recipe reads its base from a build argument, so the
digest was never in the file the builder read edges from. A build now reports what it was handed, the
control plane records it by artifact path, and the order is read from each module's newest build.
**What still can lie.** The overview compares what was built against where it was last told the
source is; the forge's event is now what moves that mark, so it is right for as long as the forge
module was listening. A merge made while that module was down is a merge the mesh does not know of
until the module next polls — it announces what merged since it last looked, so the gap closes when
it comes back, and not before. The overview does not say so.
@@ -0,0 +1,60 @@
---
status: resolved
opened: 2026-09-28
located-in: [mesh-controller cmd/mesh-controller]
fixed-by: mesh-controller — `module add` takes `--path` and `--self`, so a module handed over by hand records the whole location it came from; a record naming a repository and no directory says so in the reply; the rule is one function with a test beside it. The nine records already wrong were corrected by rebuilding each with its real directory, which is the same act through the same door.
amended-design:
---
# 132 — A module can be recorded without the directory it lives in
## What was observed
Nine modules on one mesh could not be rebuilt. Each attempt failed the same way:
> has no module.json at its root, so there is nothing saying what it is
All nine were recorded as coming from a repository that holds many modules, each in its own
directory — and each record named the repository and no directory. So every build cloned the
repository and looked for a manifest where there has never been one.
The failure only surfaced when something asked for all of them at once. Before that, the overview
said every module was current with its source, because what it compares is what was built against
what the mesh was last told the source has, and neither half knows whether the source can be found
at all.
## Why it matters beyond this instance
**A module is a repository and a directory inside it** ([ADR 0069](../../02-DECISIONS/0069-a-module-is-a-repository-and-a-path.md)),
and one of the two doors into the catalogue could record only the first half. A build records the
directory it was given, so a module that arrived by being built is always whole; a module handed over
by hand had no way to say where it lived, and the flag to say it did not exist. The rule was decided
and enforced on one path out of two.
**Half a location reads exactly like a whole one.** Nothing in the record is empty in a way a person
would notice: the repository is there, the branch is there, the commit is there. The mesh only finds
out at the moment it needs the manifest, which is the moment it is trying to rebuild — and the module
stays on whatever it last built, indefinitely, with nothing saying why.
**It is the same shape as [131](../131-nothing-tells-the-mesh-a-source-moved/00-report.md).** A
comparison between two facts the mesh holds about itself will agree with itself. Whether the source
can be found is a question only an attempt to read it answers, and the answer had nowhere to go.
## What was done
`module add` takes the directory and which forge holds the repository, so a hand-registered module
records the same whole location a built one does. What a record must say to be worth anything is one
function with a test beside it, rather than a paragraph in a help string: provenance together or not
at all, a directory needs a repository to be inside, a path on the mesh's own forge is not an address.
And a record that names a repository but no directory says so when it is made — not refused, because a
module really at a repository's root is ordinary, but said, because the person adding it is the one
who knows which it is.
The nine wrong records were corrected by building each with its real directory, which re-records it.
No row was written by hand.
## What is still true
A directory that does not exist in the repository cannot be refused when the module is added: the
control plane does not clone, and inventing a check there would mean it did. The first build says so
plainly, which is one build rather than nine, and the record it leaves behind is right from then on.
+6 -2
View File
@@ -52,8 +52,12 @@ vocabulary — *controller* (not "control plane"), *foundation* (not "substrate"
docs (`layer`, `status`, `code`, `updated`), issue reports (`status`, `located-in`,
`fixed-by`, `amended-design`) and decision records (`status`, `date`, `deciders`).
Never create a central status file; cross-cutting views are generated from frontmatter.
- **`02-DECISIONS/` records are immutable.** Supersede with a new record; never edit meaning. Fixing a
broken link or path is allowed.
- **`02-DECISIONS/` records hold their meaning.** Supersede with a new record rather than rewriting
what was decided, the options weighed, or a consequence another record relies on. Fixing a broken
link or path is allowed, and so is a **progressive insight** — a correction of *fact* that leaves
the decision standing, made in place, marked and dated in the record's own words
([`02-DECISIONS/README.md`](02-DECISIONS/README.md)). A fact going stale is not the decision going
wrong, and superseding a sound record for one buries it.
- **Design docs are prose and diagrams only** — no code. A manifest field may be named; a
manifest may not be pasted.
- **Two layers, never mixed.** [`03-DESIGN/00-as-is/`](03-DESIGN/00-as-is/) describes the mesh