Compare commits

..
Author SHA1 Message Date
jochen 5440a144a0 ADR 0201 (module state) renumbered 0202: 0201 landed first on main for a provider's derivations 2026-10-04 03:44:34 +02:00
mesh-admin be4b5777b8 Merge pull request 'Research 024 and ADR 0201: a module keeps its current state in key-value buckets' (#348) from feat/module-state-on-the-bus into main 2026-10-04 01:43:34 +00:00
mesh-admin d882b3568c Merge pull request 'Group 8: ADR 0201 (a provider declares what it derives, issue 124) and ADR 0189 (the store keeps what the records name, issue 108); issue 202' (#305) from feat/the-store-keeps-what-the-records-name into main 2026-10-04 01:30:47 +00:00
jschoubben a3523617d3 Review before merge: the multi-holder boundary, the window's open race, the sweep's bounds
ADR 0201 gains the boundary found reading it back: a consumer keeping several
holders of a deriving provider is refused, because the two ends have no way
to agree. ADR 0189 gains two consequences — the sweep is bounded because it
runs inside a build, and an apply arriving mid-window reopens it.

That last one is issue 224, recorded rather than fixed: the host's rule for a
stopped container is to replace it, and while-stopped is the first thing that
makes a stopped container intentional. Both candidate fixes are decisions with
their own cost. Nothing is worse than it was; the store has never collected.
2026-10-04 03:27:32 +02:00
jochen 6b4da63261 Research 024: the composed grants, checked against a server once built 2026-10-04 02:50:02 +02:00
jschoubben 0231974226 Rebased onto main: ADR 0188 renumbered to 0201, and issue 202's evidence re-taken
The bundles refactor took 0188 on main while this waited in a pull request,
and the mesh's own code cites that one, so this record moves. Only the number
moved; the decision is the one taken on 2026-10-02, and the record says so.

Issue 202 re-checked against the refactored main: the fault stands, and the
test that surfaced it now fails one step earlier on issue 203's new credential
guard. Proven again past both — mint the credential, compose twice, and all
eight of dnsmasq's resources appear only with the setting set. ADR 0164 is
noted as the decision that answers half of it, and is not built.
2026-10-04 02:44:58 +02:00
jschoubben 92c029d10e ADR 0189: the store keeps what the records name, and a maintenance step holds its writers still
Issue 108: the artifact store has never collected anything. Fifty-three
repositories on the machine that serves everything else, and the only outcome
of leaving it is a full disk reported as somebody else's failure.

The mesh decides what may go — from its own build records, so it never names
a digest it did not put there — and the store reclaims the bytes in a nightly
window with its server held still. Deletion on the one door takes nothing a
push did not already have.

Designs 18 and 20 amended; issue 108 resolved.

Also issue 202, found running the controller's suite: a module whose required
setting nobody set is left out of the machine in silence, and dnsmasq became
that module this morning.
2026-10-04 02:40:25 +02:00
jschoubben 2a60da821d ADR 0188: a provider declares what it derives for each consumer, and the mesh tells both ends
Issue 124: a value the mesh's own rule produced reached neither end as a
statement. The object store's provisioner derived each consumer's bucket in
its own code; all three consumers transcribed the rule into their own
definitions, one of them wrong, and each of the three also named the machine
it happens to run on.

A served value may now name the consumer the mesh is serving. Design 27
amended; issue 124 resolved.
2026-10-04 02:40:06 +02:00
jochen e1b0bbde91 Research 024 and ADR 0201: a module keeps its current state in key-value buckets
Events miss a machine that joins after them and replay history where only the
latest matters. A module now declares state it owns and reads; the controller
creates the buckets, the runtime serves them on the bundle's channel. Designs 32
and 25 amended; grants measured against a running server.
2026-10-04 02:36:59 +02:00
mesh-admin 8ca09c70d6 Merge pull request 'Issues 213 and 223 resolved' (#347) from issues/213-223-resolved into main 2026-10-04 00:19:22 +00:00
jochen db71b83711 Issues 213 and 223 resolved: the controller runs as a process, and genesis hands over to it 2026-10-04 02:19:09 +02:00
mesh-admin 9143d0b7c1 Merge pull request 'ADR 0200: genesis pivots to the controller as a container, and the first push hands it to a process' (#346) from decision/0200-genesis-pivots-to-a-container-and-hands-over into main 2026-10-03 23:41:36 +00:00
jochen 24a51a8e53 ADR 0200: genesis pivots to the controller as a container, and the first push hands it to a process 2026-10-04 01:41:30 +02:00
mesh-admin f0d7f91d90 Merge pull request 'Design 38: WP4c complete' (#345) from design/38-wp4c-complete into main 2026-10-03 23:35:31 +00:00
jochen 8578a06ca8 Design 38: WP4c complete, no module's own code runs in a container 2026-10-04 01:35:16 +02:00
mesh-admin 86083f9c9d Merge pull request 'Issues 215, 221, 222 resolved' (#344) from issues/215-221-222-resolved into main 2026-10-03 23:29:38 +00:00
jochen 63d328147e Issues 215, 221 resolved with live proof; 222 diagnosed and resolved 2026-10-04 01:29:32 +02:00
mesh-admin fead0ea440 Merge pull request 'Issues 219-223 and design 38 WP4c built' (#343) from issues/219-223-and-wp4c into main 2026-10-03 23:14:49 +00:00
jochen df503d1cff Issues 219, 220 resolved, 221 located, 222 and 223 opened; WP4c built and proven 2026-10-04 01:14:44 +02:00
mesh-admin c2fc829822 Merge pull request 'Issues 211, 212, 214, 216, 217 resolved; 215's fix recorded' (#342) from issues/211-217-resolved into main 2026-10-03 22:25:07 +00:00
jochen 8d8e5c9a7e Issues 211, 212, 214, 216, 217 resolved with their proofs; 215's fix recorded 2026-10-04 00:24:53 +02:00
mesh-admin affba60b79 Merge pull request 'Issues 219-221: the builder's ordering gaps' (#341) from issues/219-221-the-builder into main 2026-10-03 22:16:13 +00:00
jochen 9ddbc4c68e Issues 219-221: the builder's ordering gaps seen while rolling out issue 218 2026-10-04 00:16:08 +02:00
mesh-admin a972db91f0 Merge pull request 'Issue 218 resolved: the runtime follows its membership, and a refused subscription is not fatal' (#340) from issues/218-rollout into main 2026-10-03 22:14:10 +00:00
jochen 029698fdc8 Issue 218 resolved: the runtime follows its membership, and a refused subscription is not fatal 2026-10-04 00:14:04 +02:00
mesh-admin 895c2afad1 Merge pull request 'Issue 218: a mesh seat answered by a non-holder (located, fixed in mesh-controller#248)' (#339) from issues/218-a-mesh-seat-answered-by-a-non-holder into main 2026-10-03 21:32:24 +00:00
jochen 4b8c5e3b11 Issue 218 located: the controller issued a mesh seat to every claimant; fixed in mesh-controller#248 2026-10-03 23:31:13 +02:00
mesh-admin d362155401 Merge pull request 'Issue 218: a mesh seat is answered by a module on a machine that does not hold it' (#338) from issues/218-a-mesh-seat-answered-by-a-non-holder into main 2026-10-03 21:25:27 +00:00
jochen 557760e537 Issue 218: a mesh seat is answered by a module on a machine that does not hold it 2026-10-03 23:25:17 +02:00
mesh-admin 6b1ebd1d4a Merge pull request 'Issue 217: a refused announcement took down every container's runtime, and the console with it' (#336) from issues/217-a-refused-announcement into main 2026-10-03 21:20:06 +00:00
mesh-admin fa73a17ceb Merge pull request 'Issues 211, 214, 215, 216 diagnosed' (#335) from issues/211-214-216-diagnosed into main 2026-10-03 21:19:03 +00:00
jochen 08cea893bd Issue 217: a refused announcement took down every container's runtime, and the console with it 2026-10-03 22:47:28 +02:00
jochen 04625d3e35 Issues 211, 214, 215, 216 diagnosed: root causes and the branches that fix them 2026-10-03 22:28:40 +02:00
37 changed files with 1527 additions and 209 deletions
@@ -1,37 +0,0 @@
---
status: active
initiated: 2026-10-03
touches: [the seats, the seat protocol, the controller's ownership check, 03-DESIGN/01-to-be/26-the-seats.md]
---
# 023 — A seat protocol that defines what its holder owns
## What is investigated
**A seat is a definition — a protocol — and a module occupies it by implementing that protocol.**
Today the protocol is what the holder accepts, emits and serves (ADR 0118, 0129, 0132): its verbs, as MCP
tool definitions. This asks whether the protocol should also name the **files and directories the
holder owns**, so that occupying the seat means owning them: `node-resolver-config` owns
`/etc/resolv.conf`, `node-hosts-file` owns `/etc/hosts`, the intrusion prevention owns its jail file.
The direction is the protocol's, not the holder's: the seat states what any holder must own; a module
that wants the seat must declare those paths among its resources, or the controller refuses the claim
as not implementing the seat. Two seats may not name one path.
## Why
Who owns a singular file is today answered by reading every manifest, and enforced only after the fact,
when two modules on one machine both declare the same path. The question *which module owns
`/etc/resolv.conf`?* came up on 2026-10-03 with no place to look it up. A seat that names the path answers
it from the seat table, before any module is written, and makes "implements the seat" checkable.
## What it touches
- The seat definition and its table (ADR 0122) — a new part of the protocol.
- The controller's ownership check (`checkResources`), which already refuses two modules owning one path.
- Every node seat that is really about a file: `node-resolver-config`, `node-hosts-file`
([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)),
`node-intrusion-prevention`, `node-packet-filter`.
Raised by the operator during the resolver work of ADRs 0194–0199 and parked there so that work was not
widened by it.
@@ -0,0 +1,152 @@
---
status: graduated
initiated: 2026-10-04
touches: [the bus, what a module declares, the tool runtime, the SDK, the bus grants, 03-DESIGN/01-to-be/25-the-bus-on-nats.md, 03-DESIGN/01-to-be/32-what-a-module-declares.md]
became: [02-DECISIONS/0202-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md, 03-DESIGN/01-to-be/32-what-a-module-declares.md, 03-DESIGN/01-to-be/25-the-bus-on-nats.md]
---
# 024 — State a module keeps on the bus
## What is investigated
A place on the bus where a module's own code keeps **current state** — not history — that every
machine sees, including a machine that joins after the state was written: put, get, delete, list and
watch, reached through the node's runtime the way a bundle already publishes, asks and subscribes
([ADR 0198](../../02-DECISIONS/0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)).
On NATS that is a key-value bucket. The questions are what a module declares, who creates the
bucket, what the grants are, what the runtime's verbs are, and what may never be stored.
## Why
The mesh carries two kinds of module traffic and a third is missing.
- **Events** land in the EVENTS stream: limits retention, seven days, ten thousand messages per
subject, a durable consumer per consuming module that replays what it missed. Never a secret
([design 32](../../03-DESIGN/01-to-be/32-what-a-module-declares.md) §10).
- **Requests** are core request/reply — tool calls, a bundle's `mesh/ask` — and are kept nowhere.
Neither is *the current value of something*. Two cases from the first module that needs it, the
operator's agent on a machine ([design 36](../../03-DESIGN/01-to-be/36-the-operators-agent-on-a-machine.md)):
1. **An MCP server registered for every machine.** Registering emits an event every machine's copy
of the module consumes. A machine the module is assigned to *after* the registration has no
durable consumer yet — the consumer is created at assignment — so it never hears of it. Wanted
instead: one entry per server, for every machine or for one; every machine reads the whole current
set when it starts and watches for changes; unregistering is a delete; any machine can list it.
2. **Which licence a machine is bound to** ([design 39](../../03-DESIGN/01-to-be/39-the-anthropic-licence-manager.md)).
As events, a machine that was off for a day replays every rotation since and asks for a token
after each. It needs only the latest binding and its generation. The token itself stays on
request/reply and is never stored.
The design already expects this. [Design 25](../../03-DESIGN/01-to-be/25-the-bus-on-nats.md) §1:
"conditions and observed state in key-value buckets that anything may watch".
[Research 017](../017-a-mesh-that-heals-itself/01-the-intended-behaviour.md) wants a provisioner's
"what I applied" and a rotation's step kept in one rather than in memory. Nothing implements it.
## What exists, measured 2026-10-04
| | fact | where |
|---|---|---|
| streams | five kinds of mesh stream: CONTROL (work queue), NODES and ASSIGNMENTS (last per subject), EVENTS (limits: 7 days, 10 000 per subject), one work queue per seat that accepts | the controller's broker streams |
| the state relationship | [design 32](../../03-DESIGN/01-to-be/32-what-a-module-declares.md) §4 already names *state* — 1:1, last per subject — and says it is "declared: the mesh's own". Two streams use it, both written by the controller. No module can declare it | design 32, the controller |
| key-value buckets | none, anywhere | all four code repositories |
| the runtime's bus verbs | `mesh/publish`, `mesh/ask`, `mesh/subscribe`; delivery back to the bundle is `mesh/event` | the runtime's launcher |
| the runtime's principal | one bus user per machine carries every assigned module; its grant is the union of theirs. That one module's code does not act as another is the runtime's to keep: it publishes under the module's own name by construction | the controller's grant composition, the runtime's bus |
| what a bundle is issued | a membership per assignment, last per subject, read directly by the runtime: where it serves, where it emits, what it reaches | ADR 0160 |
| who creates bus objects | the controller only — mesh streams on every raise, a seat's stream at registration, a module's consumer at assignment. No module reaches the JetStream API | design 25 §3 |
### What a key-value bucket needs from a grant, against a real server
Measured against nats-server 2.10 with the Go client the runtime already uses, a bucket created by
an unrestricted user and used by two users holding only the subjects below (`B` is the bucket):
| operation | subject published | writer | reader |
|---|---|---|---|
| bind to the bucket | `$JS.API.STREAM.INFO.KV_B` | yes | yes |
| get | `$JS.API.DIRECT.GET.KV_B.>` | yes | yes |
| put, delete | `$KV.B.>` | yes | **refused** |
| list keys, watch | `$JS.API.CONSUMER.CREATE.KV_B.>` — an ordered, ephemeral consumer | yes | yes |
| stop a watch cleanly | `$JS.API.CONSUMER.DELETE.KV_B.>` | yes | yes |
| answers | its own inbox, which every principal already subscribes | — | — |
*Checked again once built, 2026-10-04:* the grants the controller composes for two machines' runtimes —
one carrying the owner, one only a reader — were loaded into a server as composed, and each operation
was run as each runtime's user. The owner's did all of them; the reader's read, listed and watched,
and its put and delete were refused by the server.
Three things the measurement showed that reading the documentation would not have:
1. **A refused put is not an error to the caller; it is a timeout.** The server reports the
permission violation asynchronously, on the connection, and the client waits out its deadline
for an acknowledgement that never comes. So a runtime that relies on the grant alone tells a
bundle "timed out" for "you may not write this" — it must refuse first, from what the module was
issued, with the reason.
2. **A watch's current values include deletions.** A key deleted earlier arrives among the initial
values as a delete marker, before the end-of-current marker. A bundle asking "what is there now"
must not be handed those.
3. **Without the consumer-delete grant, stopping a watch hangs** until its deadline, and the
ephemeral consumer lingers on the server until it times out by itself.
### Whether the events shape is enough instead
Honestly compared, because a new primitive is a cost:
- **EVENTS cannot be made last-per-subject for some subjects.** Retention is per stream, and
JetStream refuses a second stream overlapping the first (verified and recorded in design 32 §3).
A state subject inside `mesh.mod.*.event.>` keeps EVENTS' seven days: a licence binding unchanged
for a week disappears.
- **A separate last-per-subject stream per module** is possible — it is exactly what a key-value
bucket *is* on the server: a stream with one message per subject, a rollup for purge, and direct
reads. Building it by hand gives up the client's get, list, delete and watch, which are the
operations both cases need, and would be the mesh writing NATS's own key-value layer again.
- **Consumers are the wrong reader.** A durable consumer per reading module is created at
assignment and replays from where it is; state wants "everything current, now, then changes",
which an ordered ephemeral consumer from the last value per subject gives and a durable does not.
So key-value is not a convenience over events; it is the state relationship design 32 already
names, opened to modules.
## Questions, and what this effort proposes
1. **What a manifest says.** `state` names the buckets a module owns, by local name — every
instance of the module may write them and read them. `reads` names another module's bucket as
`<module>.<name>`, read-only. Names only, never a bucket or subject (design 32 §1). A bucket's
options — how many past values it keeps, how long a value lives — are the owner's to declare,
the way a seat declares its own retention (design 32 §3).
2. **Scope.** One bucket per module per name, mesh-wide. A key may carry a machine by the module's
own convention (`all.<server>`, `<machine>.<server>`). A bucket per machine was considered and
not proposed: "list every server for every machine" becomes a walk over buckets, and the grant
could only narrow writes, which nothing asked for — every instance of the owner already writes.
3. **Who creates the bucket.** The controller, from the catalogue, on every raise — a bucket exists
from registration, like a seat's stream, so a reader can watch before the owner is assigned
anywhere. Never a module.
4. **The runtime's verbs.** `mesh/state.get`, `mesh/state.put`, `mesh/state.delete`,
`mesh/state.keys`, `mesh/state.watch`, each naming the bucket as the module named it. A watch
is answered once the current values are on their way, then each change is delivered to the
bundle as a `mesh/state` request it answers — current values first (no deletions among them), an
end-of-current marker, then changes. A child that restarts watches again, as it subscribes
again. The runtime refuses, with the reason, a bucket the module was not issued, and a write to
one it only reads.
5. **Secrets.** None in a bucket, sealed or not: a bucket is a stream (design 32 §10). Sealed values
are plain base64 and cannot be recognised, so the mechanical check is partial and said to be: the
runtime refuses a value carrying a field whose name says it is a credential (`password`,
`secret`, `token`, `authorization`, …), which catches the ordinary mistake and not a determined
one. For the first consumer this has a concrete consequence: an MCP server registered with an
authorisation header keeps that header out of the bucket.
6. **History, lifetime, size.** One value per key unless the owner says more; no expiry unless it
says one; a value at most 256 KiB and a bucket at most 64 MiB, the mesh's caps rather than a
module's. **A bucket outlives its module's assignment** — what a module stored is data, and data
outlives what declared it ([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md));
unassigning is not cleaning up. A bucket whose declaration is gone is reported, never removed.
7. **Events or state.** State (above).
## The work, once decided
1. A decision record, then design 32 (*state* becomes a relationship a module declares) and design
25 (key-value buckets are part of the bus) amended.
2. The controller: the manifest's two words and their registration check; buckets asserted on every
raise; the grants for owners' and readers' runtimes; the buckets issued in each membership.
3. The runtime: the five verbs, the watch delivery, the refusals; tested against a real server.
4. The SDK, TypeScript and Go: a small state surface over the verbs.
5. Proved on a running mesh with one small module, then handed to the operator's agent, whose
registered servers move from events to a bucket.
@@ -0,0 +1,141 @@
---
topic: the mesh
status: accepted
date: 2026-10-02
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md
---
# 189. The store keeps what the records name, and a maintenance step holds its writers still
## Context
The mesh's artifact store has never collected anything
([issue 108](../04-ISSUES/108-the-registry-has-no-garbage-collection-once-it-has-two-doors/00-report.md)).
Every build pushes another layer set; nothing has ever removed one. The predecessor ran a routine
on a timer — stop the registry, collect, start it — and the conversion carried the settings that
routine depends on without the routine, because the routine was a script beside the module and not
a resource in it. The store now holds fifty-three repositories on the machine that serves
everything else, and the only outcome of leaving it is a full disk reported as somebody else's
failure.
Three things stood in the way, and the issue names all three.
**Nothing in the mesh's vocabulary expresses a maintenance window.** The collector requires every
writer stopped while it runs. A `run-once` step runs *beside* containers, not instead of them, and
a scheduled step is the same container on a cadence. There is no way for a module to say *hold this
container of mine still while this runs*.
**Deletion is not enabled, and the door it would be enabled on has no accounts.** The store is
internal, reached by name over the overlay, trusted because being on that network is the permission
([ADR 0082](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)). The predecessor
kept deletion behind an authenticated door, which it could, having one.
**Nothing says what may be removed.** The registry's own answer — collect everything no tag names —
is wrong here. The mesh pushes each artifact under one moving tag and pins machines by digest, so
every build but the newest is untagged and some machine may still be running it.
## Decision
**1. Deletion is enabled on the store's one door, and the overlay stays the permission.** The
objection dissolves on inspection: that door **already accepts a push**, and a writer who can push
can replace any tag in the store with anything it likes. Delete takes nothing a push did not
already have, and the machines that can reach the door are the ones the mesh's own filter admits
([ADR 0168](0168-a-converged-machine-is-filtered-by-the-mesh-alone.md)). Putting an authenticated
door in front of deletion while leaving push open would be a lock on the window beside an open
door, and it would cost the thing ADR 0082 bought: a store every machine can reach without a
credential to distribute first.
**2. The mesh deletes what it made and no longer keeps; the store reclaims the bytes.** Two halves,
each doing what only it can.
The **mesh** decides. It does not need to enumerate the store to do it — it has never put anything
there it did not record, so **every digest it could remove is already in its own build records**.
It deletes those manifests through the store's door, by digest, and remembers that it did.
The **store** reclaims. A deleted manifest frees no bytes until the registry's own collector walks
the storage with nothing writing to it, so the module declares that collector as a scheduled step
with the server held still for its duration. Plain collection, not `--delete-untagged`: what the
mesh keeps is still a manifest in the store, so it is still referenced, so its blobs stay — the
dangerous flag is not needed at all once the mesh is the one deciding.
**3. What the mesh keeps, stated as three reasons rather than a number.** A digest is kept because:
- **a definition names it** — every artifact reference in any module's current recorded manifest,
which is what the mesh would hand a machine now. No age limit: this is the floor;
- **the mesh can still go back to it** — every artifact of the **five most recent successful
builds** of each module, so a release that turns out wrong has somewhere to return to;
- **nothing else.** An artifact older than that, which no definition names, is what the store is
carrying for no stated reason.
A digest the mesh did not record making is never touched. That is not a safety margin, it is the
whole rule restated: the mesh removes what it put there and can account for, and the images genesis
pushed before any record existed are exactly what this must not reach
([04-ISSUES/102](../04-ISSUES/102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md), F4).
**4. A scheduled step may hold its module's own containers still while it runs** —
`while-stopped`, naming resource ids in the same module. The host stops each, runs the step, and
starts them again **whatever the step did**, including when it failed or the host was interrupted.
Three boundaries:
- **Its own module's containers only.** A module that could quiesce a neighbour could stop the
mesh; a maintenance window is a statement about one service's own insides.
- **Scheduled steps only, not `run-once`.** At apply time the host already has a window: the
declaration is applied in order and a step gates what follows, so a one-time offline migration
says *before* rather than *instead of*. A recurring window is the case order cannot express.
- **Restoring is not conditional.** A step that fails must leave the service running; the whole
risk of this field is a window that never closes.
**5. The sweep runs where the records change — after a build the mesh recorded.** That is the
moment new bytes landed and the moment the keep set moved, and it needs no new timer. The
store's collection runs nightly, because reclaiming is slow and the thing it reclaims is already
unreferenced.
## Consequences
- Disk stops growing without bound on the machine that serves the mesh. That is the whole point
and it has no other way to be true.
- A machine behind by more than five builds of a module, which recreates a container, cannot pull
what it was running. It is already a machine the mesh reports as behind, and the answer is the
one the mesh already gives it: the current declaration. Stated here rather than discovered.
- The store is a little less of a museum. A digest in an old build record may no longer be
fetchable, and the record still says what that build made — the record is history, not an
index of what is on disk. The collected mark is kept beside it so the two can be told apart.
- `while-stopped` is a second thing the host does to a container it did not start this pass. It is
deliberately the narrowest form: the module's own, by id, restored unconditionally.
- The store is briefly unavailable each night, for as long as collection takes. Everything that
pulls from it retries; nothing in the mesh treats a momentary store as a failure
([ADR 0185](0185-a-control-plane-behind-its-seats-row-serves-what-it-can.md)).
- **An apply arriving during the window reopens it**, because the host's rule for a container it
finds stopped is to replace it, and `while-stopped` is the first thing that makes a stopped
container intentional. Found by reading this before it merged, recorded as
[issue 224](../04-ISSUES/224-an-apply-reopens-a-maintenance-window-by-recreating-what-it-held-still/00-report.md)
rather than fixed here: the two candidate fixes — the window takes the apply lock, or the apply
learns which containers are held — are each a decision with its own cost, and neither belongs
inside this record. Nothing is worse than it was; the store has never collected at all.
- **The sweep is bounded**: at most two hundred artifacts and sixty seconds per build, stopping at
the first refusal, because it runs inside somebody's build. What is left over is offered again
next time. The store stops growing from the first sweep; it does not empty in one.
## How this is checked
- The host: a scheduled step with `while-stopped` stops the named containers before the run and
starts them after; it starts them again **when the step fails**; it refuses an id that is not a
container of the same module, its own id, and `while-stopped` on a `run-once` step. Each refusal
is tested for what it says, not only that it says something.
- The controller: given build records and current manifests, the keep set holds every reference a
manifest names and every reference of the five most recent builds per module, and nothing else;
a reference the mesh never recorded is never in the delete set; a delete that answers 404 is
recorded as collected rather than retried forever.
- The sweep is tested against a fake store that records what it was asked to delete, so what is
asserted is the decision and not the registry's behaviour.
- Live: the store's size before and after the first nightly collection, read from the machine.
## References
- [issue 108 — the registry has no garbage collection](../04-ISSUES/108-the-registry-has-no-garbage-collection-once-it-has-two-doors/00-report.md)
- [ADR 0082 — the registry is reached by name and trusted by the overlay](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)
- [ADR 0053 — a step that runs on a schedule](0053-a-step-that-runs-on-a-schedule.md)
- [ADR 0156 — an artifact is what a build produces, and the store is named for its scope](0156-an-artifact-is-what-a-build-produces-and-the-store-is-named-for-its-scope.md)
- [design 32 — what a module declares](../03-DESIGN/01-to-be/32-what-a-module-declares.md)
@@ -1,127 +0,0 @@
---
topic: the tiers
status: accepted
date: 2026-10-03
deciders: jochen
reconstructed: false
extends: 0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
---
# 199. A module that answers names declares its zone, and a node's hosts file is one module's
## Context
**[ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) and
[ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md) leave
one resolver holding the nodes' internal domains, and retire the resolver every node ran.** Two kinds
of names lived in those per-node resolvers that are neither a node nor a route, and both were found on
the workstation on 2026-10-03:
- **Names a module answers.** The lab raises scenario machines and gives them addresses from its
scenario files — the anchor's stand-in at a documentation address, the home server's on the LAN —
and the workstation resolved `<machine>.incus` through two wildcard lines in a drop-in file its
resolver read. The lines were written by hand; the addresses are the lab's, known only while a
scenario runs.
- **The operator's own names, unrelated to the mesh.** Twelve `<loopback> <name>` lines for a
client's development hosts, kept in `/etc/hosts` and again in `/etc/hosts.local`, which the per-node
resolver read as additional hosts.
**A manifest never names an address, a node or a domain** ([ADR 0112](0112-a-module-definition-names-no-node-mesh-or-path.md)).
So the lab cannot list `<machine>.incus → <address>` in its definition, and the operator's twelve lines
are not any module's to define.
## Considered Options
**For a module's names:**
1. **The manifest lists its records.** Refused by ADR 0112: the addresses are the lab's runtime facts
and the scenario's choice.
2. **The module reports its records at runtime to the mesh's resolver**, which writes them into its
configuration. It works, and it makes the resolver hold every module's runtime state and decide,
per call, whether the caller may write the name it sent — authorisation for a write, on the one
server every node depends on.
3. **The module declares the zone it answers and the listen that answers it; the mesh's resolver
forwards that zone there.** The definition names a zone (from a setting) and one of its own listens,
which ADR 0112 allows; the address and the port are the mesh's facts. The records stay where they
are known — in the module, at runtime. Chosen.
**For the operator's names:**
1. **Records the controller holds, served by the mesh's resolver.** They are not the mesh's: a client's
development hosts on one machine are nothing any other node should resolve, and the controller would
become the keeper of a workstation's private notes.
2. **A node-scoped module owns `/etc/hosts`, and the operator's lines live in its kept region**
([ADR 0174](0174-a-node-varies-a-module-through-settings-and-kept-regions-never-an-edit.md)), changed
through that module's tools on that machine. Chosen.
## Decision
**1. A module that answers names declares a zone.** Its definition names the zone — a single label or a
dotted name, from a setting, never a domain the mesh knows — and the listen that answers DNS for it.
The controller refuses two modules in the mesh declaring one zone, and a zone that is the mesh's suffix,
under it, or one of a node's public domains: a module may not shadow names the mesh or the public DNS
answers.
**2. The mesh's resolver forwards each zone to the module that declared it.** The controller hands the
holder of `mesh-dns-resolver` every declared zone with the private address of the node its module runs
on and the port that listen is published on; the holder places one forwarding rule per zone into its
configuration and answers nothing in that zone itself. What names exist in the zone, and their
addresses, are the module's — answered by its own long-running code
([ADR 0198](0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)),
from its own state, as they change. Whether an answered address is reachable from the asking node is
the module's matter, not the resolver's.
**3. A node's `/etc/hosts` is held by one module, through a node seat, `node-hosts-file`.** The seat is
the definition: its holder owns `/etc/hosts`, and implements three verbs — MCP tool definitions served
as `<node>/node-hosts-file.<verb>` ([ADR 0195](0195-the-meshs-tools-are-found-by-address-not-announced-whole.md)):
**`entries`** (the file's lines, the module's and the operator's, each marked whose), **`add`** (one
address and its names, into the operator's region) and **`remove`** (one name or address from it). The
module writes the machine's own lines — loopback and the machine's name — and keeps a region for the
operator, which survives every push and is given back when the module goes. Its tools change that
region on that machine, escalating as the packet filter's do
([ADR 0175](0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md) §4). **The
controller holds none of it:** an operator's line is the machine's, not a record.
**4. No other module writes `/etc/hosts`.** The private network's region goes, as ADR 0194 already has
it; a module that once wrote a line there asks the mesh's resolver instead.
## Consequences
- **The lab's names follow its scenarios.** A scenario raised is resolvable from every node at once; a
scenario torn down is gone, with no line left behind in any file.
- **The mesh's resolver holds no module's state.** It holds the nodes' domains and a table of who
answers which zone, both composed by the controller; nothing writes to it at runtime.
- **A module answering a zone needs a DNS answerer of its own** — a long-running bundle, or a resolver
it runs. The lab gains one.
- **The operator's names reach the machine's own programs, not its containers.** A container does not
read the machine's `/etc/hosts`. For names unrelated to the mesh that is the right boundary; a name a
container needs belongs in a zone.
- **Taking `/etc/hosts` keeps what is there.** The first time the module writes the file, every line
that is not the machine's own goes into the operator's region, so a workstation's twelve lines survive
the take — the same adoption [ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md) gives
every shared file.
**How each is checked:**
- **Zones:** the controller's catalogue tests refuse a second module declaring a zone, a zone under the
mesh suffix, and a zone equal to a node's public domain.
- **Forwarding:** on the holder, the resolver's configuration carries one forwarding rule per declared
zone, at the declaring node's private address and published port; asking any node's resolver for a
name in the lab's zone while a scenario runs returns the scenario's address.
- **The hosts file:** a push leaves the operator's region byte for byte; `add` followed by `entries`
shows the line as the operator's; unassigning the module gives the region back.
## References
- [ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md),
[ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md) —
the one resolver and how nodes ask it.
- [ADR 0112](0112-a-module-definition-names-no-node-mesh-or-path.md) — why a definition names no address.
- [ADR 0174](0174-a-node-varies-a-module-through-settings-and-kept-regions-never-an-edit.md),
[ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md) — kept regions and shared files.
- [ADR 0198](0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md) —
where a zone's answerer runs.
- [Connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md) and [the seats](../03-DESIGN/01-to-be/26-the-seats.md),
amended alongside.
- [Research 023](../01-RESEARCH/023-a-seat-protocol-that-defines-what-its-holder-owns/00-overview.md) —
the general form of decision 3's "the holder owns `/etc/hosts`".
@@ -0,0 +1,78 @@
---
topic: building it
status: accepted
date: 2026-10-04
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0067-genesis-is-a-pivot.md
---
# 200. Genesis pivots to the controller as a container, and the first push hands it to a process
## Context
The controller is Go, compiled to one static binary, and is the last of the mesh's own programs a
machine runs from an image ([issue 213](../04-ISSUES/213-the-controller-is-a-go-program-run-in-a-container/00-report.md)).
[ADR 0188](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md)
§1 says a module's own code is bundles, never an image, and §3 that a service bundle is a `process` the
host runs. The handover exists: a process may name the container it `replaces`, and the host removes
that container only after the process has stayed up across two checks; two controllers are safe
together for that moment, the second standing by on the controller's consumers and every plan held by
one lock.
What stands in the way is genesis ([ADR 0067](0067-genesis-is-a-pivot.md)), which
[issue 223](../04-ISSUES/223-a-new-mesh-installs-its-controller-as-a-container/00-report.md) found
assumes an image and a container at every step from its third: it builds the controller's image,
starts a temporary controller from it, publishes it, finds the controller's container in the pivot
declaration, and from then on talks to the controller through it. A process's bundle is fetched from
the artifact store, and genesis raises the artifact store only after the pivot.
## Considered Options
1. **Raise the artifact store before the pivot**, publish the controller's bundle to it, and talk to
the controller from the host's side. Rejected for now: it reorders genesis around a store that is
itself a module the controller deploys, and rewrites the steps that talk to the controller — a
larger change to the one path that is exercised least, to remove a container that exists for
minutes.
2. **Pivot to the controller as a container, as today, and let the first push hand it over to the
process**, through the handover that already exists. Chosen.
3. **Keep the controller a container.** Rejected: it is the exception to ADR 0188 that every other
module's code has now left, and it costs a container runtime on the control machine and a
container recreation in the middle of a plan.
## Decision
**Genesis raises the controller as a container, under the resource the controller's process
`replaces`, and the first declaration the controller composes for its own machine hands it over.**
The container is genesis's own shape, built from the controller's repository, and is recorded on the
control machine exactly as the manifest's `replaces` names it, so the first apply after the pivot
finds a replacement for it and removes it once the process is up. The controller's manifest declares
only the process; the image form exists for genesis alone and is not a second way to run the
controller on a live mesh.
This is the one bounded exception to ADR 0188 §1: a module's own code in an image, for the minutes
between the pivot and the first push, on a mesh being created.
## Consequences
- A new mesh ends where a running one is: the controller a process, no controller container.
- Genesis keeps its steps; what changes is that it no longer reads the controller's container from the
manifest, and that it records the container under the name the handover expects.
- The controller's repository keeps its image build for genesis and the lab.
- The handover is now on genesis's path too: a process that fails to stay up leaves the genesis
container serving, and the apply says so — the same rule as on a live mesh.
## How it is checked
The installer's test raises a mesh whose controller manifest is the process form, and asserts that the
container genesis recorded is exactly what the process `replaces`, so the first apply hands over and
leaves one controller. Live, on the running mesh: after the manifest change is pushed, the control
machine runs the controller as a process and no controller container, and the controller's seat
answers throughout.
## References
- [Issue 213](../04-ISSUES/213-the-controller-is-a-go-program-run-in-a-container/00-report.md),
[issue 223](../04-ISSUES/223-a-new-mesh-installs-its-controller-as-a-container/00-report.md)
- mesh-host#86 (the handover), mesh-controller#252 (two controllers safe together),
mesh-controller#253 (the controller's manifest as a process)
@@ -0,0 +1,149 @@
---
topic: the mesh
status: accepted
date: 2026-10-02
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md
---
# 201. A provider declares what it derives for each consumer, and the mesh tells both ends
> Written as 0188 on 2026-10-02 and renumbered to 0201 on 2026-10-04: the record of the bundles
> refactor took 0188 on main while this one waited in a pull request, and the mesh's own code now
> cites that one. Only the number moved; the decision is the one taken on the 2nd.
## Context
An arrangement between a consumer and a provider is delivered entirely by the mesh. Where the
provider is, which port it answers on, what name the consumer must present, where its password
is — each arrives as a fact the consumer reads from its binding, or as `${bound:…}` filled into a
file before the declaration leaves the control plane. The provider invents none of it and hands
none of it back ([ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)).
One kind of value escapes that. Where the **provider names the resource** — a bucket, a database,
a vhost — the name is derived from the consumer, per consumer, and the mesh has no way to carry
it. `serves` is a literal block in the provider's definition: the same values for every consumer.
A provisioner's contract takes a provision and returns nothing. So a value the mesh's own rule
produced reaches neither end as a statement; it is recomputed at one end and transcribed at the
other.
The object store is the instance ([issue 124](../04-ISSUES/124-a-consumer-cannot-be-told-what-its-provider-derived/00-report.md)).
Its provisioner normalises the login the mesh minted into a bucket name and creates, checks and
removes exactly that; the rule lives in twenty lines of the module's own TypeScript. Its three
consumers each write the answer into their own definition by hand. Two transcribed it correctly;
one named a predecessor's bucket, and would have authenticated successfully and been refused on
every object, which reads like a credential fault and is not one.
Even corrected, the transcriptions are wrong in a second way. Each is `mesh-<node>-<slug>`, so
each **names the machine the module happens to run on today** — a definition stating a fact about
one installation, which [ADR 0155](0155-a-definition-names-no-installation-and-how-that-is-checked.md)
forbids and whose check does not catch because the name is not a domain. Move any of the three to
another machine and its configuration points at a bucket its key cannot open.
The shape is not the object store's. A database provisioner that prefixed names, a queue provider
that scoped vhosts, any provider that derives a resource from who is asking: each forces the
consumer to reproduce somebody else's rule and keep it in agreement by hand.
## Decision
**1. A served value may name the consumer the mesh is serving.** A `serves` block, which is
literal today, may interpolate the mesh's own statement of who the consumer is:
- `${consumer:as}` — the identity the mesh minted for this consumer, exactly as the login it is
told to present ([ADR 0049](0049-a-consumers-identity-fits-the-tightest-backend.md));
- `${consumer:as:dns}` — the same identity written as a DNS label.
Nothing else. **The mesh learns no protocol here; it spells its own name in an alphabet it already
knows.** The identity is the mesh's, minted by the mesh, already capped at twenty characters
because of what an S3 access key accepts; `dns` is that same name with its separator written `-`
instead of `_`, which is the whole of the difference between the mesh's identifier alphabet and
the one buckets, vhosts and hostnames use. A provider that needs a prefix or a suffix writes it
around the placeholder, because a served value is a string.
The rejected alternative is **the provider returning values from provisioning** — the natural
channel, since the provider is what derived them. It is rejected for three reasons, in order of
weight. It inverts the delivery the mesh is built on: a grant would carry data the provider wrote
rather than only data the mesh minted, and [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)
removed exactly that second path once already. It makes a consumer's declaration incomplete until
its provider's reconcile loop has run, so a consumer could not be composed before a provider
answered — a bootstrap order the mesh does not have and does not want. And it puts the rule where
nothing can check it: a value that arrives from a running process cannot be refused at resolution,
only discovered wrong later, which is the failure this record exists to end.
**2. The mesh resolves it once, per consumer, and tells both ends from the one resolution.** At the
moment a consumer's declaration is composed, the mesh knows exactly who the consumer is. There, and
only there, the placeholders are filled. The result reaches:
- the **consumer**, as the served facts in its binding file and as `${bound:<provision>:<key>}` in
any file it writes — unchanged mechanisms, carrying one more key;
- the **provider**, as `serves` on that consumer's entry in its contributions file, so the
provisioner is *told* the name rather than recomputing it.
**The provider stops deriving in code and starts declaring.** One statement, filled once, delivered
to both ends: the two cannot disagree, because there is no second computation to disagree with.
**3. A served value stays settled before it is per-consumer.** Settings still compose into `serves`
([ADR 0174](0174-a-node-varies-a-module-through-settings-and-kept-regions-never-an-edit.md)), and the consumer
placeholders are filled after that, so an operator may set a prefix and the mesh still derives the
rest. A `${consumer:…}` naming a fact or an alphabet the mesh does not have is refused when the
definition is parsed, with what it may say.
**4. A consumer may no longer name the resource its provider derives.** With the value delivered,
a literal in a consumer's definition is not merely redundant — it is the one thing that can
disagree with what the provider will actually create. The three object-store consumers lose their
hand-written bucket names in this change.
**5. A consumer that keeps several holders of one provision may not be served a derived value.**
Each holder gets its own login, `…_<local>` ([ADR 0094](0094-a-module-may-hold-several-secrets-from-one-provider.md)),
and a provider derives from the login — so it would make one resource per holder, while the
consumer's side has one binding and one `${bound:<provision>:<key>}`, both derived from the
un-suffixed identity. That is this record's own failure one case to the side, and just as quiet:
the consumer would authenticate and be refused on every object. Refused at resolution, naming
both ends. Lifting it means giving the consumer's side a local dimension, which is a decision and
not an omission.
## Consequences
- One more thing a definition may say, and one less thing a module may be wrong about. The
vocabulary grows by a placeholder; the catalogue loses three literals that named this
installation's control node.
- A provider's naming rule becomes readable in its definition instead of in its source. `minio`'s
`bucketFor` goes; the manifest says `"bucket": "${consumer:as:dns}"` and the provisioner uses
what it is given.
- A provider that already serves consumers keeps serving them: the derived value equals what the
code derived, so no bucket, database or login changes name. This is a change of **who says it**,
not of **what is said**.
- A refusal here fails **that machine's push**, naming the definition, and nothing else. That is
deliberate and is the opposite of a module quietly left out: a definition that transcribes
somebody else's rule is wrong everywhere, not just here, and the loud failure is in front of
whoever can fix it.
- The mesh now holds a rule in another system's alphabet — one rule, `dns`, stated once. A second
alphabet is a decision, not an addition: the cost of each is that the mesh must be right about
somebody else's naming, and that cost is only worth paying where the mesh already mints the name.
## How this is checked
- A served value naming an unknown fact or alphabet is refused at parse, with the list of what it
may say — tested on both halves of the message.
- Resolving a consumer whose provider derives a value puts that value in the consumer's binding
file, in its `${bound:…}` substitutions, and in the provider's contributions entry for that
consumer — one test asserting the three agree, because agreeing is the whole point.
- Two consumers of one provider on one machine get two different derived values, and neither gets
the other's.
- A consumer with several holders of a deriving provider is refused, with both ends named — the
test asserts the refusal, not merely that something failed.
- A catalogue-wide test refuses a consumer definition that writes a literal where its provider
derives: the provider's `serves` names the key, so the catalogue can say which definitions
transcribe one.
- `dns` is checked against the identity the mesh actually mints, not against an invented string:
the test derives an identity with `ConsumerIdentity` and asserts the label it becomes.
## References
- [issue 124 — a consumer cannot be told a value its provider derived for it](../04-ISSUES/124-a-consumer-cannot-be-told-what-its-provider-derived/00-report.md)
- [ADR 0048 — a provider creates the credential the mesh minted, and seals nothing](0048-a-provider-creates-the-credential-the-mesh-minted.md)
- [ADR 0049 — a consumer's identity fits the tightest backend](0049-a-consumers-identity-fits-the-tightest-backend.md)
- [ADR 0174 — a node varies a module through settings and kept regions, never through an edit](0174-a-node-varies-a-module-through-settings-and-kept-regions-never-an-edit.md)
- [ADR 0155 — a definition names no installation, and how that is checked](0155-a-definition-names-no-installation-and-how-that-is-checked.md)
- [design 27 — a module requires, the mesh resolves](../03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md)
@@ -0,0 +1,106 @@
---
topic: what runs on it
status: accepted
date: 2026-10-04
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md
---
# 202. A module keeps its current state in key-value buckets it declares, and reaches them through the runtime
## Context
A module's code reaches the bus through the node's runtime: it publishes events, subscribes to them
and asks tools ([ADR 0198](0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)).
Events are kept for a week and replayed to a consumer that was away; requests are kept nowhere. What
neither gives is **the current value of something**, seen by every machine, including one that joins
after it was written. The first module to need it — the operator's agent on a machine — registers MCP
servers for every machine as events, and a machine assigned later never hears of them; and it would
replay a week of licence rotations where it needs only the binding that holds now. Research
[024](../01-RESEARCH/024-state-a-module-keeps-on-the-bus/00-overview.md) measured the alternatives and
the grants against a real server.
[Design 32](../03-DESIGN/01-to-be/32-what-a-module-declares.md) §4 already names *state* as one of the
mesh's relationships — 1:1, last per subject — and reserves it to the mesh's own declarations.
[Design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md) §1 expects key-value buckets on the bus.
## Considered Options
1. **Key-value buckets a module declares, created by the controller, reached through the runtime.**
Chosen.
2. **State as events on EVENTS, read last-per-subject.** Rejected: retention is per stream and EVENTS
keeps seven days, so a value unchanged for a week disappears; a second stream over the same subjects
is refused by the server (design 32 §3). And events give no get, list or delete.
3. **A last-per-subject stream per module, written by hand.** Rejected: it is what a key-value bucket
is on the server, without the client's get, list, delete and watch — the mesh writing NATS's
key-value layer again.
4. **State in a module's own files or database, shared by asking a tool.** Rejected for state every
machine must see: a machine joining later has to know whom to ask and poll, and an owner that is
down answers nothing — the property the bus exists to remove.
## Decision
**1. A module declares its state by name.** `state` names the buckets it owns, by local name; every
instance of the module may write and read them. `reads` names another module's bucket as
`<module>.<name>`, read-only. A bucket's options are its owner's: how many past values a key keeps,
and how long a value lives. A manifest names no bucket, stream or subject (design 32 §1).
**2. One bucket per module per name, mesh-wide.** A key may name a machine by the module's own
convention; the mesh does not scope buckets per machine.
**3. The controller creates the buckets, from the catalogue, on every raise** — from registration,
like a seat's stream, so a reader can watch a bucket whose owner is not yet assigned anywhere. A module
never creates one. The runtime's grant on each bucket is the union of what its carried modules may do:
an owner's instances write and read, a reader's read.
**4. Each assignment is issued its buckets in its membership** ([ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)),
by the name the module uses for each and whether it may write. The runtime serves `mesh/state.get`,
`put`, `delete`, `keys` and `watch` on the bundle's channel from that list, and refuses — with the
reason — a bucket the module was not issued and a write to one it only reads. A watch delivers the
current values first, without deletions, then an end-of-current marker, then every change, each as a
`mesh/state` request the bundle answers.
**5. No secret is stored in a bucket, sealed or not.** A bucket is a stream, and design 32 §10 keeps
every secret off streams. A value that needs a secret names it; the secret travels on request/reply.
**6. The mesh caps size; a bucket outlives its module.** One value per key and no expiry unless the
owner says otherwise; at most 256 KiB a value and 64 MiB a bucket. Unassigning a module leaves its
buckets and what is in them ([ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md)); a bucket
whose declaration is gone from the catalogue is reported, never removed by the mesh.
## Consequences
- A machine that joins reads the current state at once, and every machine sees a change as it
happens, with no consumer created per reader and nothing replayed.
- The runtime's channel has a sixth verb family, and the SDKs a small state surface over it — a
contract, which ADR 0039 admits: it changes when the verbs do, rarely, and every module should be
rebuilt when it does.
- What got harder: the runtime must keep each module to its own buckets, because one principal per
machine carries all of them and the server enforces only the union. A write the server refuses
surfaces to a client as a timeout, not a refusal, so the runtime's own refusal is what a module sees.
- The secrets rule is only partly mechanical. Sealed values cannot be recognised; the runtime refuses
a value with a field whose name says it is a credential, which catches the ordinary mistake and not a
determined one. For the operator's agent this means an MCP server's authorisation header stays out of
its bucket.
- Buckets accumulate as modules come and go; that they are reported rather than removed is the price
of not deleting data.
## How it is checked
| Rule | Checked by |
|---|---|
| A manifest's state names are local, and a read names a bucket its owner declares | the catalogue's registration check, per manifest; a catalogue test that every `reads` whose owner is present names a bucket that owner declares |
| Buckets exist for every declared state | the controller's raise asserts them idempotently; its test over a real bus |
| Owners write, readers only read | the composer's test of the grants, per principal kind; the runtime's refusal test over a real bus |
| A watch hands current values first, without deletions, then changes | the runtime's test over a real bus |
| No credential-named field in a value | the runtime's refusal test |
| Live | one module puts on one machine and another machine's watch sees it; a machine assigned afterwards reads it at start |
## References
- Research [024](../01-RESEARCH/024-state-a-module-keeps-on-the-bus/00-overview.md)
- [Design 32](../03-DESIGN/01-to-be/32-what-a-module-declares.md) §4 and §10, [design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md) §1 and §3
- [ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md), [ADR 0039](0039-what-the-sdk-holds-and-refuses.md),
[ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md),
[ADR 0198](0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)
+4 -1
View File
@@ -188,7 +188,9 @@ python3 00-META/checks/index.py fail if stale
- **0185** — [A control plane behind its seat's row serves what it can](0185-a-control-plane-behind-its-seats-row-serves-what-it-can.md)
- **0186** — [A ban list never holds a neighbour, and the mesh's own bans are its own wherever they hang](0186-a-ban-list-never-holds-a-neighbour.md)
- **0187** — [A dead tracker is not the machine's failure](0187-a-dead-tracker-is-not-the-machines-failure.md)
- **0189** — [The store keeps what the records name, and a maintenance step holds its writers still](0189-the-store-keeps-what-the-records-name.md)
- **0190** — [A seat's work is shared by its holders, and building is the first such role](0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md)
- **0201** — [A provider declares what it derives for each consumer, and the mesh tells both ends](0201-a-provider-declares-what-it-derives-for-each-consumer.md)
### Its tiers, from the bottom up
@@ -225,7 +227,6 @@ python3 00-META/checks/index.py fail if stale
- **0191** — [The mesh's resolver holds only the mesh's own names; a public name resolves publicly](0191-the-meshs-resolver-holds-only-the-meshs-own-names.md)
- **0194** — [The mesh has one resolver, and every node asks it for the mesh's names](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)
- **0196** — [A node asks the mesh's resolver first, and a public one only when it is silent](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md)
- **0199** — [A module that answers names declares its zone, and a node's hosts file is one module's](0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)
### What runs on them, and how it gets there
@@ -299,6 +300,7 @@ python3 00-META/checks/index.py fail if stale
- **0195** — [The mesh's tools are found by address, not announced whole](0195-the-meshs-tools-are-found-by-address-not-announced-whole.md)
- **0197** — [Every tool announces itself on the bus, in the NATS services protocol](0197-every-tool-announces-itself-on-the-bus-in-the-nats-services-protocol.md)
- **0198** — [A module's long-running code is launched by the node's runtime, and reaches the bus through it](0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)
- **0202** — [A module keeps its current state in key-value buckets it declares, and reaches them through the runtime](0202-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md)
### How it is built
@@ -321,6 +323,7 @@ python3 00-META/checks/index.py fail if stale
- **0111** — [A build source is on the mesh's git seat, or it is an external repository](0111-a-build-source-is-on-the-git-seat-or-external.md)
- **0149** — [The live mesh is the test bed](0149-the-live-mesh-is-the-test-bed.md)
- **0174** — [A node varies a module through settings and kept regions, never through an edit](0174-a-node-varies-a-module-through-settings-and-kept-regions-never-an-edit.md)
- **0200** — [Genesis pivots to the controller as a container, and the first push hands it to a process](0200-genesis-pivots-to-the-controller-as-a-container-and-the-first-push-hands-it-to-a-process.md)
### How it is checked
-18
View File
@@ -6,16 +6,10 @@ code:
- mesh-controller examples/route-proxy
- mesh-controller internal/identity/authority.go
- mesh-controller cmd/mesh-controller/plan.go (the names the roster publishes)
- mesh-controller internal/catalogue/zones.go (the zones a module answers, ADR 0199)
- mesh-controller internal/catalogue/seats.go (mesh-dns-resolver, node-hosts-file)
- mesh-catalog modules/dnsmasq (the mesh's one resolver)
- mesh-catalog modules/resolv-conf (what a node asks)
- mesh-catalog modules/hosts (a node's /etc/hosts)
- mesh-host internal/identity/serving.go
- mesh-host internal/apply (the service that reflects a rule set)
updated: 2026-10-03
decisions:
- 02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md
- 02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md
- 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
- 02-DECISIONS/0191-the-meshs-resolver-holds-only-the-meshs-own-names.md
@@ -383,16 +377,6 @@ member's resolver answers a LAN; a router pointing at one is moved first. *Check
`/etc/resolv.conf` naming `mesh-resolver` then a public resolver, by no node but the holder answering
DNS on any address, and by the router's DHCP DNS option naming the router.*
**Names that are neither a node nor a route.** A module that answers names declares a zone (a
setting) and the listen that answers it; the controller hands the `mesh-dns-resolver` holder every
zone with its module's node address and published port, and the holder forwards that zone there and
answers nothing in it itself — the lab answers `<machine>.incus` for its running scenarios this way.
An operator's own names, unrelated to the mesh, live in `/etc/hosts`'s kept region, held per node by
the `node-hosts-file` seat's holder and changed through its tools; the controller holds none of them
([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)).
*Checked by the holder's configuration carrying one forwarding rule per declared zone, and by a push
leaving the hosts file's operator region byte for byte.*
*What follows describes the per-node resolver this replaces — how it was built and why the roles were
split. The split stands; the serving role's scope is what moved.*
@@ -1037,8 +1021,6 @@ The list is worth having in one place, because it is most of the argument:
- **One resolver ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)).** Not built: every node still runs
`node-dns-resolver`. The migration's four steps are in the record, in order.
Nor are zones or the hosts file's holder ([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)): the
workstation moves to the one resolver only once both exist, its lab and operator names depending on them.
- ~~**What happens when the hub is down.**~~ **Resolved** by
[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md), together with `06`'s
+40 -1
View File
@@ -5,10 +5,11 @@ code:
- mesh-controller cmd/mesh-builder
- mesh-controller internal/builder
- mesh-catalog modules/build-agent
updated: 2026-10-03
updated: 2026-10-04
decisions:
- 02-DECISIONS/0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md
- 02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md
- 02-DECISIONS/0189-the-store-keeps-what-the-records-name.md
- 02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md
- 02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md
- 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md
@@ -317,6 +318,44 @@ build's lines reach a reader of its subject in order and the stream holds them a
against a real server); the seat verb with an id reads the log (controller test); and, live, a build
after the roll-out read line by line through the console.
## The store keeps what the records name
*2026-10-02 — [ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md),
[issue 108](../../04-ISSUES/108-the-registry-has-no-garbage-collection-once-it-has-two-doors/00-report.md).*
Every build pushes another layer set and, until this, nothing ever removed one. The registry's own
answer — collect what no tag names — is wrong for this mesh: each artifact is pushed under one
moving tag and machines are pinned by digest, so every build but the newest is untagged and some
machine may still be running it.
**The mesh decides and the store reclaims.** Deletion is enabled on the store's one door — that
door already accepts a push, and a writer who can push can replace any tag, so delete takes
nothing a push did not already have, and ADR 0082's bargain (a store every machine reaches with no
credential to distribute first) is kept. The mesh then removes what it put there and no longer
keeps, **naming it from its own build records** rather than enumerating the store: it has never
put anything there it did not record, so a digest it did not record making is never named, which
is what keeps the sweep away from the images genesis pushed before any record existed.
An artifact stays for one of two reasons and otherwise goes: a definition the mesh holds names it
(no age limit — this is the floor), or it belongs to one of the five most recent successful builds
of its module (somewhere for a wrong release to return to). The sweep runs after a build the mesh
recorded, which is the moment new bytes landed and the moment the keep set moved; it needs no
timer. Deleting a manifest frees no bytes, so the store's own collector runs nightly as a
scheduled step with the server held still — which is what `while-stopped` exists for
([design 20](20-writing-a-module.md)). Plain collection, not `--delete-untagged`: what the mesh
keeps is still a manifest and so still referenced, and the dangerous flag is not needed once the
mesh is the one deciding.
A machine behind by more than five builds of a module, recreating a container, cannot pull what it
was running. It is already a machine the mesh reports as behind, and the answer is the current
declaration.
*How it is checked:* the keep set, against records, holds what a manifest names and the five most
recent builds and nothing else; a reference the mesh never recorded is never in the delete set; an
image and an archive are asked for at their own endpoints; a store with deletion off names the
remedy rather than the status code; a store that does not have it is recorded collected rather
than retried for ever. Live: the store's size before and after the first nightly collection.
## The builder compiles the languages the mesh is written in
*2026-09-29 —
+26 -1
View File
@@ -5,11 +5,12 @@ code:
- mesh-catalog modules/showcase
- mesh-controller internal/builder
- mesh-sdk src
updated: 2026-09-30
updated: 2026-10-02
decisions:
- 02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md
- 02-DECISIONS/0099-a-step-that-runs-once-names-what-it-reads.md
- 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md
- 02-DECISIONS/0189-the-store-keeps-what-the-records-name.md
- 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md
- 02-DECISIONS/0040-what-a-module-is.md
- 02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md
@@ -208,3 +209,27 @@ is recreated with the new fact
([ADR 0099](../../02-DECISIONS/0099-a-step-that-runs-once-names-what-it-reads.md)). *How it is
checked:* the host's unit tests run a step again when its named file changed and not otherwise,
and recreate a container naming a step after the step ran.
## A recurring step may hold its own module's containers still
*Written 2026-10-02, from [ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md)
and [issue 108](../../04-ISSUES/108-the-registry-has-no-garbage-collection-once-it-has-two-doors/00-report.md).*
Some work cannot be done underneath a running service: an artifact store's collector walks the
storage and requires every writer stopped. A `run-once` step runs *beside* containers and a
scheduled one is the same container again, so until this a module had no way to say it — and the
mesh inherited a store that has never collected anything, because the predecessor said it with a
shell script and a script beside a module is not a resource in it.
A scheduled step may name `while-stopped`: resource ids of **its own module's** containers, which
the host stops before the run and starts again after it, in the reverse order, **whatever the step
did**. Three boundaries, each refused where it can be seen earliest — its own module's containers
only, because a module that could quiesce a neighbour could stop the mesh; scheduled steps only,
because at apply the declaration is applied in order and a step already gates what follows, so a
one-time offline job says *before* rather than *instead of*; and restoring that is not conditional
on anything, because the only real risk of the field is a window that never closes.
*How it is checked:* the host's unit tests assert stop–run–start in that order, the restart after a
step that **failed**, the reverse order for several containers, and a service left down said
loudly. The controller refuses, from the definition alone, a window with no schedule, one on a
run-once step, one naming a container the module does not declare, and one naming itself.
+23 -2
View File
@@ -7,8 +7,10 @@ code:
- mesh-tools src/broker-amqp.ts (to be replaced)
- mesh-catalog modules/nats (to be written)
- mesh-sdk src (the protocol's NATS binding, step 3)
updated: 2026-10-02
- mesh-tools node-tools/internal/bus (a module's state, ADR 0202)
updated: 2026-10-04
decisions:
- 02-DECISIONS/0202-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md
- 02-DECISIONS/0167-a-membership-carries-what-its-module-receives-and-who-the-mesh-is.md
- 02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md
- 02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md
@@ -48,6 +50,7 @@ mesh's own state lives, and where what a module may say is decided by what it de
| **events** — a module says something happened | 1:many | delivered to every consumer that declared it; dead-lettered when it cannot be |
| **tools** — a module or a person asks another's tool | request/reply | one answer, from one server, or a timeout |
| **work to a role** — a module submits to a capability without knowing who provides it | job | exactly one holder does it; it queues while nobody does |
| **state** — a module's current value of something, every machine reading it | key-value | the newest per key, kept until replaced or deleted; read whole by a machine that joins later |
The last two rows are the ones worth dwelling on, because they are not messaging in the sense of
carrying bytes from A to B. **A role is addressable**, so a caller names the capability and never
@@ -56,7 +59,8 @@ changing ([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md)
property of the mesh's architecture that happens to be expressed in subjects.
And more of the mesh lands here as it is built: conditions and observed state in key-value
buckets that anything may watch, the server's own advisories becoming observations like any other
buckets that anything may watch — the first of them a module's own declared state, *2026-10-04*
([ADR 0202](../../02-DECISIONS/0202-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md)) — the server's own advisories becoming observations like any other
([research 017](../../01-RESEARCH/017-a-mesh-that-heals-itself/00-overview.md)), and a person's
client speaking the bus directly rather than through a surface built over it (§7). None of that
is a message being moved; all of it is the bus being the mesh's centre.
@@ -84,8 +88,16 @@ mesh.seat.<seat>.event.<verb> a role's own event (JetStream: EVENT
mesh.seat.<seat>.tool.<verb> a role's tool (core request/reply)
mesh.ask.<node>.<command> the controller's command api (core request/reply)
mesh.assignment.<node>.<module> an assignment's membership (JetStream: ASSIGNMENTS, last-per-subject)
$KV.<module>_<name>.<key> a module's state (JetStream: a key-value bucket per declared name)
```
**Added 2026-10-04** ([ADR 0202](../../02-DECISIONS/0202-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md)):
the last row is outside `mesh.` on purpose. A key-value bucket is NATS's own construct and lives
under NATS's own prefix, which is what lets the server's key-value layer — direct reads, rollups,
delete markers, watches — do the work instead of the mesh writing it again. The bucket is named for
the module and the local name joined by an underscore, which neither may contain, so two modules can
never derive one bucket.
**Revised 2026-10-01** ([ADR 0160](../../02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)): the rows above
for a module's and a seat's tools are the shapes the controller *issues*, not rules a runtime carries.
Every assignment is published a membership — what it serves and where, in which queue, its seat verbs,
@@ -155,6 +167,7 @@ Core NATS is at-most-once. Everything the mesh must not lose lives in a JetStrea
| CONTROL | `mesh.control.>` except `alive` (a build's outcome moved to its seat, ADR 0121) | work queue, one consumer (the controller), explicit ack | the store-window guarantee ([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)): the controller `nak`s with a delay while its store is away and the message is redelivered; nothing is dropped |
| NODES | `mesh.node.>` | last per subject | one declaration per node, always the newest |
| EVENTS | `mesh.mod.*.event.>` and `mesh.seat.*.event.>` | limits (age, size), durable consumer per subscribing module | a subscriber that was down catches up; after `max-deliver` attempts the advisory feeds `mesh.events.dead` (its own small stream). *2026-10-01:* a build's whole log is here too, as the build-machine seat's `log.<build id>` events ([ADR 0157](../../02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md)) — one subject per build, a week of retention, read back by `builds --log <id>` with a consumer that is gone when the reading is done |
| `KV_<module>_<name>` | `$KV.<module>_<name>.>` | the newest value per key — as many past values as the owner declared — no age unless the owner declared one; a value at most 256 KiB, a bucket at most 64 MiB | a module's state (ADR 0202): one per name in a manifest's `state`, created from the catalogue on every raise, so it exists before its owner runs anywhere; kept when the module is unassigned, because what it holds is data |
Tool calls and heartbeats stay on core NATS: a lost heartbeat is the next heartbeat; a lost tool
call is a timeout the caller already handles.
@@ -234,6 +247,14 @@ expresses this exactly, per subject, and better than a vhost could:
permissions for each consumed event's subject, its tool subjects, and that same inbox prefix.
Nothing else. A module that tries to publish outside its emits is refused by the server, not by
convention.
- **A module's state** (ADR 0202), for whichever principal carries the module — today the machine's
runtime, whose grant is the union of its modules': binding to the bucket, reading a key directly,
and an ordered consumer for listing and watching, created and deleted on the bucket's own stream
and nothing else's; and, for the owner's instances only, publishing under the bucket's own
`$KV.<bucket>.>`. *Measured 2026-10-04 against a running server:* without the consumer-delete
grant a watch cannot be stopped cleanly, and a write the server refuses reaches the writer as a
timeout rather than a refusal — so the runtime refuses first, from the membership, and the grant
is the second line.
- **The controller's user** owns `mesh.control.>`, `mesh.node.>` and the streams, and may submit work
to the seats the mesh's own flows use — a build, for one (ADR 0121).
**A host's user** may publish its own `mesh.control.<node>.>` and subscribe its own
-2
View File
@@ -12,7 +12,6 @@ code:
- mesh-catalog modules/gitea/module.json
updated: 2026-10-03
decisions:
- 02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md
- 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
- 02-DECISIONS/0161-what-deserves-a-seat.md
- 02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
@@ -130,7 +129,6 @@ convention, which later seats departed from.
| `mesh-build-machine` | `the-build-machine` | node | — | a builder |
| `mesh-resolver` | — | mesh | — | the mesh's one resolver, holding every node's internal domain ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)) |
| ~~`mesh-dns-port`~~ | `the-dns-port` | node | — | retired by [ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md): the local resolver became the mesh's one |
| `node-hosts-file` | — | node | — | owns `/etc/hosts`: the machine's own lines and the operator's kept region, changed through its verbs `entries`, `add`, `remove` ([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)) |
| `mesh-intrusion-prevention` | `the-intrusion-prevention` | node | — | an intrusion-prevention service |
| `mesh-packet-filter` | `the-packet-filter` | node | — | the packet filter |
| `mesh-private-network` | `the-private-network` | node | — | the private network the mesh runs over |
@@ -2,7 +2,7 @@
layer: to-be
status: in-progress
code: [mesh-controller internal/catalogue]
updated: 2026-09-30
updated: 2026-10-02
decisions:
- 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md
- 02-DECISIONS/0155-a-definition-names-no-installation-and-how-that-is-checked.md
@@ -14,6 +14,7 @@ decisions:
- 02-DECISIONS/0084-which-provider-serves-a-consumer.md
- 02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md
- 02-DECISIONS/0038-the-mesh-assigns-the-port.md
- 02-DECISIONS/0201-a-provider-declares-what-it-derives-for-each-consumer.md
---
# 27 — A module requires, the mesh resolves
@@ -206,6 +207,22 @@ name when nothing sets it. That is the contract half of this design's operator p
the placeholder allows: the definition says which values reach which requirement, and nothing else
does. *How it is checked:* the unit tests named in issue 173, and the plan comparison that closed it.
*A provider says once what it derives for each consumer (2026-10-02,
[ADR 0201](../../02-DECISIONS/0201-a-provider-declares-what-it-derives-for-each-consumer.md),
[issue 124](../../04-ISSUES/124-a-consumer-cannot-be-told-what-its-provider-derived/00-report.md)):*
where a provider **names the resource** it gives each consumer — a bucket, a database, a vhost — the
name is derived per consumer, and a literal `serves` block could not carry it. A served value may
now name the consumer the mesh is serving: `${consumer:as}`, the identity the mesh minted, and
`${consumer:as:dns}`, that same identity written as a DNS label. Nothing else — **the mesh learns no
protocol here; it spells its own name in an alphabet it already knows.** Settings are laid on first,
so an operator may still set a prefix and the mesh derives the rest. The mesh fills it at the one
moment it knows who the consumer is, and the one filled value reaches both ends: the consumer, as
its binding's served facts and as `${bound:<provision>:<key>}` in any file it writes; the provider,
as `derived` on that consumer's entry in its contributions file, so its provisioner is told the name
rather than recomputing it. A consumer that writes the derived value into its own definition instead
of asking for it is refused, naming the placeholder to use. *How it is checked:* the unit tests in
ADR 0201's "how this is checked", each run against the unchanged controller first.
## How a definition reads what was resolved
**One form, naming a requirement and a field of its contract.** A definition that needs the database's
@@ -214,7 +231,9 @@ name in a configuration file writes the same thing: the requirement's name and t
controller fills it at resolution.
This one form replaces the placeholders that exist today, one per mechanism: bound values, secrets,
ports and machine facts.
ports and machine facts. It subsumes the consumer placeholder too — a value a provider derives is
read by the consumer exactly as any other field of the contract is, and `${consumer:…}` is only
how the *provider* states the rule.
**The seat placeholder stays, for the controller alone.** The controller composes its own
declaration and reaches the store and broker it made before any module existed, so it cannot be
@@ -11,8 +11,10 @@ code:
- mesh-host internal/apply/apply.go
- mesh-tools src/main.ts
- mesh-catalog modules/mesh-catalog
updated: 2026-10-02
- mesh-tools node-tools/internal/runtime (a module's state, ADR 0202)
updated: 2026-10-04
decisions:
- 02-DECISIONS/0202-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md
- 02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md
- 02-DECISIONS/0126-a-module-declares-its-own-seats.md
- 02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
@@ -67,6 +69,8 @@ the catalogue, and the mesh would have hundreds of copies of a decision it made
| `tools: status` | queue-group subscription on `mesh.mod.<module>.tool.status` |
| seat `telegram-sender`, `accepts: send` | work-queue consumer on `mesh.seat.telegram-sender.accept.send` |
| `uses: telegram-sender` | publish on that seat's `accept` subjects, and nothing else |
| `state: servers` | a key-value bucket for the module, created by the controller; its instances write and read it |
| `reads: billing.orders` | read and watch billing's `orders` bucket, and nothing else of it |
**Wildcards, and they are the mesh's rather than a bus's.** *Added 2026-09-27, from
[issue 127](../../04-ISSUES/127-a-module-event-derives-a-subject-nothing-publishes/00-report.md).* A
@@ -221,7 +225,7 @@ service" versus "one worker per machine".
| credential | sealed, per consumer | none | none | none | none |
| reply | — | none | none, or an event later | a report | awaited |
| retention | — | age and size | work queue, explicit ack | **last per subject** | none |
| declared | `provides`/`requires` | `emits`/`consumes` | seat `accepts` / `uses` | the mesh's own | `serves` |
| declared | `provides`/`requires` | `emits`/`consumes` | seat `accepts` / `uses` | the mesh's own; a module's `state` / `reads` | `serves` |
**Job** is the one [ADR 0041](../../02-DECISIONS/0041-events-are-a-relationship.md) had no room
for. Its table has events at 1:many and provisions at 1:1; a module submitting work to a service
@@ -236,6 +240,31 @@ last-per-subject retention, and a node that has seen sequence *n* refuses *n−1
That is the wire-level answer to
[issue 107](../../04-ISSUES/107-a-declaration-carries-no-order/00-report.md).
**A module declares state too.** *Added 2026-10-04,
[ADR 0202](../../02-DECISIONS/0202-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md).*
State was the mesh's alone, and modules had the same need with nowhere to put it: an MCP server
registered for every machine, sent as an event, never reached a machine assigned afterwards — its
consumer did not exist yet when the event passed — and a licence binding sent as events replays a
week of rotations where only the latest matters. So a module names the state it **owns** with
`state`, and another module's it **reads** with `reads: <module>.<name>`. Each is a key-value bucket
the controller creates from the catalogue, mesh-wide, existing from registration so a reader can
watch before the owner runs anywhere ([design 25](25-the-bus-on-nats.md) §3). Every instance of the
owner writes; a reader reads and watches. A key may name a machine by the module's own convention;
the mesh keeps one bucket per name, not one per machine, because "every server, for every machine"
is then one list rather than a walk.
What a module sees is what it named. Its assignment's membership lists its buckets by those names,
with whether it may write, and the runtime answers `get`, `put`, `delete`, `keys` and `watch` for
them on the bundle's channel — refusing, with the reason, a name it was not issued or a write to a
bucket it only reads. A watch hands the current values first, then every change: a bundle that starts
late, or starts again, has the whole of the state before it has any of the news.
The owner says how many past values a key keeps and how long a value lives, as a seat says how long
its backlog survives (§3); the mesh caps a value's size and a bucket's. **A bucket outlives its
module's assignment** — what a module stored is data
([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)) — and one whose
declaration is gone is reported, never removed by the mesh.
## 5. Seats
A module declares a seat with its protocol, and the mesh enforces one holder at its scope
@@ -454,6 +483,14 @@ sealing key leaks, that stream is an archive rather than a moment. So:
existing discipline — *fetched from it, not carried* — applied to the one payload where carrying
it is worst.
**A key-value bucket is a stream, so the same holds for it.** *Added 2026-10-04,
[ADR 0202](../../02-DECISIONS/0202-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md).*
No secret is put in a module's state, sealed or not: state is exactly what a machine joining a year
later reads in full. A value that needs a secret names it, and the secret travels on request/reply.
Sealed values are plain text to anything inspecting them, so this is checked only partly — the
runtime refuses a value carrying a field whose name says it is a credential, which catches the
ordinary mistake and not a determined one.
**The bootstrap, which is circular and has a precedent.** The vault makes every secret
([ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md)), including the bus's own
passwords. The vault is a module, and a module needs a bus account, whose password the vault
@@ -523,3 +560,9 @@ billing existing under that name.
on, and exactly those two are rebuilt.
- **A stale declaration is refused.** A bed: replay sequence *n−1* after *n*, and the node refuses
it rather than applying it.
- **A module reaches only the state it declared.** The composer's test: an owner's runtime may write
its buckets, a reader's may only read, and nothing else is granted; the runtime's test over a real
bus: a name not issued and a reader's write are refused with the reason.
- **State is current at once.** The runtime's test over a real bus: a watch hands the current values
without deletions, then an end-of-current marker, then changes. Live: a machine assigned after a put
reads it at start.
@@ -2,7 +2,7 @@
layer: to-be
status: in-progress
code: [mesh-tools, mesh-controller, mesh-host, mesh-catalog]
updated: 2026-10-03
updated: 2026-10-04
decisions:
- 02-DECISIONS/0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md
- 02-DECISIONS/0173-the-operators-machine-is-the-meshs-and-a-module-is-what-it-declares.md
@@ -309,6 +309,25 @@ nextcloud, minio), the two whose clients exist on no system (mongodb, mssql: a d
and last the mesh's own (mesh-catalog, mesh-vault, records, gitea, mailu, audit-logger, lab, and the
three mains).
*Built 2026-10-04.* Thirty-four modules no longer run their own code in a container: the first wave
(mesh-catalog#245), the mesh's own and the media modules (mesh-catalog#248, mesh-media-catalog#13),
with a step run where and as it is declared (mesh-host#85) and a process's words filled like a
container's (mesh-controller#250). **Proven live** on every machine that runs them: each moved
module's tools answer from the node's runtime, the steps run as their oneshot units, and the forge's
merge events reach the build pipeline from the runtime — the merge after the move started its own
plan. **Two corrections the machines taught:** a tool that called a broker's command-line client now
runs it inside the broker's own container, because a host package may be uninstallable on a machine
whose package index is stale (mesh-catalog#249); and a module reading its application's own key reads
it through the application's container, because that directory belongs to the account the application
runs as there, which is not the runtime's (mesh-media-catalog#14). *Completed 2026-10-04:* the last three — mesh-catalog, mongodb and mssql, whose code imports npm
packages of its own — moved once the builder installs a bundle's own dependencies before compiling,
keeping the toolchain's SDK authoritative (mesh-controller#255, mesh-catalog#250). Their database
clients are now drivers inlined into the bundle, not command-line clients fetched by a container; one
more correction the machines taught: a driver reaching its server on loopback must give TLS a host
name, since the runtime's Node refuses an address (mesh-catalog#252). **No module's own code runs in
a container any more;** every module with tools answers from its node's runtime, proven by calling a
tool of each.
## WP5 — The shell, on a server first
*mesh-catalog #224, already written. Half a day to assign and prove.*
@@ -1,9 +1,9 @@
---
status: open
status: resolved
opened: 2026-09-23
located-in: []
fixed-by:
amended-design:
located-in: [mesh-controller internal/inventory, mesh-controller internal/artifacts, mesh-host internal/apply, mesh-catalog modules/distribution]
fixed-by: 02-DECISIONS/0189-the-store-keeps-what-the-records-name.md
amended-design: 03-DESIGN/01-to-be/18-building-a-module.md
---
# 108 — The registry has no garbage collection, and two doors make it harder to add
@@ -75,3 +75,32 @@ real thing services need, and the mesh cannot express one.
images by digest and moves by version — is retention "the digests no recorded build names"?
- Who owns the routine when the store and its public door are two modules — the store, since the
volume is its?
## Answered, 2026-10-02 — [ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md)
The three open questions, answered:
- **A maintenance step, or a backend that does not need its writers stopped?** The step. A
scheduled container may name `while-stopped` — resource ids of **its own module's** containers,
which the host stops before the run and starts again after it whatever the step did. A storage
backend the mesh does not run would be a bigger thing to own than the mechanism it avoids, and
the mechanism is wanted anyway: a service that cannot have work done underneath it is a real
shape and the mesh could not express it at all.
- **Is retention "the digests no recorded build names"?** Nearly. An artifact stays because a
definition the mesh holds names it (no age limit), or because it belongs to one of the five most
recent successful builds of its module. Last-N-tags was the predecessor's rule for a registry
that knew nothing else; this mesh knows what each digest is for.
- **Who owns the routine now the second door is gone?** Both halves, each where it can be. The
**mesh** decides what may go — only it holds the records — and asks the store to drop it. The
**store** reclaims the bytes, because only it can stop its own server. Neither half can be done
by the other.
And the sharpened point — enabling deletion on a door with no accounts — dissolved on inspection:
**that door already accepts a push**, so a writer who can reach it can already replace any tag.
Delete takes nothing a push did not have. What it does not do is undo ADR 0082's bargain, which
putting an authenticated door in front of deletion would have.
The second registry process is not built, as the 2026-09-26 note says, so the shared blob cache
and the deletion-cached-by-the-other-door problem never arise. Plain `garbage-collect` is enough:
what the mesh keeps is still a manifest in the store, so `--delete-untagged` — the flag that would
delete images machines are running — is not needed at all.
@@ -1,9 +1,9 @@
---
status: located
status: resolved
opened: 2026-09-26
located-in: [mesh-controller internal/catalogue/declaration.go, mesh-sdk src/provisioner, mesh-catalog modules/minio]
fixed-by:
amended-design:
located-in: [mesh-controller internal/catalogue, mesh-sdk src/provisioner, mesh-catalog modules/minio]
fixed-by: 02-DECISIONS/0201-a-provider-declares-what-it-derives-for-each-consumer.md
amended-design: 03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md
---
# 124 — A consumer cannot be told a value its provider derived for it, so it transcribes one
@@ -63,3 +63,27 @@ compares it to what the provider will actually create. The one wrong instance wa
- What would have caught the wrong instance? A test that resolves a consumer's grant and compares the
bucket in its own configuration against the one the provider would create is a check that could
exist today, for any interface, without the mechanism above.
## Answered, 2026-10-02 — [ADR 0201](../../02-DECISIONS/0201-a-provider-declares-what-it-derives-for-each-consumer.md)
The channel is the provider's own `serves` block, which may now name the consumer the mesh is
serving: `${consumer:as}` and `${consumer:as:dns}`. The mesh fills it once, where it knows who the
consumer is, and delivers the one filled value to both ends — the consumer's binding and its
`${bound:…}` substitutions, and the provider's contributions entry, so a provisioner is told the
name rather than deriving it. Each open question above, answered:
- **Should a provider return values from provisioning?** No. It would make a grant carry data the
provider wrote, make a consumer's declaration wait on its provider's reconcile loop, and put the
rule where nothing can refuse it. The reasoning is in the record.
- **Or should `serves` say a value is derived?** Yes, and the mesh performs the derivation — but it
learns no protocol doing it. The only fact is the identity the mesh itself minted, in one of two
alphabets it already knows.
- **Should a consumer that names the resource be refused?** Yes. A consumer's file that already
contains the value the mesh is about to derive for it is refused at resolution, naming the
placeholder to write instead. That is the check this report asked for, and it is exact rather than
heuristic: a derived value carries the identity minted for this consumer on this machine, which
nothing else would spell out.
minio's `bucketFor` is gone; its manifest serves `"bucket": "${consumer:as:dns}"`. The three
consumers' hand-written bucket names are gone with it — each of them also named the machine the
module happens to run on, which is the second thing wrong with a transcription.
@@ -0,0 +1,93 @@
---
status: open
opened: 2026-10-02
located-in: [mesh-catalog modules/dnsmasq, mesh-controller cmd/mesh-controller]
fixed-by:
amended-design:
---
# 202 — A module whose required setting nobody set is left out of the machine, and the resolver is the module it happened to
## What was observed
Running the controller's own test suite against the catalogue beside it, 2026-10-02.
`TestTheResolverIsToldEveryMachineOnTheNetworkAndToldAgainWhenOneLeaves` fails with *"the resolver
was not handed the machines"*. Composing the same machine by hand and listing what it receives
shows why: **dnsmasq contributes nothing at all.** Four resources are composed for that node, all
of them the overlay's. The resolver's package, its configuration, its service and the fact that
carries every machine's name are simply not there.
The cause is one line added to `dnsmasq`'s configuration earlier the same day: the addresses it
listens on beside the machine's own became an operator setting,
`listen-address=${setting:listen-addresses}`, with no default. A `${setting:…}` nothing sets is
refused, a module that cannot be composed is **left out** rather than failing the whole machine
([ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md)), and so a node
assigned the resolver is handed a declaration with no resolver in it.
The failing test is the symptom that surfaced it. The test is not what is wrong.
**Proven rather than inferred.** Composing the same machine a second time with
`listen-addresses` set to `127.0.0.1` and nothing else changed, every one of dnsmasq's eight
resources appears — `needs-broker`, `mesh-state`, `package`, `config`, `runtime-dns`, `runtime`,
`service` and `fact-node-zones`. The only difference between a machine with a resolver and a
machine without one is whether somebody set a value that did not exist yesterday.
## Why it matters beyond this instance
**Leaving a module out is right, and being quiet about it is not.** The rule exists so one
module's broken setting cannot stop a machine converging — a good rule. But the outcome here is a
machine that applies cleanly, reports current, and is missing its DNS resolver. Every name on that
machine then resolves through whatever was there before, or not at all, and nothing in the mesh
says the resolver was dropped. That is the shape
[issue 152](../152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md) records for
routed names, here for a whole module.
**And a setting with no default is a definition that cannot be assigned.** Every other
`${setting:…}` in the catalogue names something that is genuinely particular to one installation —
a public domain, an issuer. "Which addresses besides my own do I answer on" has an obvious correct
default for every machine that is not a LAN gateway: none beside loopback. A definition that
refuses to compose until somebody sets a value most machines do not need is a definition that
breaks the next node to be assigned it, and genesis with it.
## What this does not claim
Whether the live machines are affected was not checked — those four have had the setting set, or
their resolvers would already be gone. The claim is about a machine assigned the resolver *from
now on*, and about the silence.
## Open questions
- Should the declaration say which modules it left out, where a person or the console can see it?
`left_out` already travels to the host ([ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md));
what is missing is anything that reads it back and says so.
- ~~Should a `${setting:…}` be allowed a default?~~ **Decided in principle and not built.**
[ADR 0164](../../02-DECISIONS/0164-a-setting-is-declared-with-its-default-its-meaning-and-what-changing-it-costs.md)
(proposed, 2026-10-01) says a setting with a default is a tunable and one without is the
operator's, and narrows 0155's refusal to exactly the second. `listen-addresses` is a tunable by
that rule, and dnsmasq declares no settings block at all. So this issue is, in part, 0164 waiting
to be built — and in part the silence, which 0164 does not address.
- Is leaving a module out ever right for a module a node is **assigned**, as opposed to one it
merely pulls in? An assignment is somebody saying *this machine runs this*; silently not running
it is the one answer nobody asked for.
## Still true on 2026-10-04, and the evidence had to be re-taken
Re-checked after the bundles refactor landed (fourteen records, ADRs 0188 and 0190–0200). **The
fault stands and the old evidence no longer reaches it.**
`TestTheResolverIsToldEveryMachineOnTheNetworkAndToldAgainWhenOneLeaves` still fails on
mesh-controller main, with the same message — and now for a *different first reason*. `assign` is
refused before composition ever happens:
> dnsmasq on anchor has no bus credential: nothing was issued for anchor.dnsmasq … (novox/hq issue 203)
That is [issue 203](../203-a-fresh-assignment-is-pushed-before-its-credential-exists/00-report.md)'s
new guard doing its job on a test harness that mints no credential. Two faults are stacked in one
failing test, and the second was invisible behind the first.
Proven again, past both: mint `anchor.dnsmasq` so the assignment stands, then compose the machine
twice. **Without `listen-addresses` the node composes four resources, all the overlay's. With it
set to `127.0.0.1`, all eight of dnsmasq's appear** — `needs-broker`, `mesh-state`, `package`,
`config`, `runtime-dns`, `runtime`, `service`, `fact-node-zones`. Nothing else differs.
The test is now wrong about two things and should be fixed with whichever of these is fixed first.
@@ -1,9 +1,10 @@
---
status: located
status: resolved
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by:
- mesh-controller#246
amended-design:
---
@@ -45,3 +46,7 @@ have to write — a bundle of language L depends on the module that publishes L'
bundle lands in the tier after it. A test: a merge touching the toolchain module and a TypeScript
bundle plans the bundle one tier later. Worked around on the day by building the bundle again once
the toolchain was built.
## Resolved
Every mesh-tools plan since the fix ran in two rounds: the toolchain and runtime images first, what is built in or on them after. Before it, a merge moving both tiered them together. The planner's test proves the order: a merge moving a bundle and its toolchain plans two rounds.
@@ -0,0 +1,11 @@
# 211 — Diagnosis
*2026-10-03.* The planner orders a merge's modules by `inventory.Dependencies`, whose edges come from a
manifest's `build.on`, from what a build recorded it stood on, from the repositories it read, and from
the build machine. A bundle names its toolchain by `language`; the builder takes the toolchain image
(`ToolchainFor(language)`) from what the mesh holds and records nothing of it as stood on. So no edge
ran from a bundle to the module publishing its toolchain, and a merge moving both (mesh-tools: the
images and node-tools) tiered them together. **Fix (mesh-controller, branch
`fix/issue-211-a-bundle-stands-on-its-toolchain`, commit c72f6ca):** `dependenciesOf` adds a `stands-on`
edge from every bundle artifact to its toolchain's module, read from the manifest. Tested: TypeScript
bundle → mesh-tools, Go bundle → mesh-tools-go, image → none; a merge moving both plans two tiers.
@@ -1,10 +1,11 @@
---
status: located
status: resolved
opened: 2026-10-03
located-in:
- mesh-tools
- mesh-controller
fixed-by:
- mesh-tools#44
amended-design:
---
@@ -42,3 +43,7 @@ install from a lockfile that a release updates, or build the install layer witho
the build should record which SDK version the image carries, so a bundle's record says what it was
compiled against. A check: after an SDK release and a toolchain rebuild, a bundle built on it reports
the released version.
## Resolved
The toolchain now installs the exact SDK version the mesh last published, passed in as a build argument from the SDK module's published package, and the planner orders the toolchain after the SDK. Proven 2026-10-04: the toolchain built with the published SDK, and the six TypeScript bundles built in it serve their tools and seat verbs.
@@ -1,10 +1,14 @@
---
status: open
status: resolved
opened: 2026-10-03
located-in:
- mesh-controller
- mesh-catalog
- mesh-host
fixed-by:
- mesh-host#86
- mesh-controller#252
- mesh-controller#253
- mesh-host#88
amended-design:
---
@@ -50,3 +54,27 @@ of a plan today.
Located only by owner; the move is a change of the controller's module and its deployment, not of
its code.
## Fix prepared (2026-10-04)
Three changes. mesh-host#86, merged: a process may name the container it `replaces`, and the host
removes that container only once the process has stayed up across two checks. mesh-controller#252,
awaiting the operator's merge: the controller's composition for a service process, and two
controllers safe together for the handover — the second stands by on the controller's consumers
until the first lets go, and all plan work holds one advisory lock. mesh-controller#253, held: the
controller's manifest as a Go bundle and a process. It waits on
[issue 223](../223-a-new-mesh-installs-its-controller-as-a-container/00-report.md), because with it a
new mesh cannot be installed.
## Resolved
Proven 2026-10-04 on the control machine: after the manifest change was merged and pushed, the host
created the controller's account, ran its preparation step, started the controller as a process,
found it up across both checks and removed the container. The controller now runs as its own account
from its bundle, no controller container remains, its seat answered throughout, and the merge's own
plan finished all three tiers under the new process.
One fault on the way, fixed before it could leave two controllers or none: the host read the account
not existing yet as a user database that did not answer — it matched the exit as text in a wording
its own runner did not use — so the first apply stopped at the account and the container kept serving,
which is the handover's safe failure (mesh-host#88).
@@ -1,9 +1,10 @@
---
status: open
status: resolved
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by:
- mesh-controller#246
amended-design:
---
@@ -34,3 +35,7 @@ Owner mesh-controller (the planner). **Fix direction:** on start, and whenever a
build, the plan settles an `asked` build against the build records — a build recorded as built from
the plan's commit is that tier's outcome — so a plan resumes after the controller replaced itself.
A test: a plan whose build outcome was recorded while no controller followed it resumes on start.
## Resolved
A plan settles an asked build from the build records, whoever heard the outcome. Proven 2026-10-03 and 2026-10-04: both controller merge plans since the fix finished all three rounds, including the round that replaced the controller, without being stopped by hand.
@@ -0,0 +1,10 @@
# 214 — Diagnosis
*2026-10-03.* A plan learns a tier's outcome only through `planBuilt`, called when a controller takes
in a build result off the bus. A merge to the controller's repository replaces the controller in tier
0; the build that produced the new controller was recorded, but the plan state the new controller
loaded still read `asked`, and no path ever revisited it. **Fix (branch
`fix/issue-214-a-plan-settles-from-the-build-records`, commit d86baeb):** `advanceOnce` settles every
still-asked module from its build records — a build recorded after the ask is that ask's outcome,
built or failed — on every advance and on the 30-second ticker. Tested with a pure helper. Live
proof: the next merge to mesh-controller passes tier 0 on its own.
@@ -1,9 +1,10 @@
---
status: open
status: resolved
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by:
- mesh-controller#246
amended-design:
---
@@ -33,3 +34,9 @@ Owner mesh-controller. **Fix direction:** a build asked at a commit does not cha
module follows; or, if pinning is meant, the pin is said — in `module list`, in `status`, and by a
merge's plan naming the module it leaves out and why. A test: building a module at a commit and then
merging a change to it plans it.
## Resolved
Proven 2026-10-04: the catalogue merges since the fix rebuilt the module that had been pinned at an
old commit, at the merge's commit, by an ordinary plan — the same as every other module of the
catalogue. Nothing was asked for it by hand.
@@ -0,0 +1,10 @@
# 215 — Diagnosis
*2026-10-03.* `takeIn` registers a build's `Ref` as the branch the module follows. unifi was once
built with `ref=9c97a8a`, which became its followed ref. `sourceIs` matches a merge only to modules
whose ref is empty or the merged base — so every merge into main left unifi out — and `askTier`
re-asks `Source.Ref`, so every plan that rebuilt unifi built the same old commit again (its build
records all read "at 9c97a8a"). **Fix (branch `fix/issue-215-a-commit-is-never-a-branch-to-follow`,
commit 6784efa):** registration keeps the followed branch when a build names a commit; matching and
re-asking read a recorded commit as the default branch, healing existing records; a merge names the
modules of its repository it leaves out. Store-backed test fails without the fix.
@@ -1,9 +1,10 @@
---
status: open
status: resolved
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by:
- mesh-controller#246
amended-design:
---
@@ -28,3 +29,7 @@ Owner mesh-controller (the catalogue's registration check). **Fix direction:** a
loads, runs or unpacks — no `loads`, no `tools` list on its module, no resource naming it — is refused
at registration, naming the field that would deliver it. A test: such a manifest is refused; adding
`loads` admits it.
## Resolved
A bundle nothing would deliver is refused at registration, by name. Proven by the controller's tests; the catalogue's bundles all name what the runtime loads (mesh-catalog#243).
@@ -0,0 +1,9 @@
# 216 — Diagnosis
*2026-10-03.* The composer delivers a bundle as an archive only when its `Loads` is non-empty, and
`Loads` derives from the artifact's `loads` or, failing that, from the module's `tools` list. The
seven modules had neither, so their bundles were recorded and never composed into any declaration;
nothing checked it. **Fix (branch `fix/issue-216-a-bundle-nothing-delivers-is-refused`, commit
cf2bb3b):** registration refuses a bundle that nothing loads, runs or unpacks — no `loads`, no `tools`,
no resource naming it, and not the runtime — naming the field that would deliver it. The current
catalogue passes the check.
@@ -0,0 +1,58 @@
---
status: resolved
opened: 2026-10-03
located-in:
- mesh-tools
fixed-by:
- mesh-tools#42
- mesh-tools#43
amended-design:
---
# 217 — A refused announcement took down every container's runtime, and the console with it
## What was observed
2026-10-03, rolling out [ADR 0197](../../02-DECISIONS/0197-every-tool-announces-itself-on-the-bus-in-the-nats-services-protocol.md).
The tool runtimes announced themselves by subscribing `$SRV.<verb>.>`; the grants composed for them
allowed only `$SRV.<verb>` and the service's own name and instance. The bus refused the wildcard:
```
Subscription Violation - User "novox.gitea", Subject "$SRV.PING.>"
NatsError: 'Permissions Violation for Subscription to "$SRV.PING.>"'
```
The TypeScript runtime the per-module containers run treats a refused subscription as fatal, so on
the one machine that had received the new images nine containers crash-looped — gitea's runtime,
postgres's, the catalogue, the vault, mongodb, mssql, keycloak, mailu, nextcloud. Gitea's tools went
with them, which closed the usual path for merging the fix.
Then the console stopped answering the controller's verbs, though the controller held every
subscription: the console builds the list that tells a seat's verb from a module's tool by asking the
controller *and* the catalogue, and gives up on both when the catalogue does not answer — so
`mesh-controller.status` was asked of a module subject nobody serves.
## Why it matters beyond this instance
Two properties, each worse than the mistake that exposed it:
- **A runtime dies for an optional subscription.** Announcing is discovery; serving tools and running
provisioners is the work. A refusal of the first should never stop the second.
- **The console's view of the mesh's own verbs depended on a module.** The controller's verbs are how
the operator repairs the mesh; they must not become unreachable because the catalogue is down.
## Diagnosis
Owner mesh-tools. The wildcard is fixed on branch `fix/announce-only-what-the-grants-allow`: both
runtimes subscribe exactly what the grants allow. The console that discovers from what announces itself
(ADR 0195, 0197, on main) asks the bus and the controller, not the catalogue, which removes the second
property once it is deployed. **Still to do:** the TypeScript runtime treats a refused announcement
subscription as a logged warning, not as fatal; a test against a bus with real grants proves the
announcement subscriptions are allowed for every principal kind.
Recovered on the day without the forge's API: the toolchain images built by the controller straight
from the fix branch, every module image rebuilt on them, and the machine pushed.
## Resolved
A refused announcement is logged and the runtime serves on, and every runtime subscribes only the discovery subjects its grants allow. Proven 2026-10-04: after rolling out to every machine, no container restarts anywhere, and the discovery console lists no runtime as not answering. A refused tool subscription was made non-fatal the same way afterwards, under issue 218 (mesh-tools#46).
@@ -0,0 +1,92 @@
---
status: resolved
opened: 2026-10-03
located-in:
- mesh-controller
- mesh-tools
fixed-by:
- mesh-controller#248
- mesh-tools#45
- mesh-tools#46
amended-design:
---
# 218 — A seat held once for the mesh is answered by a module on a machine that does not hold it
## What was observed
2026-10-03. Asked which databases the mesh's store holds, `mesh-store.databases` answered from the
postgres on one machine with that machine's application databases; the controller's own database
lives on the postgres of the other machine, which the controller's records name as the seat's one
holder:
```
mesh-store scope: mesh delivers: postgres-database holders: [ {node: <the control machine>, module: postgres} ]
```
The discovery console, which reads what answers on the bus ([ADR 0197](../../02-DECISIONS/0197-every-tool-announces-itself-on-the-bus-in-the-nats-services-protocol.md)),
shows the same seat announced from **both** machines running postgres.
## Why it matters beyond this instance
A seat held once for the mesh promises one answerer: the role's holder. A module that implements a
seat's verbs on every machine it runs on, and is let serve them on each, turns "the mesh's store" into
"whichever postgres replied first" — a read against the wrong database that looks like a right one,
and a write would be worse. Every mesh-scoped seat whose implementing module runs on more than one
machine has this shape.
## Where to look
Whether the runtime serves a seat's verbs where its module merely *claims* the seat rather than where
the mesh made it the holder ([ADR 0159](../../02-DECISIONS/0159-a-tool-call-names-the-machine-and-a-holder-serves-its-seats-verbs.md),
[ADR 0160](../../02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)):
the membership the controller issues each assignment, and what the runtime admits from it. A check:
a mesh-scoped seat's verbs are served by exactly the holder the records name, on every machine.
## Root cause
The controller composed each assignment's held seats from what its module *claims*, once per module
and not once per machine. Every machine running postgres was therefore given the store seat's grants
and issued its subjects, and each runtime served the seat's verbs because it serves what it is issued
([ADR 0160](../../02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)).
The runtime behaved as designed. The fault was in what it was issued.
The seat's verbs are not the module's tools. The store's `databases` and `query` are a separate
implementation registered under the seat's name ([ADR 0159](../../02-DECISIONS/0159-a-tool-call-names-the-machine-and-a-holder-serves-its-seats-verbs.md)).
Only that implementation should be withdrawn where the module does not hold the seat. postgres's own
tools stay served on every machine it runs on.
## Fix
The controller now reads the recorded seat holdings when it composes grants and memberships. A seat
held once for the mesh is issued only to the machine and module the records name as its holder. A
seat held once per machine, and a mesh seat with no holder on record, are issued as before. Grants
and memberships come from the same list, so they cannot disagree.
**How it is checked.** A controller test asserts that a claimant on another machine keeps its node
seats and loses the recorded mesh seat. Live, the discovery console's overview must show each
mesh-scoped seat announced from exactly the holder the records name. Status moves to `resolved` once
that holds after the fix is rolled out.
## A second cause, and what the rollout broke (2026-10-04)
With the grants corrected, calls to the store reached only the holder, yet the console still showed
the seat announced from both machines. The module's runtime added every seat its start-up credential
claims, even after the mesh issued a membership that left the seat out. Once a membership exists,
it now alone decides which seat verbs a runtime serves (mesh-tools#45).
The rollout then exposed a third fault. The module on the machine that does not hold the seat was
still running an image built before #45, so it subscribed to the seat's subject. The corrected grants
refused that subscription, and the refusal ended the process. Its runtime crash-looped until the
module was rebuilt on the new runtime image. The database itself kept running. A refused tool
subscription is now logged and costs only that subject (mesh-tools#46), as a refused announcement
already did ([issue 217](../217-a-refused-announcement-took-down-every-containers-runtime/00-report.md)).
The module was not rebuilt by the plan that rebuilt the runtime image. This is the ordering gap of
[issue 211](../211-a-bundle-is-built-before-the-toolchain-it-is-compiled-in/00-report.md) seen
from a container module.
**Proven 2026-10-04.** The discovery console's overview shows the store seat announced from the
recorded holder only. Repeated calls to the store are answered by that machine, and the answers
include the controller's own database. The non-holder still answers its own module tools. No
container restarts on any machine.
@@ -0,0 +1,48 @@
---
status: resolved
opened: 2026-10-04
located-in:
- mesh-controller
fixed-by:
- mesh-controller#249
amended-design:
---
# 219 — An older build that finishes later replaces a newer one
## What was observed
2026-10-03. Two merges to the runtime module came minutes apart. Each plan asked for every module
built on the runtime image to be rebuilt. One module's two builds were of the same source and
differed only in the runtime image they stood on:
| Build | Requested | Finished | Stood on |
|---|---|---|---|
| asked by the first plan | 21:33 | 22:04 | the runtime image before the fix |
| asked by the second plan | 21:49 | 21:56 | the runtime image with the fix |
The older request finished last, and its image became the module's current artifact. The next push
deployed it, and the module's runtime crash-looped on a fault the newer image had already fixed
([issue 218](../218-a-mesh-seat-is-answered-by-a-module-that-does-not-hold-it/00-report.md)).
## Why it matters beyond this instance
Which build is current should follow what it was built from, not which build machine was slowest.
Whenever two plans overlap, which happens on any busy evening, a fix can be silently reverted by a
build that started before it existed. Every passing check still passes.
## Where to look
How a finished build is recorded and how a module's current artifact is chosen. **How it is checked:**
a test in which an older request completes after a newer one for the same artifact, and the newer
stays current.
## Resolved
A build is ordered by when it was requested, read from the id the controller gives it, and a
module's registered manifest is replaced only by a build requested at or after the one it came from.
An older request finishing later is recorded and changes nothing; a plan settles only from builds it
asked for itself. **How it is checked:** store-backed tests replay the incident — the newer request
stays what the module is — and fail without the fix. Live since 2026-10-04: the rebuilds of every
runtime-image module and the three waves of module code moves since then each registered the build
they asked for.
@@ -0,0 +1,41 @@
---
status: resolved
opened: 2026-10-04
located-in:
- mesh-host
fixed-by:
- mesh-host#84
amended-design:
---
# 220 — A delivered bundle keeps the files of the one before
## What was observed
2026-10-04. A tools bundle was rebuilt as one self-contained file per entrypoint
([ADR 0193](../../02-DECISIONS/0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md)),
so it no longer carries a package directory. On the machine it was delivered to, its directory still
held the package directory and a compiled file from the earlier delivery, dated hours before the
new files. The new files were written over the old directory, and nothing removed what the new
bundle no longer contains.
## Why it matters beyond this instance
A bundle on disk should be exactly the artifact that was built. Leftover files can be imported by
code that should no longer find them. A fix that removes a file then works on a fresh machine and
fails on every machine that ran an earlier version. It also makes "what runs here" impossible to
read from the artifact.
## Where to look
How the host unpacks a bundle into its directory. **How it is checked:** deliver a bundle, then a
version without one of its files, and the file is gone.
## Resolved
An archive is unpacked into a fresh directory beside the old one and swapped in by rename; a refused
or failed unpack leaves the old tree whole. **How it is checked:** a second delivery without a file
removes it, and nothing is left beside the directory; both tests fail without the fix. Proven
2026-10-04: a bundle rebuilt and delivered after the fix holds exactly the new build and nothing
beside it. A bundle that has not changed since keeps its old leftovers until its next version, by
design: an unchanged archive is not unpacked again.
@@ -0,0 +1,52 @@
---
status: resolved
opened: 2026-10-04
located-in:
- mesh-controller
fixed-by:
- policy: `upgrade build-agent roll-out` (2026-10-04)
amended-design:
---
# 221 — A build machine learns a new builder only from a push
## What was observed
2026-10-04. A controller merge changed how TypeScript bundles are built: one file per entrypoint
instead of a package directory. Its plan finished, and six bundles were rebuilt right after. They
came out in the old shape, because the build machines still ran the previous builder. They got the
new one only from the next push. Rebuilt after that push, the same six came out right.
The same order showed in the bus grants the same night. A push sent while the controller was still
the previous build composed grants with the previous code. A second push was needed after the new
controller had started.
## Why it matters beyond this instance
A merge to the controller is not in effect when its plan says done. Anything built or pushed in the
gap uses the old code and looks current. Today only the operator knows to push first and build
second, and even the operator forgot.
## Where to look
Whether a controller plan should end by delivering itself to the build machines and the control
machine, or whether a build should refuse a builder older than the controller that asked for it.
**How it is checked:** after a controller merge, a build asked right after its plan finishes runs
the new builder.
## Located
The mechanism existed: a module whose upgrade policy is `roll-out` is sent to its machines when its
tier is built, and the plan's next tier waits until it is applied. The build agent's policy was
`record` — built, never sent — as was that of 76 other modules, which is why every rollout on
2026-10-03 needed a push by hand. The build agent was set to `roll-out`, one machine at a time,
stopping at the first failure. **How it is checked:** the next controller merge's plan sends the
build agent to the build machines before its last tier, and a build asked right after the plan
finishes runs the new builder; then this moves to resolved. Whether the other modules roll out is
the operator's policy, not this issue's.
## Resolved
Proven 2026-10-04: the next controller merge's plan logged that its first tier was built and the
build agent sent to all four build machines, and only then asked its next tier. The builder change
in that merge reached the build machines without a hand push.
@@ -0,0 +1,51 @@
---
status: resolved
opened: 2026-10-04
located-in:
- mesh-tools
fixed-by:
- mesh-tools#47
amended-design:
---
# 222 — An assignment is refused on the bus until the bus's machine is pushed
## What was observed
2026-10-04. A module was assigned to the laptop, and the laptop alone was pushed. The module's two
bundles arrived and the node's runtime launched both, and then the bus refused every one of the
module's subjects:
```
nats: permissions violation: Permissions Violation for Subscription to "mesh.mod.<module>.tool.<tool>.<node>"
```
The tools were unreachable until a later push that included the machine running the bus. Then the
runtime served them without a restart, because a new membership arrived and it re-subscribed.
## Why it matters beyond this instance
What an account may answer lives in the bus's user list, and the controller writes that list only
into the declaration of the machine that runs the bus. Assigning anything to any machine changes
that list, so `push <machine>` after `assign <machine> <module>`, which is what the controller itself
tells the operator to run, leaves the module running and unreachable, with nothing reporting a fault.
## Where to look
Whether a push to one machine should also send the bus's machine when the user list it would
compose differs from the one that machine holds. **How it is checked:** assign a module with tools
to a machine that does not run the bus, push only that machine, and its tools answer.
## Diagnosed and resolved
The push did send the bus's machine: the controller's log shows both machines applying in the same
second. The fault was the order within that second. The node's runtime subscribed before the bus
had reloaded its user list, the bus refused, and the bus client marks a refused subscription dead.
Nothing asked again until a later membership happened to re-serve the module.
A subject the runtime answers on is now asked for again when the bus refuses it, after waits from
two seconds to two minutes, and given up and said after about five minutes. **How it is checked:**
a test refuses a subject and finds it asked for again and answering, given up past its attempts,
and not asked again once stopped; and by hand against a bus whose permissions were reloaded while
connected, the runtime answered two seconds after the grant arrived, where the runtime before the
fix never answered.
@@ -0,0 +1,60 @@
---
status: resolved
opened: 2026-10-04
located-in:
- mesh-host
fixed-by:
- mesh-host#87
amended-design:
---
# 223 — A new mesh installs its controller as a container
## What was observed
2026-10-04. Preparing [issue 213](../213-the-controller-is-a-go-program-run-in-a-container/00-report.md),
the controller's own manifest was changed to a Go bundle run as a process. The installer that raises a
new mesh assumes the controller is an image and a container at every step from its third on:
- it requires the controller's build to produce exactly one image;
- it starts a temporary controller from that image, and publishes the image to the registry;
- at the pivot, it finds the controller's container in the declaration, reads its environment and
volumes, waits for it, and from then on talks to the controller only through the container.
## Why it matters beyond this instance
With the controller's manifest changed, a new mesh cannot be installed: the pivot fails. The deeper
constraint is ordering. A process's bundle is fetched from the artifact store, and the installer
raises the artifact store only after the pivot, so the controller's first declaration names a bundle
nothing can serve yet.
## What a fix has to settle
One of two shapes, and it is a decision, not a repair:
1. raise the artifact store before the pivot, publish the controller's bundle to it, and talk to the
controller from the host's side rather than through a container; or
2. pivot to the image form as today, and let the first push hand over to the process, which
requires the controller's manifest to carry both forms.
Until it is settled, the change of the controller's manifest (mesh-controller#253) is held. The
handover itself is built and merged (mesh-host#86); the controller's half (mesh-controller#252) waits
on the operator. **How it is checked:** the installer's test raises a mesh whose controller manifest
is the process form, and the controller answers its seat's verbs at the end.
## Decided (2026-10-04)
Option 2, [ADR 0200](../../02-DECISIONS/0200-genesis-pivots-to-the-controller-as-a-container-and-the-first-push-hands-it-to-a-process.md):
genesis pivots to the controller as a container recorded under the name the process `replaces`, and
the first push hands it over.
## Resolved
Built as [ADR 0200](../../02-DECISIONS/0200-genesis-pivots-to-the-controller-as-a-container-and-the-first-push-hands-it-to-a-process.md)
decided: genesis reads only the controller's process from the manifest, builds the controller's image
from its repository, and pivots to a container of its own shape recorded under the id the process
`replaces`; an older controller in the image form still builds as before. **How it is checked:** the
installer's tests build the genesis form from the controller's real manifest, and apply the
controller's first process declaration over the recorded container with the host's own apply — the
process starts, the container is removed, one controller remains. A real install from nothing has not
been run since; the handover it ends in was proven live on the running mesh (issue 213).
@@ -0,0 +1,64 @@
---
status: open
opened: 2026-10-04
located-in: [mesh-host internal/apply/apply.go, mesh-host internal/apply/schedule.go]
fixed-by:
amended-design:
---
# 224 — An apply arriving during a maintenance window reopens it, by recreating the container the window is holding still
## What was observed
Reviewing [ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md)'s
`while-stopped` before merging it, 2026-10-04. Found by reading, not by running.
A scheduled step may hold its module's containers still while it runs. The host stops them, runs
the step, starts them again. Nothing tells the **apply** that a window is open, and the apply's
rule for a container it finds stopped is to replace it:
```
case existed && (before.Spec == want || legacy) && before.Running && len(reasons) == 0:
out.Action = "unchanged"
case existed:
rm -f
```
`before.Running` is false for a container a window is holding, so the second branch takes it:
the container is removed and recreated, **running**, in the middle of the step that required it
to be still.
## Why it matters
For the store, which is what the field was built for, the chain is: a push lands at 03:30 → the
apply recreates the registry → the registry accepts an upload from a build running at the same
time → `garbage-collect`, already past its mark phase, sweeps the blob that upload just wrote.
The image is then in the store with a layer missing, and the build that made it reported success.
Two things have to coincide, so it is not likely. It is also not rare enough to leave unsaid: the
mesh pushes on every merge, at any hour, and a collection over a store this size is minutes rather
than seconds.
**The general shape is the one that matters.** `while-stopped` is the first thing in the mesh that
makes a container's stopped state *intentional*. Everything else in the host reads "stopped" as
"broken, fix it", which is right everywhere else and wrong here. Any future use of the field
inherits this.
## What this is not
Not a regression. The store has never collected anything, so nothing is worse than it was; this
is a hole in something new rather than something that broke.
## Open questions
- **Should a window take the apply lock?** The daemon already serialises applies with `applying`
and, across processes, with `store.Lock`. A window that held it would make the race impossible.
The cost is that a push arriving mid-window waits for minutes, and a push that waits is what
[issue 185](../185-a-refused-membership-publish-stops-the-controller/00-report.md)'s
family of outages looked like from outside.
- **Or should the apply learn that a container is held?** Narrower: the scheduler says which
containers a window currently holds, and `applyContainer` reports those unchanged instead of
recreating them. Nothing blocks, and the apply tells the truth for the minutes it matters —
at the cost of a second source for "is this container meant to be running".
- Either way: should the *report* say a window is open, so a machine that looks half-stopped at
03:31 reads as working rather than broken?