Compare commits

..
Author SHA1 Message Date
jschoubben bc64c5c187 ADR 0201 → 0202, and three issues from the night it shipped
The derived-value record is renumbered a second time: the key-value-buckets
record took 0201 while this waited to merge, as the bundles record took 0188
before it. Both times free when chosen, taken by the time it landed. cycle.py
caught it; three repositories cite this record, so the number matters.

225 — a provisioner has not been able to read its grant secrets since 01:30,
when a module's own code left its container and the files stayed root's. Four
thousand refusals, each worded as patience, and two consumers unserved. Not
from ADR 0202 or 0189, which landed hours later; dates in the report.

226 — the store's sweep stops at the first reference recorded with an address
and collects nothing. A guard that cannot tell 'I will not ask about this'
from 'it would not answer' stops the wrong amount of work.

227 — the photo app's admin client asks for the port the proxy holds. A module
pinned months behind carries everything its branch gained, the first time
anything makes it move.
2026-10-04 04:33:45 +02:00
mesh-admin be4b5777b8 Merge pull request 'Research 024 and ADR 0201: a module keeps its current state in key-value buckets' (#348) from feat/module-state-on-the-bus into main 2026-10-04 01:43:34 +00:00
mesh-admin d882b3568c Merge pull request 'Group 8: ADR 0201 (a provider declares what it derives, issue 124) and ADR 0189 (the store keeps what the records name, issue 108); issue 202' (#305) from feat/the-store-keeps-what-the-records-name into main 2026-10-04 01:30:47 +00:00
jschoubben a3523617d3 Review before merge: the multi-holder boundary, the window's open race, the sweep's bounds
ADR 0201 gains the boundary found reading it back: a consumer keeping several
holders of a deriving provider is refused, because the two ends have no way
to agree. ADR 0189 gains two consequences — the sweep is bounded because it
runs inside a build, and an apply arriving mid-window reopens it.

That last one is issue 224, recorded rather than fixed: the host's rule for a
stopped container is to replace it, and while-stopped is the first thing that
makes a stopped container intentional. Both candidate fixes are decisions with
their own cost. Nothing is worse than it was; the store has never collected.
2026-10-04 03:27:32 +02:00
jochen 6b4da63261 Research 024: the composed grants, checked against a server once built 2026-10-04 02:50:02 +02:00
jschoubben 0231974226 Rebased onto main: ADR 0188 renumbered to 0201, and issue 202's evidence re-taken
The bundles refactor took 0188 on main while this waited in a pull request,
and the mesh's own code cites that one, so this record moves. Only the number
moved; the decision is the one taken on 2026-10-02, and the record says so.

Issue 202 re-checked against the refactored main: the fault stands, and the
test that surfaced it now fails one step earlier on issue 203's new credential
guard. Proven again past both — mint the credential, compose twice, and all
eight of dnsmasq's resources appear only with the setting set. ADR 0164 is
noted as the decision that answers half of it, and is not built.
2026-10-04 02:44:58 +02:00
jschoubben 92c029d10e ADR 0189: the store keeps what the records name, and a maintenance step holds its writers still
Issue 108: the artifact store has never collected anything. Fifty-three
repositories on the machine that serves everything else, and the only outcome
of leaving it is a full disk reported as somebody else's failure.

The mesh decides what may go — from its own build records, so it never names
a digest it did not put there — and the store reclaims the bytes in a nightly
window with its server held still. Deletion on the one door takes nothing a
push did not already have.

Designs 18 and 20 amended; issue 108 resolved.

Also issue 202, found running the controller's suite: a module whose required
setting nobody set is left out of the machine in silence, and dnsmasq became
that module this morning.
2026-10-04 02:40:25 +02:00
jschoubben 2a60da821d ADR 0188: a provider declares what it derives for each consumer, and the mesh tells both ends
Issue 124: a value the mesh's own rule produced reached neither end as a
statement. The object store's provisioner derived each consumer's bucket in
its own code; all three consumers transcribed the rule into their own
definitions, one of them wrong, and each of the three also named the machine
it happens to run on.

A served value may now name the consumer the mesh is serving. Design 27
amended; issue 124 resolved.
2026-10-04 02:40:06 +02:00
jochen e1b0bbde91 Research 024 and ADR 0201: a module keeps its current state in key-value buckets
Events miss a machine that joins after them and replay history where only the
latest matters. A module now declares state it owns and reads; the controller
creates the buckets, the runtime serves them on the bundle's channel. Designs 32
and 25 amended; grants measured against a running server.
2026-10-04 02:36:59 +02:00
mesh-admin 8ca09c70d6 Merge pull request 'Issues 213 and 223 resolved' (#347) from issues/213-223-resolved into main 2026-10-04 00:19:22 +00:00
jochen db71b83711 Issues 213 and 223 resolved: the controller runs as a process, and genesis hands over to it 2026-10-04 02:19:09 +02:00
mesh-admin 9143d0b7c1 Merge pull request 'ADR 0200: genesis pivots to the controller as a container, and the first push hands it to a process' (#346) from decision/0200-genesis-pivots-to-a-container-and-hands-over into main 2026-10-03 23:41:36 +00:00
jochen 24a51a8e53 ADR 0200: genesis pivots to the controller as a container, and the first push hands it to a process 2026-10-04 01:41:30 +02:00
mesh-admin f0d7f91d90 Merge pull request 'Design 38: WP4c complete' (#345) from design/38-wp4c-complete into main 2026-10-03 23:35:31 +00:00
jochen 8578a06ca8 Design 38: WP4c complete, no module's own code runs in a container 2026-10-04 01:35:16 +02:00
mesh-admin 86083f9c9d Merge pull request 'Issues 215, 221, 222 resolved' (#344) from issues/215-221-222-resolved into main 2026-10-03 23:29:38 +00:00
jochen 63d328147e Issues 215, 221 resolved with live proof; 222 diagnosed and resolved 2026-10-04 01:29:32 +02:00
mesh-admin fead0ea440 Merge pull request 'Issues 219-223 and design 38 WP4c built' (#343) from issues/219-223-and-wp4c into main 2026-10-03 23:14:49 +00:00
jochen df503d1cff Issues 219, 220 resolved, 221 located, 222 and 223 opened; WP4c built and proven 2026-10-04 01:14:44 +02:00
mesh-admin c2fc829822 Merge pull request 'Issues 211, 212, 214, 216, 217 resolved; 215's fix recorded' (#342) from issues/211-217-resolved into main 2026-10-03 22:25:07 +00:00
jochen 8d8e5c9a7e Issues 211, 212, 214, 216, 217 resolved with their proofs; 215's fix recorded 2026-10-04 00:24:53 +02:00
mesh-admin affba60b79 Merge pull request 'Issues 219-221: the builder's ordering gaps' (#341) from issues/219-221-the-builder into main 2026-10-03 22:16:13 +00:00
jochen 9ddbc4c68e Issues 219-221: the builder's ordering gaps seen while rolling out issue 218 2026-10-04 00:16:08 +02:00
mesh-admin a972db91f0 Merge pull request 'Issue 218 resolved: the runtime follows its membership, and a refused subscription is not fatal' (#340) from issues/218-rollout into main 2026-10-03 22:14:10 +00:00
jochen 029698fdc8 Issue 218 resolved: the runtime follows its membership, and a refused subscription is not fatal 2026-10-04 00:14:04 +02:00
mesh-admin 895c2afad1 Merge pull request 'Issue 218: a mesh seat answered by a non-holder (located, fixed in mesh-controller#248)' (#339) from issues/218-a-mesh-seat-answered-by-a-non-holder into main 2026-10-03 21:32:24 +00:00
mesh-admin d362155401 Merge pull request 'Issue 218: a mesh seat is answered by a module on a machine that does not hold it' (#338) from issues/218-a-mesh-seat-answered-by-a-non-holder into main 2026-10-03 21:25:27 +00:00
32 changed files with 1553 additions and 27 deletions
@@ -0,0 +1,152 @@
---
status: graduated
initiated: 2026-10-04
touches: [the bus, what a module declares, the tool runtime, the SDK, the bus grants, 03-DESIGN/01-to-be/25-the-bus-on-nats.md, 03-DESIGN/01-to-be/32-what-a-module-declares.md]
became: [02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md, 03-DESIGN/01-to-be/32-what-a-module-declares.md, 03-DESIGN/01-to-be/25-the-bus-on-nats.md]
---
# 024 — State a module keeps on the bus
## What is investigated
A place on the bus where a module's own code keeps **current state** — not history — that every
machine sees, including a machine that joins after the state was written: put, get, delete, list and
watch, reached through the node's runtime the way a bundle already publishes, asks and subscribes
([ADR 0198](../../02-DECISIONS/0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)).
On NATS that is a key-value bucket. The questions are what a module declares, who creates the
bucket, what the grants are, what the runtime's verbs are, and what may never be stored.
## Why
The mesh carries two kinds of module traffic and a third is missing.
- **Events** land in the EVENTS stream: limits retention, seven days, ten thousand messages per
subject, a durable consumer per consuming module that replays what it missed. Never a secret
([design 32](../../03-DESIGN/01-to-be/32-what-a-module-declares.md) §10).
- **Requests** are core request/reply — tool calls, a bundle's `mesh/ask` — and are kept nowhere.
Neither is *the current value of something*. Two cases from the first module that needs it, the
operator's agent on a machine ([design 36](../../03-DESIGN/01-to-be/36-the-operators-agent-on-a-machine.md)):
1. **An MCP server registered for every machine.** Registering emits an event every machine's copy
of the module consumes. A machine the module is assigned to *after* the registration has no
durable consumer yet — the consumer is created at assignment — so it never hears of it. Wanted
instead: one entry per server, for every machine or for one; every machine reads the whole current
set when it starts and watches for changes; unregistering is a delete; any machine can list it.
2. **Which licence a machine is bound to** ([design 39](../../03-DESIGN/01-to-be/39-the-anthropic-licence-manager.md)).
As events, a machine that was off for a day replays every rotation since and asks for a token
after each. It needs only the latest binding and its generation. The token itself stays on
request/reply and is never stored.
The design already expects this. [Design 25](../../03-DESIGN/01-to-be/25-the-bus-on-nats.md) §1:
"conditions and observed state in key-value buckets that anything may watch".
[Research 017](../017-a-mesh-that-heals-itself/01-the-intended-behaviour.md) wants a provisioner's
"what I applied" and a rotation's step kept in one rather than in memory. Nothing implements it.
## What exists, measured 2026-10-04
| | fact | where |
|---|---|---|
| streams | five kinds of mesh stream: CONTROL (work queue), NODES and ASSIGNMENTS (last per subject), EVENTS (limits: 7 days, 10 000 per subject), one work queue per seat that accepts | the controller's broker streams |
| the state relationship | [design 32](../../03-DESIGN/01-to-be/32-what-a-module-declares.md) §4 already names *state* — 1:1, last per subject — and says it is "declared: the mesh's own". Two streams use it, both written by the controller. No module can declare it | design 32, the controller |
| key-value buckets | none, anywhere | all four code repositories |
| the runtime's bus verbs | `mesh/publish`, `mesh/ask`, `mesh/subscribe`; delivery back to the bundle is `mesh/event` | the runtime's launcher |
| the runtime's principal | one bus user per machine carries every assigned module; its grant is the union of theirs. That one module's code does not act as another is the runtime's to keep: it publishes under the module's own name by construction | the controller's grant composition, the runtime's bus |
| what a bundle is issued | a membership per assignment, last per subject, read directly by the runtime: where it serves, where it emits, what it reaches | ADR 0160 |
| who creates bus objects | the controller only — mesh streams on every raise, a seat's stream at registration, a module's consumer at assignment. No module reaches the JetStream API | design 25 §3 |
### What a key-value bucket needs from a grant, against a real server
Measured against nats-server 2.10 with the Go client the runtime already uses, a bucket created by
an unrestricted user and used by two users holding only the subjects below (`B` is the bucket):
| operation | subject published | writer | reader |
|---|---|---|---|
| bind to the bucket | `$JS.API.STREAM.INFO.KV_B` | yes | yes |
| get | `$JS.API.DIRECT.GET.KV_B.>` | yes | yes |
| put, delete | `$KV.B.>` | yes | **refused** |
| list keys, watch | `$JS.API.CONSUMER.CREATE.KV_B.>` — an ordered, ephemeral consumer | yes | yes |
| stop a watch cleanly | `$JS.API.CONSUMER.DELETE.KV_B.>` | yes | yes |
| answers | its own inbox, which every principal already subscribes | — | — |
*Checked again once built, 2026-10-04:* the grants the controller composes for two machines' runtimes —
one carrying the owner, one only a reader — were loaded into a server as composed, and each operation
was run as each runtime's user. The owner's did all of them; the reader's read, listed and watched,
and its put and delete were refused by the server.
Three things the measurement showed that reading the documentation would not have:
1. **A refused put is not an error to the caller; it is a timeout.** The server reports the
permission violation asynchronously, on the connection, and the client waits out its deadline
for an acknowledgement that never comes. So a runtime that relies on the grant alone tells a
bundle "timed out" for "you may not write this" — it must refuse first, from what the module was
issued, with the reason.
2. **A watch's current values include deletions.** A key deleted earlier arrives among the initial
values as a delete marker, before the end-of-current marker. A bundle asking "what is there now"
must not be handed those.
3. **Without the consumer-delete grant, stopping a watch hangs** until its deadline, and the
ephemeral consumer lingers on the server until it times out by itself.
### Whether the events shape is enough instead
Honestly compared, because a new primitive is a cost:
- **EVENTS cannot be made last-per-subject for some subjects.** Retention is per stream, and
JetStream refuses a second stream overlapping the first (verified and recorded in design 32 §3).
A state subject inside `mesh.mod.*.event.>` keeps EVENTS' seven days: a licence binding unchanged
for a week disappears.
- **A separate last-per-subject stream per module** is possible — it is exactly what a key-value
bucket *is* on the server: a stream with one message per subject, a rollup for purge, and direct
reads. Building it by hand gives up the client's get, list, delete and watch, which are the
operations both cases need, and would be the mesh writing NATS's own key-value layer again.
- **Consumers are the wrong reader.** A durable consumer per reading module is created at
assignment and replays from where it is; state wants "everything current, now, then changes",
which an ordered ephemeral consumer from the last value per subject gives and a durable does not.
So key-value is not a convenience over events; it is the state relationship design 32 already
names, opened to modules.
## Questions, and what this effort proposes
1. **What a manifest says.** `state` names the buckets a module owns, by local name — every
instance of the module may write them and read them. `reads` names another module's bucket as
`<module>.<name>`, read-only. Names only, never a bucket or subject (design 32 §1). A bucket's
options — how many past values it keeps, how long a value lives — are the owner's to declare,
the way a seat declares its own retention (design 32 §3).
2. **Scope.** One bucket per module per name, mesh-wide. A key may carry a machine by the module's
own convention (`all.<server>`, `<machine>.<server>`). A bucket per machine was considered and
not proposed: "list every server for every machine" becomes a walk over buckets, and the grant
could only narrow writes, which nothing asked for — every instance of the owner already writes.
3. **Who creates the bucket.** The controller, from the catalogue, on every raise — a bucket exists
from registration, like a seat's stream, so a reader can watch before the owner is assigned
anywhere. Never a module.
4. **The runtime's verbs.** `mesh/state.get`, `mesh/state.put`, `mesh/state.delete`,
`mesh/state.keys`, `mesh/state.watch`, each naming the bucket as the module named it. A watch
is answered once the current values are on their way, then each change is delivered to the
bundle as a `mesh/state` request it answers — current values first (no deletions among them), an
end-of-current marker, then changes. A child that restarts watches again, as it subscribes
again. The runtime refuses, with the reason, a bucket the module was not issued, and a write to
one it only reads.
5. **Secrets.** None in a bucket, sealed or not: a bucket is a stream (design 32 §10). Sealed values
are plain base64 and cannot be recognised, so the mechanical check is partial and said to be: the
runtime refuses a value carrying a field whose name says it is a credential (`password`,
`secret`, `token`, `authorization`, …), which catches the ordinary mistake and not a determined
one. For the first consumer this has a concrete consequence: an MCP server registered with an
authorisation header keeps that header out of the bucket.
6. **History, lifetime, size.** One value per key unless the owner says more; no expiry unless it
says one; a value at most 256 KiB and a bucket at most 64 MiB, the mesh's caps rather than a
module's. **A bucket outlives its module's assignment** — what a module stored is data, and data
outlives what declared it ([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md));
unassigning is not cleaning up. A bucket whose declaration is gone is reported, never removed.
7. **Events or state.** State (above).
## The work, once decided
1. A decision record, then design 32 (*state* becomes a relationship a module declares) and design
25 (key-value buckets are part of the bus) amended.
2. The controller: the manifest's two words and their registration check; buckets asserted on every
raise; the grants for owners' and readers' runtimes; the buckets issued in each membership.
3. The runtime: the five verbs, the watch delivery, the refusals; tested against a real server.
4. The SDK, TypeScript and Go: a small state surface over the verbs.
5. Proved on a running mesh with one small module, then handed to the operator's agent, whose
registered servers move from events to a bucket.
@@ -0,0 +1,141 @@
---
topic: the mesh
status: accepted
date: 2026-10-02
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md
---
# 189. The store keeps what the records name, and a maintenance step holds its writers still
## Context
The mesh's artifact store has never collected anything
([issue 108](../04-ISSUES/108-the-registry-has-no-garbage-collection-once-it-has-two-doors/00-report.md)).
Every build pushes another layer set; nothing has ever removed one. The predecessor ran a routine
on a timer — stop the registry, collect, start it — and the conversion carried the settings that
routine depends on without the routine, because the routine was a script beside the module and not
a resource in it. The store now holds fifty-three repositories on the machine that serves
everything else, and the only outcome of leaving it is a full disk reported as somebody else's
failure.
Three things stood in the way, and the issue names all three.
**Nothing in the mesh's vocabulary expresses a maintenance window.** The collector requires every
writer stopped while it runs. A `run-once` step runs *beside* containers, not instead of them, and
a scheduled step is the same container on a cadence. There is no way for a module to say *hold this
container of mine still while this runs*.
**Deletion is not enabled, and the door it would be enabled on has no accounts.** The store is
internal, reached by name over the overlay, trusted because being on that network is the permission
([ADR 0082](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)). The predecessor
kept deletion behind an authenticated door, which it could, having one.
**Nothing says what may be removed.** The registry's own answer — collect everything no tag names —
is wrong here. The mesh pushes each artifact under one moving tag and pins machines by digest, so
every build but the newest is untagged and some machine may still be running it.
## Decision
**1. Deletion is enabled on the store's one door, and the overlay stays the permission.** The
objection dissolves on inspection: that door **already accepts a push**, and a writer who can push
can replace any tag in the store with anything it likes. Delete takes nothing a push did not
already have, and the machines that can reach the door are the ones the mesh's own filter admits
([ADR 0168](0168-a-converged-machine-is-filtered-by-the-mesh-alone.md)). Putting an authenticated
door in front of deletion while leaving push open would be a lock on the window beside an open
door, and it would cost the thing ADR 0082 bought: a store every machine can reach without a
credential to distribute first.
**2. The mesh deletes what it made and no longer keeps; the store reclaims the bytes.** Two halves,
each doing what only it can.
The **mesh** decides. It does not need to enumerate the store to do it — it has never put anything
there it did not record, so **every digest it could remove is already in its own build records**.
It deletes those manifests through the store's door, by digest, and remembers that it did.
The **store** reclaims. A deleted manifest frees no bytes until the registry's own collector walks
the storage with nothing writing to it, so the module declares that collector as a scheduled step
with the server held still for its duration. Plain collection, not `--delete-untagged`: what the
mesh keeps is still a manifest in the store, so it is still referenced, so its blobs stay — the
dangerous flag is not needed at all once the mesh is the one deciding.
**3. What the mesh keeps, stated as three reasons rather than a number.** A digest is kept because:
- **a definition names it** — every artifact reference in any module's current recorded manifest,
which is what the mesh would hand a machine now. No age limit: this is the floor;
- **the mesh can still go back to it** — every artifact of the **five most recent successful
builds** of each module, so a release that turns out wrong has somewhere to return to;
- **nothing else.** An artifact older than that, which no definition names, is what the store is
carrying for no stated reason.
A digest the mesh did not record making is never touched. That is not a safety margin, it is the
whole rule restated: the mesh removes what it put there and can account for, and the images genesis
pushed before any record existed are exactly what this must not reach
([04-ISSUES/102](../04-ISSUES/102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md), F4).
**4. A scheduled step may hold its module's own containers still while it runs** —
`while-stopped`, naming resource ids in the same module. The host stops each, runs the step, and
starts them again **whatever the step did**, including when it failed or the host was interrupted.
Three boundaries:
- **Its own module's containers only.** A module that could quiesce a neighbour could stop the
mesh; a maintenance window is a statement about one service's own insides.
- **Scheduled steps only, not `run-once`.** At apply time the host already has a window: the
declaration is applied in order and a step gates what follows, so a one-time offline migration
says *before* rather than *instead of*. A recurring window is the case order cannot express.
- **Restoring is not conditional.** A step that fails must leave the service running; the whole
risk of this field is a window that never closes.
**5. The sweep runs where the records change — after a build the mesh recorded.** That is the
moment new bytes landed and the moment the keep set moved, and it needs no new timer. The
store's collection runs nightly, because reclaiming is slow and the thing it reclaims is already
unreferenced.
## Consequences
- Disk stops growing without bound on the machine that serves the mesh. That is the whole point
and it has no other way to be true.
- A machine behind by more than five builds of a module, which recreates a container, cannot pull
what it was running. It is already a machine the mesh reports as behind, and the answer is the
one the mesh already gives it: the current declaration. Stated here rather than discovered.
- The store is a little less of a museum. A digest in an old build record may no longer be
fetchable, and the record still says what that build made — the record is history, not an
index of what is on disk. The collected mark is kept beside it so the two can be told apart.
- `while-stopped` is a second thing the host does to a container it did not start this pass. It is
deliberately the narrowest form: the module's own, by id, restored unconditionally.
- The store is briefly unavailable each night, for as long as collection takes. Everything that
pulls from it retries; nothing in the mesh treats a momentary store as a failure
([ADR 0185](0185-a-control-plane-behind-its-seats-row-serves-what-it-can.md)).
- **An apply arriving during the window reopens it**, because the host's rule for a container it
finds stopped is to replace it, and `while-stopped` is the first thing that makes a stopped
container intentional. Found by reading this before it merged, recorded as
[issue 224](../04-ISSUES/224-an-apply-reopens-a-maintenance-window-by-recreating-what-it-held-still/00-report.md)
rather than fixed here: the two candidate fixes — the window takes the apply lock, or the apply
learns which containers are held — are each a decision with its own cost, and neither belongs
inside this record. Nothing is worse than it was; the store has never collected at all.
- **The sweep is bounded**: at most two hundred artifacts and sixty seconds per build, stopping at
the first refusal, because it runs inside somebody's build. What is left over is offered again
next time. The store stops growing from the first sweep; it does not empty in one.
## How this is checked
- The host: a scheduled step with `while-stopped` stops the named containers before the run and
starts them after; it starts them again **when the step fails**; it refuses an id that is not a
container of the same module, its own id, and `while-stopped` on a `run-once` step. Each refusal
is tested for what it says, not only that it says something.
- The controller: given build records and current manifests, the keep set holds every reference a
manifest names and every reference of the five most recent builds per module, and nothing else;
a reference the mesh never recorded is never in the delete set; a delete that answers 404 is
recorded as collected rather than retried forever.
- The sweep is tested against a fake store that records what it was asked to delete, so what is
asserted is the decision and not the registry's behaviour.
- Live: the store's size before and after the first nightly collection, read from the machine.
## References
- [issue 108 — the registry has no garbage collection](../04-ISSUES/108-the-registry-has-no-garbage-collection-once-it-has-two-doors/00-report.md)
- [ADR 0082 — the registry is reached by name and trusted by the overlay](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)
- [ADR 0053 — a step that runs on a schedule](0053-a-step-that-runs-on-a-schedule.md)
- [ADR 0156 — an artifact is what a build produces, and the store is named for its scope](0156-an-artifact-is-what-a-build-produces-and-the-store-is-named-for-its-scope.md)
- [design 32 — what a module declares](../03-DESIGN/01-to-be/32-what-a-module-declares.md)
@@ -0,0 +1,78 @@
---
topic: building it
status: accepted
date: 2026-10-04
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0067-genesis-is-a-pivot.md
---
# 200. Genesis pivots to the controller as a container, and the first push hands it to a process
## Context
The controller is Go, compiled to one static binary, and is the last of the mesh's own programs a
machine runs from an image ([issue 213](../04-ISSUES/213-the-controller-is-a-go-program-run-in-a-container/00-report.md)).
[ADR 0188](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md)
§1 says a module's own code is bundles, never an image, and §3 that a service bundle is a `process` the
host runs. The handover exists: a process may name the container it `replaces`, and the host removes
that container only after the process has stayed up across two checks; two controllers are safe
together for that moment, the second standing by on the controller's consumers and every plan held by
one lock.
What stands in the way is genesis ([ADR 0067](0067-genesis-is-a-pivot.md)), which
[issue 223](../04-ISSUES/223-a-new-mesh-installs-its-controller-as-a-container/00-report.md) found
assumes an image and a container at every step from its third: it builds the controller's image,
starts a temporary controller from it, publishes it, finds the controller's container in the pivot
declaration, and from then on talks to the controller through it. A process's bundle is fetched from
the artifact store, and genesis raises the artifact store only after the pivot.
## Considered Options
1. **Raise the artifact store before the pivot**, publish the controller's bundle to it, and talk to
the controller from the host's side. Rejected for now: it reorders genesis around a store that is
itself a module the controller deploys, and rewrites the steps that talk to the controller — a
larger change to the one path that is exercised least, to remove a container that exists for
minutes.
2. **Pivot to the controller as a container, as today, and let the first push hand it over to the
process**, through the handover that already exists. Chosen.
3. **Keep the controller a container.** Rejected: it is the exception to ADR 0188 that every other
module's code has now left, and it costs a container runtime on the control machine and a
container recreation in the middle of a plan.
## Decision
**Genesis raises the controller as a container, under the resource the controller's process
`replaces`, and the first declaration the controller composes for its own machine hands it over.**
The container is genesis's own shape, built from the controller's repository, and is recorded on the
control machine exactly as the manifest's `replaces` names it, so the first apply after the pivot
finds a replacement for it and removes it once the process is up. The controller's manifest declares
only the process; the image form exists for genesis alone and is not a second way to run the
controller on a live mesh.
This is the one bounded exception to ADR 0188 §1: a module's own code in an image, for the minutes
between the pivot and the first push, on a mesh being created.
## Consequences
- A new mesh ends where a running one is: the controller a process, no controller container.
- Genesis keeps its steps; what changes is that it no longer reads the controller's container from the
manifest, and that it records the container under the name the handover expects.
- The controller's repository keeps its image build for genesis and the lab.
- The handover is now on genesis's path too: a process that fails to stay up leaves the genesis
container serving, and the apply says so — the same rule as on a live mesh.
## How it is checked
The installer's test raises a mesh whose controller manifest is the process form, and asserts that the
container genesis recorded is exactly what the process `replaces`, so the first apply hands over and
leaves one controller. Live, on the running mesh: after the manifest change is pushed, the control
machine runs the controller as a process and no controller container, and the controller's seat
answers throughout.
## References
- [Issue 213](../04-ISSUES/213-the-controller-is-a-go-program-run-in-a-container/00-report.md),
[issue 223](../04-ISSUES/223-a-new-mesh-installs-its-controller-as-a-container/00-report.md)
- mesh-host#86 (the handover), mesh-controller#252 (two controllers safe together),
mesh-controller#253 (the controller's manifest as a process)
@@ -0,0 +1,106 @@
---
topic: what runs on it
status: accepted
date: 2026-10-04
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md
---
# 201. A module keeps its current state in key-value buckets it declares, and reaches them through the runtime
## Context
A module's code reaches the bus through the node's runtime: it publishes events, subscribes to them
and asks tools ([ADR 0198](0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)).
Events are kept for a week and replayed to a consumer that was away; requests are kept nowhere. What
neither gives is **the current value of something**, seen by every machine, including one that joins
after it was written. The first module to need it — the operator's agent on a machine — registers MCP
servers for every machine as events, and a machine assigned later never hears of them; and it would
replay a week of licence rotations where it needs only the binding that holds now. Research
[024](../01-RESEARCH/024-state-a-module-keeps-on-the-bus/00-overview.md) measured the alternatives and
the grants against a real server.
[Design 32](../03-DESIGN/01-to-be/32-what-a-module-declares.md) §4 already names *state* as one of the
mesh's relationships — 1:1, last per subject — and reserves it to the mesh's own declarations.
[Design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md) §1 expects key-value buckets on the bus.
## Considered Options
1. **Key-value buckets a module declares, created by the controller, reached through the runtime.**
Chosen.
2. **State as events on EVENTS, read last-per-subject.** Rejected: retention is per stream and EVENTS
keeps seven days, so a value unchanged for a week disappears; a second stream over the same subjects
is refused by the server (design 32 §3). And events give no get, list or delete.
3. **A last-per-subject stream per module, written by hand.** Rejected: it is what a key-value bucket
is on the server, without the client's get, list, delete and watch — the mesh writing NATS's
key-value layer again.
4. **State in a module's own files or database, shared by asking a tool.** Rejected for state every
machine must see: a machine joining later has to know whom to ask and poll, and an owner that is
down answers nothing — the property the bus exists to remove.
## Decision
**1. A module declares its state by name.** `state` names the buckets it owns, by local name; every
instance of the module may write and read them. `reads` names another module's bucket as
`<module>.<name>`, read-only. A bucket's options are its owner's: how many past values a key keeps,
and how long a value lives. A manifest names no bucket, stream or subject (design 32 §1).
**2. One bucket per module per name, mesh-wide.** A key may name a machine by the module's own
convention; the mesh does not scope buckets per machine.
**3. The controller creates the buckets, from the catalogue, on every raise** — from registration,
like a seat's stream, so a reader can watch a bucket whose owner is not yet assigned anywhere. A module
never creates one. The runtime's grant on each bucket is the union of what its carried modules may do:
an owner's instances write and read, a reader's read.
**4. Each assignment is issued its buckets in its membership** ([ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)),
by the name the module uses for each and whether it may write. The runtime serves `mesh/state.get`,
`put`, `delete`, `keys` and `watch` on the bundle's channel from that list, and refuses — with the
reason — a bucket the module was not issued and a write to one it only reads. A watch delivers the
current values first, without deletions, then an end-of-current marker, then every change, each as a
`mesh/state` request the bundle answers.
**5. No secret is stored in a bucket, sealed or not.** A bucket is a stream, and design 32 §10 keeps
every secret off streams. A value that needs a secret names it; the secret travels on request/reply.
**6. The mesh caps size; a bucket outlives its module.** One value per key and no expiry unless the
owner says otherwise; at most 256 KiB a value and 64 MiB a bucket. Unassigning a module leaves its
buckets and what is in them ([ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md)); a bucket
whose declaration is gone from the catalogue is reported, never removed by the mesh.
## Consequences
- A machine that joins reads the current state at once, and every machine sees a change as it
happens, with no consumer created per reader and nothing replayed.
- The runtime's channel has a sixth verb family, and the SDKs a small state surface over it — a
contract, which ADR 0039 admits: it changes when the verbs do, rarely, and every module should be
rebuilt when it does.
- What got harder: the runtime must keep each module to its own buckets, because one principal per
machine carries all of them and the server enforces only the union. A write the server refuses
surfaces to a client as a timeout, not a refusal, so the runtime's own refusal is what a module sees.
- The secrets rule is only partly mechanical. Sealed values cannot be recognised; the runtime refuses
a value with a field whose name says it is a credential, which catches the ordinary mistake and not a
determined one. For the operator's agent this means an MCP server's authorisation header stays out of
its bucket.
- Buckets accumulate as modules come and go; that they are reported rather than removed is the price
of not deleting data.
## How it is checked
| Rule | Checked by |
|---|---|
| A manifest's state names are local, and a read names a bucket its owner declares | the catalogue's registration check, per manifest; a catalogue test that every `reads` whose owner is present names a bucket that owner declares |
| Buckets exist for every declared state | the controller's raise asserts them idempotently; its test over a real bus |
| Owners write, readers only read | the composer's test of the grants, per principal kind; the runtime's refusal test over a real bus |
| A watch hands current values first, without deletions, then changes | the runtime's test over a real bus |
| No credential-named field in a value | the runtime's refusal test |
| Live | one module puts on one machine and another machine's watch sees it; a machine assigned afterwards reads it at start |
## References
- Research [024](../01-RESEARCH/024-state-a-module-keeps-on-the-bus/00-overview.md)
- [Design 32](../03-DESIGN/01-to-be/32-what-a-module-declares.md) §4 and §10, [design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md) §1 and §3
- [ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md), [ADR 0039](0039-what-the-sdk-holds-and-refuses.md),
[ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md),
[ADR 0198](0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)
@@ -0,0 +1,152 @@
---
topic: the mesh
status: accepted
date: 2026-10-02
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md
---
# 202. A provider declares what it derives for each consumer, and the mesh tells both ends
> **Written as 0188 on 2026-10-02, renumbered to 0201, and to 0202 on 2026-10-04.** Twice, for the
> same reason twice: the bundles refactor took 0188 while this waited in a pull request, and the
> key-value-buckets record took 0201 while this waited again. Both times the number was free when
> it was chosen and taken by the time this merged. Only the number moved; the decision is the one
> taken on the 2nd. The check that refuses two records sharing a number is what caught it, both
> times — a number is how a record is cited, and three repositories cite this one.
## Context
An arrangement between a consumer and a provider is delivered entirely by the mesh. Where the
provider is, which port it answers on, what name the consumer must present, where its password
is — each arrives as a fact the consumer reads from its binding, or as `${bound:…}` filled into a
file before the declaration leaves the control plane. The provider invents none of it and hands
none of it back ([ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)).
One kind of value escapes that. Where the **provider names the resource** — a bucket, a database,
a vhost — the name is derived from the consumer, per consumer, and the mesh has no way to carry
it. `serves` is a literal block in the provider's definition: the same values for every consumer.
A provisioner's contract takes a provision and returns nothing. So a value the mesh's own rule
produced reaches neither end as a statement; it is recomputed at one end and transcribed at the
other.
The object store is the instance ([issue 124](../04-ISSUES/124-a-consumer-cannot-be-told-what-its-provider-derived/00-report.md)).
Its provisioner normalises the login the mesh minted into a bucket name and creates, checks and
removes exactly that; the rule lives in twenty lines of the module's own TypeScript. Its three
consumers each write the answer into their own definition by hand. Two transcribed it correctly;
one named a predecessor's bucket, and would have authenticated successfully and been refused on
every object, which reads like a credential fault and is not one.
Even corrected, the transcriptions are wrong in a second way. Each is `mesh-<node>-<slug>`, so
each **names the machine the module happens to run on today** — a definition stating a fact about
one installation, which [ADR 0155](0155-a-definition-names-no-installation-and-how-that-is-checked.md)
forbids and whose check does not catch because the name is not a domain. Move any of the three to
another machine and its configuration points at a bucket its key cannot open.
The shape is not the object store's. A database provisioner that prefixed names, a queue provider
that scoped vhosts, any provider that derives a resource from who is asking: each forces the
consumer to reproduce somebody else's rule and keep it in agreement by hand.
## Decision
**1. A served value may name the consumer the mesh is serving.** A `serves` block, which is
literal today, may interpolate the mesh's own statement of who the consumer is:
- `${consumer:as}` — the identity the mesh minted for this consumer, exactly as the login it is
told to present ([ADR 0049](0049-a-consumers-identity-fits-the-tightest-backend.md));
- `${consumer:as:dns}` — the same identity written as a DNS label.
Nothing else. **The mesh learns no protocol here; it spells its own name in an alphabet it already
knows.** The identity is the mesh's, minted by the mesh, already capped at twenty characters
because of what an S3 access key accepts; `dns` is that same name with its separator written `-`
instead of `_`, which is the whole of the difference between the mesh's identifier alphabet and
the one buckets, vhosts and hostnames use. A provider that needs a prefix or a suffix writes it
around the placeholder, because a served value is a string.
The rejected alternative is **the provider returning values from provisioning** — the natural
channel, since the provider is what derived them. It is rejected for three reasons, in order of
weight. It inverts the delivery the mesh is built on: a grant would carry data the provider wrote
rather than only data the mesh minted, and [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)
removed exactly that second path once already. It makes a consumer's declaration incomplete until
its provider's reconcile loop has run, so a consumer could not be composed before a provider
answered — a bootstrap order the mesh does not have and does not want. And it puts the rule where
nothing can check it: a value that arrives from a running process cannot be refused at resolution,
only discovered wrong later, which is the failure this record exists to end.
**2. The mesh resolves it once, per consumer, and tells both ends from the one resolution.** At the
moment a consumer's declaration is composed, the mesh knows exactly who the consumer is. There, and
only there, the placeholders are filled. The result reaches:
- the **consumer**, as the served facts in its binding file and as `${bound:<provision>:<key>}` in
any file it writes — unchanged mechanisms, carrying one more key;
- the **provider**, as `serves` on that consumer's entry in its contributions file, so the
provisioner is *told* the name rather than recomputing it.
**The provider stops deriving in code and starts declaring.** One statement, filled once, delivered
to both ends: the two cannot disagree, because there is no second computation to disagree with.
**3. A served value stays settled before it is per-consumer.** Settings still compose into `serves`
([ADR 0174](0174-a-node-varies-a-module-through-settings-and-kept-regions-never-an-edit.md)), and the consumer
placeholders are filled after that, so an operator may set a prefix and the mesh still derives the
rest. A `${consumer:…}` naming a fact or an alphabet the mesh does not have is refused when the
definition is parsed, with what it may say.
**4. A consumer may no longer name the resource its provider derives.** With the value delivered,
a literal in a consumer's definition is not merely redundant — it is the one thing that can
disagree with what the provider will actually create. The three object-store consumers lose their
hand-written bucket names in this change.
**5. A consumer that keeps several holders of one provision may not be served a derived value.**
Each holder gets its own login, `…_<local>` ([ADR 0094](0094-a-module-may-hold-several-secrets-from-one-provider.md)),
and a provider derives from the login — so it would make one resource per holder, while the
consumer's side has one binding and one `${bound:<provision>:<key>}`, both derived from the
un-suffixed identity. That is this record's own failure one case to the side, and just as quiet:
the consumer would authenticate and be refused on every object. Refused at resolution, naming
both ends. Lifting it means giving the consumer's side a local dimension, which is a decision and
not an omission.
## Consequences
- One more thing a definition may say, and one less thing a module may be wrong about. The
vocabulary grows by a placeholder; the catalogue loses three literals that named this
installation's control node.
- A provider's naming rule becomes readable in its definition instead of in its source. `minio`'s
`bucketFor` goes; the manifest says `"bucket": "${consumer:as:dns}"` and the provisioner uses
what it is given.
- A provider that already serves consumers keeps serving them: the derived value equals what the
code derived, so no bucket, database or login changes name. This is a change of **who says it**,
not of **what is said**.
- A refusal here fails **that machine's push**, naming the definition, and nothing else. That is
deliberate and is the opposite of a module quietly left out: a definition that transcribes
somebody else's rule is wrong everywhere, not just here, and the loud failure is in front of
whoever can fix it.
- The mesh now holds a rule in another system's alphabet — one rule, `dns`, stated once. A second
alphabet is a decision, not an addition: the cost of each is that the mesh must be right about
somebody else's naming, and that cost is only worth paying where the mesh already mints the name.
## How this is checked
- A served value naming an unknown fact or alphabet is refused at parse, with the list of what it
may say — tested on both halves of the message.
- Resolving a consumer whose provider derives a value puts that value in the consumer's binding
file, in its `${bound:…}` substitutions, and in the provider's contributions entry for that
consumer — one test asserting the three agree, because agreeing is the whole point.
- Two consumers of one provider on one machine get two different derived values, and neither gets
the other's.
- A consumer with several holders of a deriving provider is refused, with both ends named — the
test asserts the refusal, not merely that something failed.
- A catalogue-wide test refuses a consumer definition that writes a literal where its provider
derives: the provider's `serves` names the key, so the catalogue can say which definitions
transcribe one.
- `dns` is checked against the identity the mesh actually mints, not against an invented string:
the test derives an identity with `ConsumerIdentity` and asserts the label it becomes.
## References
- [issue 124 — a consumer cannot be told a value its provider derived for it](../04-ISSUES/124-a-consumer-cannot-be-told-what-its-provider-derived/00-report.md)
- [ADR 0048 — a provider creates the credential the mesh minted, and seals nothing](0048-a-provider-creates-the-credential-the-mesh-minted.md)
- [ADR 0049 — a consumer's identity fits the tightest backend](0049-a-consumers-identity-fits-the-tightest-backend.md)
- [ADR 0174 — a node varies a module through settings and kept regions, never through an edit](0174-a-node-varies-a-module-through-settings-and-kept-regions-never-an-edit.md)
- [ADR 0155 — a definition names no installation, and how that is checked](0155-a-definition-names-no-installation-and-how-that-is-checked.md)
- [design 27 — a module requires, the mesh resolves](../03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md)
+4
View File
@@ -188,7 +188,9 @@ python3 00-META/checks/index.py fail if stale
- **0185** — [A control plane behind its seat's row serves what it can](0185-a-control-plane-behind-its-seats-row-serves-what-it-can.md)
- **0186** — [A ban list never holds a neighbour, and the mesh's own bans are its own wherever they hang](0186-a-ban-list-never-holds-a-neighbour.md)
- **0187** — [A dead tracker is not the machine's failure](0187-a-dead-tracker-is-not-the-machines-failure.md)
- **0189** — [The store keeps what the records name, and a maintenance step holds its writers still](0189-the-store-keeps-what-the-records-name.md)
- **0190** — [A seat's work is shared by its holders, and building is the first such role](0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md)
- **0202** — [A provider declares what it derives for each consumer, and the mesh tells both ends](0202-a-provider-declares-what-it-derives-for-each-consumer.md)
### Its tiers, from the bottom up
@@ -298,6 +300,7 @@ python3 00-META/checks/index.py fail if stale
- **0195** — [The mesh's tools are found by address, not announced whole](0195-the-meshs-tools-are-found-by-address-not-announced-whole.md)
- **0197** — [Every tool announces itself on the bus, in the NATS services protocol](0197-every-tool-announces-itself-on-the-bus-in-the-nats-services-protocol.md)
- **0198** — [A module's long-running code is launched by the node's runtime, and reaches the bus through it](0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)
- **0201** — [A module keeps its current state in key-value buckets it declares, and reaches them through the runtime](0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md)
### How it is built
@@ -320,6 +323,7 @@ python3 00-META/checks/index.py fail if stale
- **0111** — [A build source is on the mesh's git seat, or it is an external repository](0111-a-build-source-is-on-the-git-seat-or-external.md)
- **0149** — [The live mesh is the test bed](0149-the-live-mesh-is-the-test-bed.md)
- **0174** — [A node varies a module through settings and kept regions, never through an edit](0174-a-node-varies-a-module-through-settings-and-kept-regions-never-an-edit.md)
- **0200** — [Genesis pivots to the controller as a container, and the first push hands it to a process](0200-genesis-pivots-to-the-controller-as-a-container-and-the-first-push-hands-it-to-a-process.md)
### How it is checked
+40 -1
View File
@@ -5,10 +5,11 @@ code:
- mesh-controller cmd/mesh-builder
- mesh-controller internal/builder
- mesh-catalog modules/build-agent
updated: 2026-10-03
updated: 2026-10-04
decisions:
- 02-DECISIONS/0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md
- 02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md
- 02-DECISIONS/0189-the-store-keeps-what-the-records-name.md
- 02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md
- 02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md
- 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md
@@ -317,6 +318,44 @@ build's lines reach a reader of its subject in order and the stream holds them a
against a real server); the seat verb with an id reads the log (controller test); and, live, a build
after the roll-out read line by line through the console.
## The store keeps what the records name
*2026-10-02 — [ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md),
[issue 108](../../04-ISSUES/108-the-registry-has-no-garbage-collection-once-it-has-two-doors/00-report.md).*
Every build pushes another layer set and, until this, nothing ever removed one. The registry's own
answer — collect what no tag names — is wrong for this mesh: each artifact is pushed under one
moving tag and machines are pinned by digest, so every build but the newest is untagged and some
machine may still be running it.
**The mesh decides and the store reclaims.** Deletion is enabled on the store's one door — that
door already accepts a push, and a writer who can push can replace any tag, so delete takes
nothing a push did not already have, and ADR 0082's bargain (a store every machine reaches with no
credential to distribute first) is kept. The mesh then removes what it put there and no longer
keeps, **naming it from its own build records** rather than enumerating the store: it has never
put anything there it did not record, so a digest it did not record making is never named, which
is what keeps the sweep away from the images genesis pushed before any record existed.
An artifact stays for one of two reasons and otherwise goes: a definition the mesh holds names it
(no age limit — this is the floor), or it belongs to one of the five most recent successful builds
of its module (somewhere for a wrong release to return to). The sweep runs after a build the mesh
recorded, which is the moment new bytes landed and the moment the keep set moved; it needs no
timer. Deleting a manifest frees no bytes, so the store's own collector runs nightly as a
scheduled step with the server held still — which is what `while-stopped` exists for
([design 20](20-writing-a-module.md)). Plain collection, not `--delete-untagged`: what the mesh
keeps is still a manifest and so still referenced, and the dangerous flag is not needed once the
mesh is the one deciding.
A machine behind by more than five builds of a module, recreating a container, cannot pull what it
was running. It is already a machine the mesh reports as behind, and the answer is the current
declaration.
*How it is checked:* the keep set, against records, holds what a manifest names and the five most
recent builds and nothing else; a reference the mesh never recorded is never in the delete set; an
image and an archive are asked for at their own endpoints; a store with deletion off names the
remedy rather than the status code; a store that does not have it is recorded collected rather
than retried for ever. Live: the store's size before and after the first nightly collection.
## The builder compiles the languages the mesh is written in
*2026-09-29 —
+26 -1
View File
@@ -5,11 +5,12 @@ code:
- mesh-catalog modules/showcase
- mesh-controller internal/builder
- mesh-sdk src
updated: 2026-09-30
updated: 2026-10-02
decisions:
- 02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md
- 02-DECISIONS/0099-a-step-that-runs-once-names-what-it-reads.md
- 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md
- 02-DECISIONS/0189-the-store-keeps-what-the-records-name.md
- 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md
- 02-DECISIONS/0040-what-a-module-is.md
- 02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md
@@ -208,3 +209,27 @@ is recreated with the new fact
([ADR 0099](../../02-DECISIONS/0099-a-step-that-runs-once-names-what-it-reads.md)). *How it is
checked:* the host's unit tests run a step again when its named file changed and not otherwise,
and recreate a container naming a step after the step ran.
## A recurring step may hold its own module's containers still
*Written 2026-10-02, from [ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md)
and [issue 108](../../04-ISSUES/108-the-registry-has-no-garbage-collection-once-it-has-two-doors/00-report.md).*
Some work cannot be done underneath a running service: an artifact store's collector walks the
storage and requires every writer stopped. A `run-once` step runs *beside* containers and a
scheduled one is the same container again, so until this a module had no way to say it — and the
mesh inherited a store that has never collected anything, because the predecessor said it with a
shell script and a script beside a module is not a resource in it.
A scheduled step may name `while-stopped`: resource ids of **its own module's** containers, which
the host stops before the run and starts again after it, in the reverse order, **whatever the step
did**. Three boundaries, each refused where it can be seen earliest — its own module's containers
only, because a module that could quiesce a neighbour could stop the mesh; scheduled steps only,
because at apply the declaration is applied in order and a step already gates what follows, so a
one-time offline job says *before* rather than *instead of*; and restoring that is not conditional
on anything, because the only real risk of the field is a window that never closes.
*How it is checked:* the host's unit tests assert stop–run–start in that order, the restart after a
step that **failed**, the reverse order for several containers, and a service left down said
loudly. The controller refuses, from the definition alone, a window with no schedule, one on a
run-once step, one naming a container the module does not declare, and one naming itself.
+23 -2
View File
@@ -7,8 +7,10 @@ code:
- mesh-tools src/broker-amqp.ts (to be replaced)
- mesh-catalog modules/nats (to be written)
- mesh-sdk src (the protocol's NATS binding, step 3)
updated: 2026-10-02
- mesh-tools node-tools/internal/bus (a module's state, ADR 0201)
updated: 2026-10-04
decisions:
- 02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md
- 02-DECISIONS/0167-a-membership-carries-what-its-module-receives-and-who-the-mesh-is.md
- 02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md
- 02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md
@@ -48,6 +50,7 @@ mesh's own state lives, and where what a module may say is decided by what it de
| **events** — a module says something happened | 1:many | delivered to every consumer that declared it; dead-lettered when it cannot be |
| **tools** — a module or a person asks another's tool | request/reply | one answer, from one server, or a timeout |
| **work to a role** — a module submits to a capability without knowing who provides it | job | exactly one holder does it; it queues while nobody does |
| **state** — a module's current value of something, every machine reading it | key-value | the newest per key, kept until replaced or deleted; read whole by a machine that joins later |
The last two rows are the ones worth dwelling on, because they are not messaging in the sense of
carrying bytes from A to B. **A role is addressable**, so a caller names the capability and never
@@ -56,7 +59,8 @@ changing ([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md)
property of the mesh's architecture that happens to be expressed in subjects.
And more of the mesh lands here as it is built: conditions and observed state in key-value
buckets that anything may watch, the server's own advisories becoming observations like any other
buckets that anything may watch — the first of them a module's own declared state, *2026-10-04*
([ADR 0201](../../02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md)) — the server's own advisories becoming observations like any other
([research 017](../../01-RESEARCH/017-a-mesh-that-heals-itself/00-overview.md)), and a person's
client speaking the bus directly rather than through a surface built over it (§7). None of that
is a message being moved; all of it is the bus being the mesh's centre.
@@ -84,8 +88,16 @@ mesh.seat.<seat>.event.<verb> a role's own event (JetStream: EVENT
mesh.seat.<seat>.tool.<verb> a role's tool (core request/reply)
mesh.ask.<node>.<command> the controller's command api (core request/reply)
mesh.assignment.<node>.<module> an assignment's membership (JetStream: ASSIGNMENTS, last-per-subject)
$KV.<module>_<name>.<key> a module's state (JetStream: a key-value bucket per declared name)
```
**Added 2026-10-04** ([ADR 0201](../../02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md)):
the last row is outside `mesh.` on purpose. A key-value bucket is NATS's own construct and lives
under NATS's own prefix, which is what lets the server's key-value layer — direct reads, rollups,
delete markers, watches — do the work instead of the mesh writing it again. The bucket is named for
the module and the local name joined by an underscore, which neither may contain, so two modules can
never derive one bucket.
**Revised 2026-10-01** ([ADR 0160](../../02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)): the rows above
for a module's and a seat's tools are the shapes the controller *issues*, not rules a runtime carries.
Every assignment is published a membership — what it serves and where, in which queue, its seat verbs,
@@ -155,6 +167,7 @@ Core NATS is at-most-once. Everything the mesh must not lose lives in a JetStrea
| CONTROL | `mesh.control.>` except `alive` (a build's outcome moved to its seat, ADR 0121) | work queue, one consumer (the controller), explicit ack | the store-window guarantee ([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)): the controller `nak`s with a delay while its store is away and the message is redelivered; nothing is dropped |
| NODES | `mesh.node.>` | last per subject | one declaration per node, always the newest |
| EVENTS | `mesh.mod.*.event.>` and `mesh.seat.*.event.>` | limits (age, size), durable consumer per subscribing module | a subscriber that was down catches up; after `max-deliver` attempts the advisory feeds `mesh.events.dead` (its own small stream). *2026-10-01:* a build's whole log is here too, as the build-machine seat's `log.<build id>` events ([ADR 0157](../../02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md)) — one subject per build, a week of retention, read back by `builds --log <id>` with a consumer that is gone when the reading is done |
| `KV_<module>_<name>` | `$KV.<module>_<name>.>` | the newest value per key — as many past values as the owner declared — no age unless the owner declared one; a value at most 256 KiB, a bucket at most 64 MiB | a module's state (ADR 0201): one per name in a manifest's `state`, created from the catalogue on every raise, so it exists before its owner runs anywhere; kept when the module is unassigned, because what it holds is data |
Tool calls and heartbeats stay on core NATS: a lost heartbeat is the next heartbeat; a lost tool
call is a timeout the caller already handles.
@@ -234,6 +247,14 @@ expresses this exactly, per subject, and better than a vhost could:
permissions for each consumed event's subject, its tool subjects, and that same inbox prefix.
Nothing else. A module that tries to publish outside its emits is refused by the server, not by
convention.
- **A module's state** (ADR 0201), for whichever principal carries the module — today the machine's
runtime, whose grant is the union of its modules': binding to the bucket, reading a key directly,
and an ordered consumer for listing and watching, created and deleted on the bucket's own stream
and nothing else's; and, for the owner's instances only, publishing under the bucket's own
`$KV.<bucket>.>`. *Measured 2026-10-04 against a running server:* without the consumer-delete
grant a watch cannot be stopped cleanly, and a write the server refuses reaches the writer as a
timeout rather than a refusal — so the runtime refuses first, from the membership, and the grant
is the second line.
- **The controller's user** owns `mesh.control.>`, `mesh.node.>` and the streams, and may submit work
to the seats the mesh's own flows use — a build, for one (ADR 0121).
**A host's user** may publish its own `mesh.control.<node>.>` and subscribe its own
@@ -2,7 +2,7 @@
layer: to-be
status: in-progress
code: [mesh-controller internal/catalogue]
updated: 2026-09-30
updated: 2026-10-02
decisions:
- 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md
- 02-DECISIONS/0155-a-definition-names-no-installation-and-how-that-is-checked.md
@@ -14,6 +14,7 @@ decisions:
- 02-DECISIONS/0084-which-provider-serves-a-consumer.md
- 02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md
- 02-DECISIONS/0038-the-mesh-assigns-the-port.md
- 02-DECISIONS/0202-a-provider-declares-what-it-derives-for-each-consumer.md
---
# 27 — A module requires, the mesh resolves
@@ -206,6 +207,22 @@ name when nothing sets it. That is the contract half of this design's operator p
the placeholder allows: the definition says which values reach which requirement, and nothing else
does. *How it is checked:* the unit tests named in issue 173, and the plan comparison that closed it.
*A provider says once what it derives for each consumer (2026-10-02,
[ADR 0202](../../02-DECISIONS/0202-a-provider-declares-what-it-derives-for-each-consumer.md),
[issue 124](../../04-ISSUES/124-a-consumer-cannot-be-told-what-its-provider-derived/00-report.md)):*
where a provider **names the resource** it gives each consumer — a bucket, a database, a vhost — the
name is derived per consumer, and a literal `serves` block could not carry it. A served value may
now name the consumer the mesh is serving: `${consumer:as}`, the identity the mesh minted, and
`${consumer:as:dns}`, that same identity written as a DNS label. Nothing else — **the mesh learns no
protocol here; it spells its own name in an alphabet it already knows.** Settings are laid on first,
so an operator may still set a prefix and the mesh derives the rest. The mesh fills it at the one
moment it knows who the consumer is, and the one filled value reaches both ends: the consumer, as
its binding's served facts and as `${bound:<provision>:<key>}` in any file it writes; the provider,
as `derived` on that consumer's entry in its contributions file, so its provisioner is told the name
rather than recomputing it. A consumer that writes the derived value into its own definition instead
of asking for it is refused, naming the placeholder to use. *How it is checked:* the unit tests in
ADR 0202's "how this is checked", each run against the unchanged controller first.
## How a definition reads what was resolved
**One form, naming a requirement and a field of its contract.** A definition that needs the database's
@@ -214,7 +231,9 @@ name in a configuration file writes the same thing: the requirement's name and t
controller fills it at resolution.
This one form replaces the placeholders that exist today, one per mechanism: bound values, secrets,
ports and machine facts.
ports and machine facts. It subsumes the consumer placeholder too — a value a provider derives is
read by the consumer exactly as any other field of the contract is, and `${consumer:…}` is only
how the *provider* states the rule.
**The seat placeholder stays, for the controller alone.** The controller composes its own
declaration and reaches the store and broker it made before any module existed, so it cannot be
@@ -11,8 +11,10 @@ code:
- mesh-host internal/apply/apply.go
- mesh-tools src/main.ts
- mesh-catalog modules/mesh-catalog
updated: 2026-10-02
- mesh-tools node-tools/internal/runtime (a module's state, ADR 0201)
updated: 2026-10-04
decisions:
- 02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md
- 02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md
- 02-DECISIONS/0126-a-module-declares-its-own-seats.md
- 02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
@@ -67,6 +69,8 @@ the catalogue, and the mesh would have hundreds of copies of a decision it made
| `tools: status` | queue-group subscription on `mesh.mod.<module>.tool.status` |
| seat `telegram-sender`, `accepts: send` | work-queue consumer on `mesh.seat.telegram-sender.accept.send` |
| `uses: telegram-sender` | publish on that seat's `accept` subjects, and nothing else |
| `state: servers` | a key-value bucket for the module, created by the controller; its instances write and read it |
| `reads: billing.orders` | read and watch billing's `orders` bucket, and nothing else of it |
**Wildcards, and they are the mesh's rather than a bus's.** *Added 2026-09-27, from
[issue 127](../../04-ISSUES/127-a-module-event-derives-a-subject-nothing-publishes/00-report.md).* A
@@ -221,7 +225,7 @@ service" versus "one worker per machine".
| credential | sealed, per consumer | none | none | none | none |
| reply | — | none | none, or an event later | a report | awaited |
| retention | — | age and size | work queue, explicit ack | **last per subject** | none |
| declared | `provides`/`requires` | `emits`/`consumes` | seat `accepts` / `uses` | the mesh's own | `serves` |
| declared | `provides`/`requires` | `emits`/`consumes` | seat `accepts` / `uses` | the mesh's own; a module's `state` / `reads` | `serves` |
**Job** is the one [ADR 0041](../../02-DECISIONS/0041-events-are-a-relationship.md) had no room
for. Its table has events at 1:many and provisions at 1:1; a module submitting work to a service
@@ -236,6 +240,31 @@ last-per-subject retention, and a node that has seen sequence *n* refuses *n−1
That is the wire-level answer to
[issue 107](../../04-ISSUES/107-a-declaration-carries-no-order/00-report.md).
**A module declares state too.** *Added 2026-10-04,
[ADR 0201](../../02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md).*
State was the mesh's alone, and modules had the same need with nowhere to put it: an MCP server
registered for every machine, sent as an event, never reached a machine assigned afterwards — its
consumer did not exist yet when the event passed — and a licence binding sent as events replays a
week of rotations where only the latest matters. So a module names the state it **owns** with
`state`, and another module's it **reads** with `reads: <module>.<name>`. Each is a key-value bucket
the controller creates from the catalogue, mesh-wide, existing from registration so a reader can
watch before the owner runs anywhere ([design 25](25-the-bus-on-nats.md) §3). Every instance of the
owner writes; a reader reads and watches. A key may name a machine by the module's own convention;
the mesh keeps one bucket per name, not one per machine, because "every server, for every machine"
is then one list rather than a walk.
What a module sees is what it named. Its assignment's membership lists its buckets by those names,
with whether it may write, and the runtime answers `get`, `put`, `delete`, `keys` and `watch` for
them on the bundle's channel — refusing, with the reason, a name it was not issued or a write to a
bucket it only reads. A watch hands the current values first, then every change: a bundle that starts
late, or starts again, has the whole of the state before it has any of the news.
The owner says how many past values a key keeps and how long a value lives, as a seat says how long
its backlog survives (§3); the mesh caps a value's size and a bucket's. **A bucket outlives its
module's assignment** — what a module stored is data
([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)) — and one whose
declaration is gone is reported, never removed by the mesh.
## 5. Seats
A module declares a seat with its protocol, and the mesh enforces one holder at its scope
@@ -454,6 +483,14 @@ sealing key leaks, that stream is an archive rather than a moment. So:
existing discipline — *fetched from it, not carried* — applied to the one payload where carrying
it is worst.
**A key-value bucket is a stream, so the same holds for it.** *Added 2026-10-04,
[ADR 0201](../../02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md).*
No secret is put in a module's state, sealed or not: state is exactly what a machine joining a year
later reads in full. A value that needs a secret names it, and the secret travels on request/reply.
Sealed values are plain text to anything inspecting them, so this is checked only partly — the
runtime refuses a value carrying a field whose name says it is a credential, which catches the
ordinary mistake and not a determined one.
**The bootstrap, which is circular and has a precedent.** The vault makes every secret
([ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md)), including the bus's own
passwords. The vault is a module, and a module needs a bus account, whose password the vault
@@ -523,3 +560,9 @@ billing existing under that name.
on, and exactly those two are rebuilt.
- **A stale declaration is refused.** A bed: replay sequence *n−1* after *n*, and the node refuses
it rather than applying it.
- **A module reaches only the state it declared.** The composer's test: an owner's runtime may write
its buckets, a reader's may only read, and nothing else is granted; the runtime's test over a real
bus: a name not issued and a reader's write are refused with the reason.
- **State is current at once.** The runtime's test over a real bus: a watch hands the current values
without deletions, then an end-of-current marker, then changes. Live: a machine assigned after a put
reads it at start.
@@ -2,7 +2,7 @@
layer: to-be
status: in-progress
code: [mesh-tools, mesh-controller, mesh-host, mesh-catalog]
updated: 2026-10-03
updated: 2026-10-04
decisions:
- 02-DECISIONS/0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md
- 02-DECISIONS/0173-the-operators-machine-is-the-meshs-and-a-module-is-what-it-declares.md
@@ -309,6 +309,25 @@ nextcloud, minio), the two whose clients exist on no system (mongodb, mssql: a d
and last the mesh's own (mesh-catalog, mesh-vault, records, gitea, mailu, audit-logger, lab, and the
three mains).
*Built 2026-10-04.* Thirty-four modules no longer run their own code in a container: the first wave
(mesh-catalog#245), the mesh's own and the media modules (mesh-catalog#248, mesh-media-catalog#13),
with a step run where and as it is declared (mesh-host#85) and a process's words filled like a
container's (mesh-controller#250). **Proven live** on every machine that runs them: each moved
module's tools answer from the node's runtime, the steps run as their oneshot units, and the forge's
merge events reach the build pipeline from the runtime — the merge after the move started its own
plan. **Two corrections the machines taught:** a tool that called a broker's command-line client now
runs it inside the broker's own container, because a host package may be uninstallable on a machine
whose package index is stale (mesh-catalog#249); and a module reading its application's own key reads
it through the application's container, because that directory belongs to the account the application
runs as there, which is not the runtime's (mesh-media-catalog#14). *Completed 2026-10-04:* the last three — mesh-catalog, mongodb and mssql, whose code imports npm
packages of its own — moved once the builder installs a bundle's own dependencies before compiling,
keeping the toolchain's SDK authoritative (mesh-controller#255, mesh-catalog#250). Their database
clients are now drivers inlined into the bundle, not command-line clients fetched by a container; one
more correction the machines taught: a driver reaching its server on loopback must give TLS a host
name, since the runtime's Node refuses an address (mesh-catalog#252). **No module's own code runs in
a container any more;** every module with tools answers from its node's runtime, proven by calling a
tool of each.
## WP5 — The shell, on a server first
*mesh-catalog #224, already written. Half a day to assign and prove.*
@@ -1,9 +1,9 @@
---
status: open
status: resolved
opened: 2026-09-23
located-in: []
fixed-by:
amended-design:
located-in: [mesh-controller internal/inventory, mesh-controller internal/artifacts, mesh-host internal/apply, mesh-catalog modules/distribution]
fixed-by: 02-DECISIONS/0189-the-store-keeps-what-the-records-name.md
amended-design: 03-DESIGN/01-to-be/18-building-a-module.md
---
# 108 — The registry has no garbage collection, and two doors make it harder to add
@@ -75,3 +75,32 @@ real thing services need, and the mesh cannot express one.
images by digest and moves by version — is retention "the digests no recorded build names"?
- Who owns the routine when the store and its public door are two modules — the store, since the
volume is its?
## Answered, 2026-10-02 — [ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md)
The three open questions, answered:
- **A maintenance step, or a backend that does not need its writers stopped?** The step. A
scheduled container may name `while-stopped` — resource ids of **its own module's** containers,
which the host stops before the run and starts again after it whatever the step did. A storage
backend the mesh does not run would be a bigger thing to own than the mechanism it avoids, and
the mechanism is wanted anyway: a service that cannot have work done underneath it is a real
shape and the mesh could not express it at all.
- **Is retention "the digests no recorded build names"?** Nearly. An artifact stays because a
definition the mesh holds names it (no age limit), or because it belongs to one of the five most
recent successful builds of its module. Last-N-tags was the predecessor's rule for a registry
that knew nothing else; this mesh knows what each digest is for.
- **Who owns the routine now the second door is gone?** Both halves, each where it can be. The
**mesh** decides what may go — only it holds the records — and asks the store to drop it. The
**store** reclaims the bytes, because only it can stop its own server. Neither half can be done
by the other.
And the sharpened point — enabling deletion on a door with no accounts — dissolved on inspection:
**that door already accepts a push**, so a writer who can reach it can already replace any tag.
Delete takes nothing a push did not have. What it does not do is undo ADR 0082's bargain, which
putting an authenticated door in front of deletion would have.
The second registry process is not built, as the 2026-09-26 note says, so the shared blob cache
and the deletion-cached-by-the-other-door problem never arise. Plain `garbage-collect` is enough:
what the mesh keeps is still a manifest in the store, so `--delete-untagged` — the flag that would
delete images machines are running — is not needed at all.
@@ -1,9 +1,9 @@
---
status: located
status: resolved
opened: 2026-09-26
located-in: [mesh-controller internal/catalogue/declaration.go, mesh-sdk src/provisioner, mesh-catalog modules/minio]
fixed-by:
amended-design:
located-in: [mesh-controller internal/catalogue, mesh-sdk src/provisioner, mesh-catalog modules/minio]
fixed-by: 02-DECISIONS/0202-a-provider-declares-what-it-derives-for-each-consumer.md
amended-design: 03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md
---
# 124 — A consumer cannot be told a value its provider derived for it, so it transcribes one
@@ -63,3 +63,27 @@ compares it to what the provider will actually create. The one wrong instance wa
- What would have caught the wrong instance? A test that resolves a consumer's grant and compares the
bucket in its own configuration against the one the provider would create is a check that could
exist today, for any interface, without the mechanism above.
## Answered, 2026-10-02 — [ADR 0202](../../02-DECISIONS/0202-a-provider-declares-what-it-derives-for-each-consumer.md)
The channel is the provider's own `serves` block, which may now name the consumer the mesh is
serving: `${consumer:as}` and `${consumer:as:dns}`. The mesh fills it once, where it knows who the
consumer is, and delivers the one filled value to both ends — the consumer's binding and its
`${bound:…}` substitutions, and the provider's contributions entry, so a provisioner is told the
name rather than deriving it. Each open question above, answered:
- **Should a provider return values from provisioning?** No. It would make a grant carry data the
provider wrote, make a consumer's declaration wait on its provider's reconcile loop, and put the
rule where nothing can refuse it. The reasoning is in the record.
- **Or should `serves` say a value is derived?** Yes, and the mesh performs the derivation — but it
learns no protocol doing it. The only fact is the identity the mesh itself minted, in one of two
alphabets it already knows.
- **Should a consumer that names the resource be refused?** Yes. A consumer's file that already
contains the value the mesh is about to derive for it is refused at resolution, naming the
placeholder to write instead. That is the check this report asked for, and it is exact rather than
heuristic: a derived value carries the identity minted for this consumer on this machine, which
nothing else would spell out.
minio's `bucketFor` is gone; its manifest serves `"bucket": "${consumer:as:dns}"`. The three
consumers' hand-written bucket names are gone with it — each of them also named the machine the
module happens to run on, which is the second thing wrong with a transcription.
@@ -0,0 +1,93 @@
---
status: open
opened: 2026-10-02
located-in: [mesh-catalog modules/dnsmasq, mesh-controller cmd/mesh-controller]
fixed-by:
amended-design:
---
# 202 — A module whose required setting nobody set is left out of the machine, and the resolver is the module it happened to
## What was observed
Running the controller's own test suite against the catalogue beside it, 2026-10-02.
`TestTheResolverIsToldEveryMachineOnTheNetworkAndToldAgainWhenOneLeaves` fails with *"the resolver
was not handed the machines"*. Composing the same machine by hand and listing what it receives
shows why: **dnsmasq contributes nothing at all.** Four resources are composed for that node, all
of them the overlay's. The resolver's package, its configuration, its service and the fact that
carries every machine's name are simply not there.
The cause is one line added to `dnsmasq`'s configuration earlier the same day: the addresses it
listens on beside the machine's own became an operator setting,
`listen-address=${setting:listen-addresses}`, with no default. A `${setting:…}` nothing sets is
refused, a module that cannot be composed is **left out** rather than failing the whole machine
([ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md)), and so a node
assigned the resolver is handed a declaration with no resolver in it.
The failing test is the symptom that surfaced it. The test is not what is wrong.
**Proven rather than inferred.** Composing the same machine a second time with
`listen-addresses` set to `127.0.0.1` and nothing else changed, every one of dnsmasq's eight
resources appears — `needs-broker`, `mesh-state`, `package`, `config`, `runtime-dns`, `runtime`,
`service` and `fact-node-zones`. The only difference between a machine with a resolver and a
machine without one is whether somebody set a value that did not exist yesterday.
## Why it matters beyond this instance
**Leaving a module out is right, and being quiet about it is not.** The rule exists so one
module's broken setting cannot stop a machine converging — a good rule. But the outcome here is a
machine that applies cleanly, reports current, and is missing its DNS resolver. Every name on that
machine then resolves through whatever was there before, or not at all, and nothing in the mesh
says the resolver was dropped. That is the shape
[issue 152](../152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md) records for
routed names, here for a whole module.
**And a setting with no default is a definition that cannot be assigned.** Every other
`${setting:…}` in the catalogue names something that is genuinely particular to one installation —
a public domain, an issuer. "Which addresses besides my own do I answer on" has an obvious correct
default for every machine that is not a LAN gateway: none beside loopback. A definition that
refuses to compose until somebody sets a value most machines do not need is a definition that
breaks the next node to be assigned it, and genesis with it.
## What this does not claim
Whether the live machines are affected was not checked — those four have had the setting set, or
their resolvers would already be gone. The claim is about a machine assigned the resolver *from
now on*, and about the silence.
## Open questions
- Should the declaration say which modules it left out, where a person or the console can see it?
`left_out` already travels to the host ([ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md));
what is missing is anything that reads it back and says so.
- ~~Should a `${setting:…}` be allowed a default?~~ **Decided in principle and not built.**
[ADR 0164](../../02-DECISIONS/0164-a-setting-is-declared-with-its-default-its-meaning-and-what-changing-it-costs.md)
(proposed, 2026-10-01) says a setting with a default is a tunable and one without is the
operator's, and narrows 0155's refusal to exactly the second. `listen-addresses` is a tunable by
that rule, and dnsmasq declares no settings block at all. So this issue is, in part, 0164 waiting
to be built — and in part the silence, which 0164 does not address.
- Is leaving a module out ever right for a module a node is **assigned**, as opposed to one it
merely pulls in? An assignment is somebody saying *this machine runs this*; silently not running
it is the one answer nobody asked for.
## Still true on 2026-10-04, and the evidence had to be re-taken
Re-checked after the bundles refactor landed (fourteen records, ADRs 0188 and 0190–0200). **The
fault stands and the old evidence no longer reaches it.**
`TestTheResolverIsToldEveryMachineOnTheNetworkAndToldAgainWhenOneLeaves` still fails on
mesh-controller main, with the same message — and now for a *different first reason*. `assign` is
refused before composition ever happens:
> dnsmasq on anchor has no bus credential: nothing was issued for anchor.dnsmasq … (novox/hq issue 203)
That is [issue 203](../203-a-fresh-assignment-is-pushed-before-its-credential-exists/00-report.md)'s
new guard doing its job on a test harness that mints no credential. Two faults are stacked in one
failing test, and the second was invisible behind the first.
Proven again, past both: mint `anchor.dnsmasq` so the assignment stands, then compose the machine
twice. **Without `listen-addresses` the node composes four resources, all the overlay's. With it
set to `127.0.0.1`, all eight of dnsmasq's appear** — `needs-broker`, `mesh-state`, `package`,
`config`, `runtime-dns`, `runtime`, `service`, `fact-node-zones`. Nothing else differs.
The test is now wrong about two things and should be fixed with whichever of these is fixed first.
@@ -1,9 +1,10 @@
---
status: located
status: resolved
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by:
- mesh-controller#246
amended-design:
---
@@ -45,3 +46,7 @@ have to write — a bundle of language L depends on the module that publishes L'
bundle lands in the tier after it. A test: a merge touching the toolchain module and a TypeScript
bundle plans the bundle one tier later. Worked around on the day by building the bundle again once
the toolchain was built.
## Resolved
Every mesh-tools plan since the fix ran in two rounds: the toolchain and runtime images first, what is built in or on them after. Before it, a merge moving both tiered them together. The planner's test proves the order: a merge moving a bundle and its toolchain plans two rounds.
@@ -1,10 +1,11 @@
---
status: located
status: resolved
opened: 2026-10-03
located-in:
- mesh-tools
- mesh-controller
fixed-by:
- mesh-tools#44
amended-design:
---
@@ -42,3 +43,7 @@ install from a lockfile that a release updates, or build the install layer witho
the build should record which SDK version the image carries, so a bundle's record says what it was
compiled against. A check: after an SDK release and a toolchain rebuild, a bundle built on it reports
the released version.
## Resolved
The toolchain now installs the exact SDK version the mesh last published, passed in as a build argument from the SDK module's published package, and the planner orders the toolchain after the SDK. Proven 2026-10-04: the toolchain built with the published SDK, and the six TypeScript bundles built in it serve their tools and seat verbs.
@@ -1,10 +1,14 @@
---
status: open
status: resolved
opened: 2026-10-03
located-in:
- mesh-controller
- mesh-catalog
- mesh-host
fixed-by:
- mesh-host#86
- mesh-controller#252
- mesh-controller#253
- mesh-host#88
amended-design:
---
@@ -50,3 +54,27 @@ of a plan today.
Located only by owner; the move is a change of the controller's module and its deployment, not of
its code.
## Fix prepared (2026-10-04)
Three changes. mesh-host#86, merged: a process may name the container it `replaces`, and the host
removes that container only once the process has stayed up across two checks. mesh-controller#252,
awaiting the operator's merge: the controller's composition for a service process, and two
controllers safe together for the handover — the second stands by on the controller's consumers
until the first lets go, and all plan work holds one advisory lock. mesh-controller#253, held: the
controller's manifest as a Go bundle and a process. It waits on
[issue 223](../223-a-new-mesh-installs-its-controller-as-a-container/00-report.md), because with it a
new mesh cannot be installed.
## Resolved
Proven 2026-10-04 on the control machine: after the manifest change was merged and pushed, the host
created the controller's account, ran its preparation step, started the controller as a process,
found it up across both checks and removed the container. The controller now runs as its own account
from its bundle, no controller container remains, its seat answered throughout, and the merge's own
plan finished all three tiers under the new process.
One fault on the way, fixed before it could leave two controllers or none: the host read the account
not existing yet as a user database that did not answer — it matched the exit as text in a wording
its own runner did not use — so the first apply stopped at the account and the container kept serving,
which is the handover's safe failure (mesh-host#88).
@@ -1,9 +1,10 @@
---
status: located
status: resolved
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by:
- mesh-controller#246
amended-design:
---
@@ -34,3 +35,7 @@ Owner mesh-controller (the planner). **Fix direction:** on start, and whenever a
build, the plan settles an `asked` build against the build records — a build recorded as built from
the plan's commit is that tier's outcome — so a plan resumes after the controller replaced itself.
A test: a plan whose build outcome was recorded while no controller followed it resumes on start.
## Resolved
A plan settles an asked build from the build records, whoever heard the outcome. Proven 2026-10-03 and 2026-10-04: both controller merge plans since the fix finished all three rounds, including the round that replaced the controller, without being stopped by hand.
@@ -1,9 +1,10 @@
---
status: located
status: resolved
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by:
- mesh-controller#246
amended-design:
---
@@ -33,3 +34,9 @@ Owner mesh-controller. **Fix direction:** a build asked at a commit does not cha
module follows; or, if pinning is meant, the pin is said — in `module list`, in `status`, and by a
merge's plan naming the module it leaves out and why. A test: building a module at a commit and then
merging a change to it plans it.
## Resolved
Proven 2026-10-04: the catalogue merges since the fix rebuilt the module that had been pinned at an
old commit, at the merge's commit, by an ordinary plan — the same as every other module of the
catalogue. Nothing was asked for it by hand.
@@ -1,9 +1,10 @@
---
status: located
status: resolved
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by:
- mesh-controller#246
amended-design:
---
@@ -28,3 +29,7 @@ Owner mesh-controller (the catalogue's registration check). **Fix direction:** a
loads, runs or unpacks — no `loads`, no `tools` list on its module, no resource naming it — is refused
at registration, naming the field that would deliver it. A test: such a manifest is refused; adding
`loads` admits it.
## Resolved
A bundle nothing would deliver is refused at registration, by name. Proven by the controller's tests; the catalogue's bundles all name what the runtime loads (mesh-catalog#243).
@@ -1,9 +1,11 @@
---
status: located
status: resolved
opened: 2026-10-03
located-in:
- mesh-tools
fixed-by:
- mesh-tools#42
- mesh-tools#43
amended-design:
---
@@ -50,3 +52,7 @@ announcement subscriptions are allowed for every principal kind.
Recovered on the day without the forge's API: the toolchain images built by the controller straight
from the fix branch, every module image rebuilt on them, and the machine pushed.
## Resolved
A refused announcement is logged and the runtime serves on, and every runtime subscribes only the discovery subjects its grants allow. Proven 2026-10-04: after rolling out to every machine, no container restarts anywhere, and the discovery console lists no runtime as not answering. A refused tool subscription was made non-fatal the same way afterwards, under issue 218 (mesh-tools#46).
@@ -1,9 +1,13 @@
---
status: located
status: resolved
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by: mesh-controller#248
- mesh-tools
fixed-by:
- mesh-controller#248
- mesh-tools#45
- mesh-tools#46
amended-design:
---
@@ -63,3 +67,26 @@ and memberships come from the same list, so they cannot disagree.
seats and loses the recorded mesh seat. Live, the discovery console's overview must show each
mesh-scoped seat announced from exactly the holder the records name. Status moves to `resolved` once
that holds after the fix is rolled out.
## A second cause, and what the rollout broke (2026-10-04)
With the grants corrected, calls to the store reached only the holder, yet the console still showed
the seat announced from both machines. The module's runtime added every seat its start-up credential
claims, even after the mesh issued a membership that left the seat out. Once a membership exists,
it now alone decides which seat verbs a runtime serves (mesh-tools#45).
The rollout then exposed a third fault. The module on the machine that does not hold the seat was
still running an image built before #45, so it subscribed to the seat's subject. The corrected grants
refused that subscription, and the refusal ended the process. Its runtime crash-looped until the
module was rebuilt on the new runtime image. The database itself kept running. A refused tool
subscription is now logged and costs only that subject (mesh-tools#46), as a refused announcement
already did ([issue 217](../217-a-refused-announcement-took-down-every-containers-runtime/00-report.md)).
The module was not rebuilt by the plan that rebuilt the runtime image. This is the ordering gap of
[issue 211](../211-a-bundle-is-built-before-the-toolchain-it-is-compiled-in/00-report.md) seen
from a container module.
**Proven 2026-10-04.** The discovery console's overview shows the store seat announced from the
recorded holder only. Repeated calls to the store are answered by that machine, and the answers
include the controller's own database. The non-holder still answers its own module tools. No
container restarts on any machine.
@@ -0,0 +1,48 @@
---
status: resolved
opened: 2026-10-04
located-in:
- mesh-controller
fixed-by:
- mesh-controller#249
amended-design:
---
# 219 — An older build that finishes later replaces a newer one
## What was observed
2026-10-03. Two merges to the runtime module came minutes apart. Each plan asked for every module
built on the runtime image to be rebuilt. One module's two builds were of the same source and
differed only in the runtime image they stood on:
| Build | Requested | Finished | Stood on |
|---|---|---|---|
| asked by the first plan | 21:33 | 22:04 | the runtime image before the fix |
| asked by the second plan | 21:49 | 21:56 | the runtime image with the fix |
The older request finished last, and its image became the module's current artifact. The next push
deployed it, and the module's runtime crash-looped on a fault the newer image had already fixed
([issue 218](../218-a-mesh-seat-is-answered-by-a-module-that-does-not-hold-it/00-report.md)).
## Why it matters beyond this instance
Which build is current should follow what it was built from, not which build machine was slowest.
Whenever two plans overlap, which happens on any busy evening, a fix can be silently reverted by a
build that started before it existed. Every passing check still passes.
## Where to look
How a finished build is recorded and how a module's current artifact is chosen. **How it is checked:**
a test in which an older request completes after a newer one for the same artifact, and the newer
stays current.
## Resolved
A build is ordered by when it was requested, read from the id the controller gives it, and a
module's registered manifest is replaced only by a build requested at or after the one it came from.
An older request finishing later is recorded and changes nothing; a plan settles only from builds it
asked for itself. **How it is checked:** store-backed tests replay the incident — the newer request
stays what the module is — and fail without the fix. Live since 2026-10-04: the rebuilds of every
runtime-image module and the three waves of module code moves since then each registered the build
they asked for.
@@ -0,0 +1,41 @@
---
status: resolved
opened: 2026-10-04
located-in:
- mesh-host
fixed-by:
- mesh-host#84
amended-design:
---
# 220 — A delivered bundle keeps the files of the one before
## What was observed
2026-10-04. A tools bundle was rebuilt as one self-contained file per entrypoint
([ADR 0193](../../02-DECISIONS/0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md)),
so it no longer carries a package directory. On the machine it was delivered to, its directory still
held the package directory and a compiled file from the earlier delivery, dated hours before the
new files. The new files were written over the old directory, and nothing removed what the new
bundle no longer contains.
## Why it matters beyond this instance
A bundle on disk should be exactly the artifact that was built. Leftover files can be imported by
code that should no longer find them. A fix that removes a file then works on a fresh machine and
fails on every machine that ran an earlier version. It also makes "what runs here" impossible to
read from the artifact.
## Where to look
How the host unpacks a bundle into its directory. **How it is checked:** deliver a bundle, then a
version without one of its files, and the file is gone.
## Resolved
An archive is unpacked into a fresh directory beside the old one and swapped in by rename; a refused
or failed unpack leaves the old tree whole. **How it is checked:** a second delivery without a file
removes it, and nothing is left beside the directory; both tests fail without the fix. Proven
2026-10-04: a bundle rebuilt and delivered after the fix holds exactly the new build and nothing
beside it. A bundle that has not changed since keeps its old leftovers until its next version, by
design: an unchanged archive is not unpacked again.
@@ -0,0 +1,52 @@
---
status: resolved
opened: 2026-10-04
located-in:
- mesh-controller
fixed-by:
- policy: `upgrade build-agent roll-out` (2026-10-04)
amended-design:
---
# 221 — A build machine learns a new builder only from a push
## What was observed
2026-10-04. A controller merge changed how TypeScript bundles are built: one file per entrypoint
instead of a package directory. Its plan finished, and six bundles were rebuilt right after. They
came out in the old shape, because the build machines still ran the previous builder. They got the
new one only from the next push. Rebuilt after that push, the same six came out right.
The same order showed in the bus grants the same night. A push sent while the controller was still
the previous build composed grants with the previous code. A second push was needed after the new
controller had started.
## Why it matters beyond this instance
A merge to the controller is not in effect when its plan says done. Anything built or pushed in the
gap uses the old code and looks current. Today only the operator knows to push first and build
second, and even the operator forgot.
## Where to look
Whether a controller plan should end by delivering itself to the build machines and the control
machine, or whether a build should refuse a builder older than the controller that asked for it.
**How it is checked:** after a controller merge, a build asked right after its plan finishes runs
the new builder.
## Located
The mechanism existed: a module whose upgrade policy is `roll-out` is sent to its machines when its
tier is built, and the plan's next tier waits until it is applied. The build agent's policy was
`record` — built, never sent — as was that of 76 other modules, which is why every rollout on
2026-10-03 needed a push by hand. The build agent was set to `roll-out`, one machine at a time,
stopping at the first failure. **How it is checked:** the next controller merge's plan sends the
build agent to the build machines before its last tier, and a build asked right after the plan
finishes runs the new builder; then this moves to resolved. Whether the other modules roll out is
the operator's policy, not this issue's.
## Resolved
Proven 2026-10-04: the next controller merge's plan logged that its first tier was built and the
build agent sent to all four build machines, and only then asked its next tier. The builder change
in that merge reached the build machines without a hand push.
@@ -0,0 +1,51 @@
---
status: resolved
opened: 2026-10-04
located-in:
- mesh-tools
fixed-by:
- mesh-tools#47
amended-design:
---
# 222 — An assignment is refused on the bus until the bus's machine is pushed
## What was observed
2026-10-04. A module was assigned to the laptop, and the laptop alone was pushed. The module's two
bundles arrived and the node's runtime launched both, and then the bus refused every one of the
module's subjects:
```
nats: permissions violation: Permissions Violation for Subscription to "mesh.mod.<module>.tool.<tool>.<node>"
```
The tools were unreachable until a later push that included the machine running the bus. Then the
runtime served them without a restart, because a new membership arrived and it re-subscribed.
## Why it matters beyond this instance
What an account may answer lives in the bus's user list, and the controller writes that list only
into the declaration of the machine that runs the bus. Assigning anything to any machine changes
that list, so `push <machine>` after `assign <machine> <module>`, which is what the controller itself
tells the operator to run, leaves the module running and unreachable, with nothing reporting a fault.
## Where to look
Whether a push to one machine should also send the bus's machine when the user list it would
compose differs from the one that machine holds. **How it is checked:** assign a module with tools
to a machine that does not run the bus, push only that machine, and its tools answer.
## Diagnosed and resolved
The push did send the bus's machine: the controller's log shows both machines applying in the same
second. The fault was the order within that second. The node's runtime subscribed before the bus
had reloaded its user list, the bus refused, and the bus client marks a refused subscription dead.
Nothing asked again until a later membership happened to re-serve the module.
A subject the runtime answers on is now asked for again when the bus refuses it, after waits from
two seconds to two minutes, and given up and said after about five minutes. **How it is checked:**
a test refuses a subject and finds it asked for again and answering, given up past its attempts,
and not asked again once stopped; and by hand against a bus whose permissions were reloaded while
connected, the runtime answered two seconds after the grant arrived, where the runtime before the
fix never answered.
@@ -0,0 +1,60 @@
---
status: resolved
opened: 2026-10-04
located-in:
- mesh-host
fixed-by:
- mesh-host#87
amended-design:
---
# 223 — A new mesh installs its controller as a container
## What was observed
2026-10-04. Preparing [issue 213](../213-the-controller-is-a-go-program-run-in-a-container/00-report.md),
the controller's own manifest was changed to a Go bundle run as a process. The installer that raises a
new mesh assumes the controller is an image and a container at every step from its third on:
- it requires the controller's build to produce exactly one image;
- it starts a temporary controller from that image, and publishes the image to the registry;
- at the pivot, it finds the controller's container in the declaration, reads its environment and
volumes, waits for it, and from then on talks to the controller only through the container.
## Why it matters beyond this instance
With the controller's manifest changed, a new mesh cannot be installed: the pivot fails. The deeper
constraint is ordering. A process's bundle is fetched from the artifact store, and the installer
raises the artifact store only after the pivot, so the controller's first declaration names a bundle
nothing can serve yet.
## What a fix has to settle
One of two shapes, and it is a decision, not a repair:
1. raise the artifact store before the pivot, publish the controller's bundle to it, and talk to the
controller from the host's side rather than through a container; or
2. pivot to the image form as today, and let the first push hand over to the process, which
requires the controller's manifest to carry both forms.
Until it is settled, the change of the controller's manifest (mesh-controller#253) is held. The
handover itself is built and merged (mesh-host#86); the controller's half (mesh-controller#252) waits
on the operator. **How it is checked:** the installer's test raises a mesh whose controller manifest
is the process form, and the controller answers its seat's verbs at the end.
## Decided (2026-10-04)
Option 2, [ADR 0200](../../02-DECISIONS/0200-genesis-pivots-to-the-controller-as-a-container-and-the-first-push-hands-it-to-a-process.md):
genesis pivots to the controller as a container recorded under the name the process `replaces`, and
the first push hands it over.
## Resolved
Built as [ADR 0200](../../02-DECISIONS/0200-genesis-pivots-to-the-controller-as-a-container-and-the-first-push-hands-it-to-a-process.md)
decided: genesis reads only the controller's process from the manifest, builds the controller's image
from its repository, and pivots to a container of its own shape recorded under the id the process
`replaces`; an older controller in the image form still builds as before. **How it is checked:** the
installer's tests build the genesis form from the controller's real manifest, and apply the
controller's first process declaration over the recorded container with the host's own apply — the
process starts, the container is removed, one controller remains. A real install from nothing has not
been run since; the handover it ends in was proven live on the running mesh (issue 213).
@@ -0,0 +1,64 @@
---
status: open
opened: 2026-10-04
located-in: [mesh-host internal/apply/apply.go, mesh-host internal/apply/schedule.go]
fixed-by:
amended-design:
---
# 224 — An apply arriving during a maintenance window reopens it, by recreating the container the window is holding still
## What was observed
Reviewing [ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md)'s
`while-stopped` before merging it, 2026-10-04. Found by reading, not by running.
A scheduled step may hold its module's containers still while it runs. The host stops them, runs
the step, starts them again. Nothing tells the **apply** that a window is open, and the apply's
rule for a container it finds stopped is to replace it:
```
case existed && (before.Spec == want || legacy) && before.Running && len(reasons) == 0:
out.Action = "unchanged"
case existed:
rm -f
```
`before.Running` is false for a container a window is holding, so the second branch takes it:
the container is removed and recreated, **running**, in the middle of the step that required it
to be still.
## Why it matters
For the store, which is what the field was built for, the chain is: a push lands at 03:30 → the
apply recreates the registry → the registry accepts an upload from a build running at the same
time → `garbage-collect`, already past its mark phase, sweeps the blob that upload just wrote.
The image is then in the store with a layer missing, and the build that made it reported success.
Two things have to coincide, so it is not likely. It is also not rare enough to leave unsaid: the
mesh pushes on every merge, at any hour, and a collection over a store this size is minutes rather
than seconds.
**The general shape is the one that matters.** `while-stopped` is the first thing in the mesh that
makes a container's stopped state *intentional*. Everything else in the host reads "stopped" as
"broken, fix it", which is right everywhere else and wrong here. Any future use of the field
inherits this.
## What this is not
Not a regression. The store has never collected anything, so nothing is worse than it was; this
is a hole in something new rather than something that broke.
## Open questions
- **Should a window take the apply lock?** The daemon already serialises applies with `applying`
and, across processes, with `store.Lock`. A window that held it would make the race impossible.
The cost is that a push arriving mid-window waits for minutes, and a push that waits is what
[issue 185](../185-a-refused-membership-publish-stops-the-controller/00-report.md)'s
family of outages looked like from outside.
- **Or should the apply learn that a container is held?** Narrower: the scheduler says which
containers a window currently holds, and `applyContainer` reports those unchanged instead of
recreating them. Nothing blocks, and the apply tells the truth for the minutes it matters —
at the cost of a second source for "is this container meant to be running".
- Either way: should the *report* say a window is open, so a machine that looks half-stopped at
03:31 reads as working rather than broken?
@@ -0,0 +1,75 @@
---
status: open
opened: 2026-10-04
located-in: [mesh-host internal/apply, mesh-controller internal/catalogue]
fixed-by:
amended-design:
---
# 225 — A provisioner cannot read the grant secrets since its code left the container, and every consumer of it is unserved
## What was observed
On the control machine, 2026-10-04, found while looking at why one app was restarting:
```
[mongodb] [provisioner:mongodb-database] mesh_novox_photos: secret not readable yet
(/var/lib/mongodb/grants/novox.photos.secret):
Error: EACCES: permission denied, open '/var/lib/mongodb/grants/novox.photos.secret'
```
**4330 times, every five seconds, since 01:30:20.** The consequence is not a log line: the
provisioner never reads the password, so it never creates the user, so the consumer never
connects —
```
UserNotFound: Could not find user "mesh_novox_photos" for db "admin"
```
— and the app crash-loops. Two consumers on this machine are in that state.
## Why
The grant secrets are what the mesh seals for each consumer and the host unseals beside the
provider's contributions file. They are written `-rw------- root root`, which was right while a
module's own code ran in a container as root.
[ADR 0198](../../02-DECISIONS/0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)
moved a module's long-running code out of its container and under the node's runtime, which runs
as the operator's account. The provisioner is now that account; the secret is still root's. The
timestamps say it exactly: the files are dated 2026-09-26, the first refusal is 01:30:20 on the
day the runtime rolled.
**Nothing reports it.** The machine applies cleanly and reads as current; the provisioner says
`secret not readable yet`, whose wording is for a real and ordinary race on the first pass — the
host has not written the file yet — and which is indistinguishable, in the log, from a permanent
refusal. Four thousand occurrences of a message that means "wait a moment" is the shape to
recognise.
## Why it matters beyond this instance
This is every provider that provisions. The grant secret is the one file the sdk's harness reads
for every consumer, so a provider that cannot read it serves nobody — and says so only in a line
that reads like patience.
It is also the general question the runtime move leaves: **what the mesh seals for a module is
owned for the shape that module's code used to have.** Each module whose code moved is a module
whose files may now be unreadable to it, and ownership is the mesh's to state, not the module's
to work around.
## What this is not
Not caused by [ADR 0202](../../02-DECISIONS/0202-a-provider-declares-what-it-derives-for-each-consumer.md)
or [ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md), which landed two
to three hours after the first refusal. Those rebuilt the two affected consumers, which recreated
their containers and made a silent fault visible as a restarting one. The dates are above.
## Open questions
- Who owns a grant secret now — the module's account, as `secrets-owner` already says for a
module's own secrets? Then the host writes it so, and this is a one-line statement in the
declaration rather than a convention.
- Should `secret not readable yet` stop saying "yet" after the first few passes? A message that
is right once and wrong four thousand times is a message that hides its own meaning.
- Which other modules' files did the runtime move leave behind? The sweep is the same question
for every path the mesh writes for a module: directories, bundles, received files.
@@ -0,0 +1,58 @@
---
status: open
opened: 2026-10-04
located-in: [mesh-controller cmd/mesh-controller/collect.go, mesh-controller internal/inventory/collection.go]
fixed-by:
amended-design:
---
# 226 — The store's sweep stops at the first reference recorded with an address, so it collects nothing at all
## What was observed
The first live run of [ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md)'s
sweep, 2026-10-04, printed on every build:
```
the artifact store kept 127.0.0.1:5100/mesh-tools/build@sha256:0de48cd3…, so nothing more was
asked of it: 127.0.0.1:5100/mesh-tools/build@sha256:0de48cd3… is not a reference into the
mesh's artifact store
1681 more to collect; the next build asks again
```
Nothing is collected, and nothing ever will be. The store holds 1681 artifacts the mesh no longer
keeps and the feature that exists to remove them is inert.
## Why
Two correct decisions meeting badly.
**A reference recorded before references were kept without an address** is
`127.0.0.1:5100/<path>@sha256:…` rather than `artifact-store://<path>@sha256:…`
([04-ISSUES/102](../102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md)).
`LetGo` rightly refuses to compose a delete for a reference whose shape it does not recognise —
that refusal is what keeps the sweep from reaching something that is not the mesh's.
**The sweep stops at the first refusal**, because "a store that refuses one refuses all of them"
— deletion disabled, the store down, the network gone — and pushing through would mean a hundred
identical failures in front of whoever was building something. That reasoning is right for the
store refusing. It is wrong for *this* record being unreadable.
So one old record, early in the oldest-first order, halts the whole sweep for ever.
## Why it matters beyond this instance
**A guard that cannot tell "I will not ask about this" from "it would not answer" stops the wrong
amount of work.** The two deserve opposite responses: skip one, abandon the other. Collapsing
them into "an error" is how a bounded, cautious loop becomes a loop that does nothing — and it
reports the right number while doing it, which is what made it look healthy.
## What a fix has to settle
- A reference the sweep cannot address is **skipped, and the sweep goes on** — it is a fact about
that record, not about the store.
- `Recorded()` already normalises the old form to the kept one, and is what the rest of the mesh
uses for exactly these references. The sweep should normalise before asking rather than refuse.
- Only a refusal *by the store* ends a sweep.
- **How it is checked:** a sweep over records holding one address-recorded reference and one kept
one collects the second; a sweep against a store that refuses stops at the first.
@@ -0,0 +1,44 @@
---
status: open
opened: 2026-10-04
located-in: [mesh-catalog modules/photos]
fixed-by:
amended-design:
---
# 227 — The photo app's admin client asks for port 80, which the reverse proxy holds, so it cannot start
## What was observed
Applying the control machine, 2026-10-04:
```
applying "photos.admin-client": starting container photos-admin-client:
failed to bind host port 0.0.0.0:80/tcp: address already in use
```
Port 80 on that machine belongs to the reverse proxy (`mesh-route-proxy`, confirmed with `ss`),
which is the whole arrangement: the proxy holds the public ports and every module is reached
through it. A module that publishes 80 itself can never start beside it.
Everything else on the machine applied; this one resource fails every pass.
## How it surfaced
`photos` had been pinned at a commit from 2026-09-28 and was rebuilt to `main` on 2026-10-04 —
forced by [ADR 0202](../../02-DECISIONS/0202-a-provider-declares-what-it-derives-for-each-consumer.md)'s
refusal of its transcribed bucket name. The admin client is one of the changes that came with the
rest of `main`. The rebuild did not create the conflict; it delivered it.
**A module pinned months behind carries whatever its branch gained, all at once, the first time
something makes it move.** That is the cost of a pin, and it is paid in full rather than
gradually.
## What a fix has to settle
- Which port the admin client should ask for, or whether it should be reached through the proxy
like everything else and publish nothing.
- Whether a module declaring a port the machine's proxy already holds should be refused when it
is composed, rather than failing on the machine every pass. The mesh assigns ports
([ADR 0038](../../02-DECISIONS/0038-the-mesh-assigns-the-port.md)); a fixed 80 beside a proxy is
a statement it could check.