Merge main: the bus design, issues 113/114/117/120 and research 017 landed

# Conflicts:
#	04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md
This commit is contained in:
jochen
2026-09-26 14:16:05 +02:00
10 changed files with 965 additions and 3 deletions
@@ -0,0 +1,47 @@
---
status: active
initiated: 2026-09-26
touches:
- 00-META/mission.md
- 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0010-delivery.md
- 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
- 03-DESIGN/01-to-be/06-the-controller.md
- 03-DESIGN/01-to-be/09-the-node-lifecycle.md
- 03-DESIGN/00-as-is/09-interfaces-and-observability.md
---
# 017 — A mesh that heals itself
**What.** The behaviour the operator wants: a mesh that runs itself. It notices what is wrong,
repairs what it can, and hands what it cannot repair to someone who can, with the reason. This effort
writes that wish down as intended behaviour, designed for the bus the mesh is moving to
([ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md): NATS). It also records what can be done
pragmatically before that move.
**Why.** The mission is *a mesh that controls itself* ([mission](../../00-META/mission.md)). The
mesh can tell whether it is up. It cannot tell whether it is right. The as-is page on observability says so
([as-is 09](../../03-DESIGN/00-as-is/09-interfaces-and-observability.md)). To-be 06 names an
`observability` context in the controller and leaves its store undecided. Nothing routes a condition
the mesh cannot fix to anyone. The cost is measurable: **46 of the 116 issue reports in this
repository describe a failure that was silent.** A mesh that heals itself is, first, a mesh that stops
failing silently.
**What it touches.** The controller's observability context, the node lifecycle's liveness, delivery
([ADR 0010](../../02-DECISIONS/0010-delivery.md), [ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)),
the provisioner harness, and rotation, which is proposed alongside to-be 27 as ADR 0114.
**Documents.**
- [01 — The intended behaviour](01-the-intended-behaviour.md): the wish, as principles and as how
the mesh behaves once the bus is NATS.
- [02 — Now, pragmatically](02-now-pragmatically.md): what is done before NATS, why it does not
build anything the move would throw away, and what has been done already.
**Next.** Two measurements this effort owes before it can graduate:
1. **Every loop in the mesh**: what it converges, and whether it compares against observed state or
against its own memory. Issue 120 found the provisioner harness trusting memory. The same pattern is
expected elsewhere.
2. **The 46 silent failures, classified**: a missing observation, a loop trusting memory, or a missing
escalation. That shows which mechanism removes the most of them.
@@ -0,0 +1,85 @@
# 01 — The intended behaviour
The operator's wish, written as behaviour: what a person or an agent sees the mesh doing. This is a
target to design toward, not a design. Every part of it is to be decided through a record before it
is built.
## Principles
**1. Every loop compares what should be with what is, never with what it did.** Desired state is the
mesh's: assignments, requirements, seats. Observed state is read from the thing itself: the container,
the backend, the node. A loop that compares against its own memory of what it applied is blind to
anything that changed behind its back. That is issue 120, and it is the pattern this whole effort is
written against.
**2. Healing is the ordinary path run again, never a second path.** Repairing a lost login is
provisioning it. Repairing a dead container is converging the node. Repairing a stale declaration is
delivering it. A repair that needs its own code is a second way of doing something, which is exactly
what the mesh is removing everywhere else.
**3. A repair never destroys.** Healing may recreate, re-provision, re-deliver and restart. It may
never delete a consumer's data, retire a credential someone still uses, or pick a winner between two
contradictory states. Where the only repair is destructive, it is escalated.
**4. Nothing fails silently.** Every condition the mesh cannot repair within its budget becomes
visible. It is named, it says since when, why, and who can resolve it. It is visible until it is
resolved, and resolved by observation, not by someone clicking it away.
**5. What the mesh cannot fix goes to an agent.** Per the mission, an agent may be human or not. A
condition that needs judgement is handed to one, as work, with what the mesh knows. It is not handed
over as a notification that someone may or may not read.
**6. Correctness, not only liveness.** A running process that authenticates with a dead credential,
serves an old version, or routes nowhere is not healthy. What a provision's contract promises is what
is checked: the credential authenticates, the route answers, the version is the declared one.
## The loop, everywhere
Every part of the mesh that owns something runs the same loop:
1. **know** what should be true: from assignments, requirements and seats;
2. **observe** what is true: from the thing itself, on its own cadence;
3. **repair** the difference by running the ordinary path again, within a budget of attempts and
time;
4. **raise** a *condition* when the budget is spent or the only repair is destructive;
5. **clear** the condition when observation shows it resolved.
A **condition** is a durable fact about something the mesh owns, such as a node, an assignment, a
provision, a seat or a rotation: what is wrong, since when, the evidence, what was tried, and who can
resolve it. Conditions are the one thing a person or an agent looks at to know whether the mesh is
right. `status` is the list of open conditions. When it is empty, the mesh is right, not just up.
## On NATS
[ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md) moves the bus to NATS, and NATS makes most of
this cheaper, because observation becomes something every component publishes rather than something
a central process polls.
| the wish needs | on NATS |
|---|---|
| every component says it is alive | a heartbeat on a subject per node and assignment; silence past its interval is a condition, and nobody polls |
| every component says what it observed | observations published on subjects (`mesh.observed.<node>.<assignment>`, for instance), consumed by whoever owns the comparison |
| the last known state survives restarts | a JetStream key-value bucket of observed state per owner; the provisioner's "what I applied" and a rotation's step live there, not in memory |
| conditions are durable and watchable | conditions as entries in a key-value bucket, watched by anyone who cares: a surface, an agent, the controller |
| the bus itself is observed | the server's advisories (a consumer exceeding its deliveries, a slow consumer, a client disconnecting) and its monitoring endpoint become observations like any other |
| a repair is retried, not lost | JetStream redelivery with delay, which is the same mechanism [ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)'s guarantee moves to |
| work handed to an agent | a condition that needs judgement published as a task on a subject an agent's queue group consumes |
**Who compares.** Each owner compares its own: the host for its node's containers and files, a
provisioner for its backend, the vault for rotations, the controller for delivery and seats. The
controller's observability context does not repair anything. It holds conditions, their history,
and the view across the mesh. It notices what no owner can see about itself: an owner gone silent.
## What stays human
Some repairs need the operator's key, and the mesh says so rather than pretending otherwise:
re-raising the vault or the broker, and recovering a node's identity. These are conditions too, with
the procedure named, and they are the only ones that can never clear themselves.
## Open
- The budgets: how many attempts, over how long, per kind of repair.
- How a condition that needs judgement reaches an agent, and how the agent's action is recorded.
- Where the observability context stores history (to-be 06 left it open; volume argues against the
relational store).
- Which correctness probe each provision's contract offers, and how often it runs.
@@ -0,0 +1,46 @@
# 02 — Now, pragmatically
The intended behaviour lands on NATS. The bus moves after the migration's core
([ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md)). Until then, work toward it is chosen by one
test:
**Does it survive the move?** A change to what a loop compares against, or to what an adapter can
tell about its backend, survives, because it is independent of the bus. A new AMQP queue for health
reports, a poller written against the broker's management API, or a condition store built on the
current broker does not survive, and is not built.
## Done
**The provisioner asks the backend, not memory** (issue 120). The SDK's harness gained an optional
`holds` on the adapter, asked for every applied consumer every minute. A consumer the backend no
longer holds is provisioned again. Being unable to ask is not treated as loss. The cache module
implements it first, because its server keeps its users in memory and forgets them all on a restart.
That was verified against a real server: a restart erases every consumer's user, and `holds` answers
correctly for absent, present, wrong-password, disabled and deleted.
Changes: mesh-sdk PR #7 (0.1.1) and mesh-catalog PR #84.
This is principle 1 applied to one loop. It survives the move unchanged. On NATS, the harness's
record of what it applied moves from memory into a key-value bucket, and `holds` stays as it is.
## Next, in order of silent failures removed
1. **`holds` for the other credential providers.** Each backend can answer whether a login exists
with the mesh's password without changing anything. Where a backend cannot check a password without
logging in, logging in is the check.
2. **The harness's other blind spot.** A consumer that goes away while its provisioner is down is never
removed. The fix is the same principle in reverse: list what the backend holds, and compare it with
what the mesh asks for. Removal stays subject to ADR 0114's rule that it never follows from a login
changing.
3. **`status` reports what owners already know.** The host knows which containers it recreated and
why. A rotation knows who it waits on. Delivery knows what is outstanding. Surfacing those as
conditions in the existing `status` needs no new transport. It is the shape the NATS condition store
will hold.
4. **The loop inventory and the classification** in [00](00-overview.md). They decide what comes after
these three.
## Not now
- Heartbeats, observation subjects, key-value state, advisories: all NATS, all after the move.
- Handing conditions to agents: designed with NATS, where a task on a subject is native.
- Choosing the observability store: decided when there is something to store, which is after the
move.
+312
View File
@@ -0,0 +1,312 @@
---
layer: to-be
status: proposed
code:
- mesh-controller internal/link (to be replaced)
- mesh-host internal/link (to be replaced)
- mesh-tools src/broker-amqp.ts (to be replaced)
- mesh-catalog modules/nats (to be written)
updated: 2026-09-24
decisions:
- 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md
- 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
- 02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md
---
# 25. The bus on NATS
**Status: proposed — a design to be reviewed before any code.** This is the architecture
[ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md) asks for. It says what rides the bus,
under which subject, with which guarantee, under whose account; how a node joins; how a person
reaches a tool; how the mesh moves from the bus it has to this one; and how each claim is checked.
Prose and diagrams only; no configuration is pasted.
## 1. What the bus is for
The bus carries five kinds of traffic today, and this design keeps the five, renaming nothing a
module can see:
| Traffic | Today | Guarantee it needs |
|---|---|---|
| **control** — a node's report, its heartbeat, a build's outcome, an enrolment | queues `control`, `.upgrades`, `.catchup` | nothing lost while the store restarts; retried; in order per node |
| **declarations** — the controller tells a node what to be | queue `node.<name>` | the node gets the newest; a stale one is never applied |
| **builds** — the controller asks the build machine to build | queue `builds` | at least once, one builder at a time |
| **events** — a module says something happened | topic exchange `mesh.events`, keys `<module>.<event>` | delivered to every consumer that declared it; dead-lettered when it cannot be |
| **tools** — one module or person asks another's tool a question | exchange `mesh.rpc`, per-tool service queues `serve.<module>.<tool>` | one answer, from one server, or a timeout |
The sdk's contract — `request`, `handle`, `publish`, `subscribe`, `close` — is the whole surface a
module sees, and it does not change ([ADR 0039](../../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md)).
## 2. Subjects
NATS addresses everything by subject. The mesh's subject space is one tree, and every account's
permissions are expressed as which branches of it that account may publish to and subscribe from.
```
mesh.control.<node>.report a node's report (JetStream: CONTROL)
mesh.control.<node>.alive heartbeat (core, no persistence)
mesh.control.enrol an enrolment request (JetStream: CONTROL)
mesh.control.built a build's outcome (JetStream: CONTROL)
mesh.node.<node>.declare a declaration for a node (JetStream: NODES, last-per-subject)
mesh.build.request work for the build machine (JetStream: BUILDS, work queue)
mesh.events.<module>.<event> an event (JetStream: EVENTS)
mesh.tools.<module>.<tool> a tool invocation (core request/reply)
mesh.ask.<node>.<command> the controller's command api (core request/reply)
```
Two things this buys over the exchanges: **request/reply is native** — a tool call is one
`request` on `mesh.tools.<module>.<tool>` answered by whichever runtime serves it (a queue group per
tool, so several nodes may serve one tool); and **a declaration is last-per-subject** — the NODES
stream keeps only the newest message on `mesh.node.<node>.declare`, so a node that was away gets
exactly the current declaration and nothing older. That is the wire-level answer to
[issue 107](../../04-ISSUES/107-a-declaration-carries-no-order/00-report.md): the stream's sequence
*is* the order, and a node that sees sequence n refuses n−1 by construction.
**A reply-to travelling through a JetStream stream is carried in the payload, never in the
transport `Reply` field.** Revision, first review: core NATS request/reply sets the requester's
ephemeral inbox as the message's `Reply` field, and a plain responder answers it directly — but a
message a JetStream consumer delivers has already had that field claimed for the consumer's own
ack address (`$JS.ACK.<stream>.<consumer>...`), so by the time the controller (§3's CONTROL
consumer) sees the message, `Reply` names where *it* must ack, not where the original caller is
waiting. `mesh.control.enrol` is the case that matters: a synchronous-feeling caller waiting on an
ephemeral inbox, over a subject the store-window guarantee may legitimately delay by several
`nak` cycles — exactly the combination that would otherwise deliver the answer to a caller who
has long since timed out and unsubscribed. So every CONTROL message that expects an answer states
its reply subject as an ordinary field of its own payload; the controller reads it from there and
publishes the answer to it explicitly, never via `Respond()`. Nothing else in this design routes
a reply through a stream — tools and heartbeats stay on core NATS, where `Reply` means what it has
always meant.
## 3. Streams, and the guarantees they carry
Core NATS is at-most-once. Everything the mesh must not lose lives in a JetStream stream:
| Stream | Subjects | Retention | Why |
|---|---|---|---|
| CONTROL | `mesh.control.>` except `alive` | work queue, one consumer (the controller), explicit ack | the store-window guarantee ([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)): the controller `nak`s with a delay while its store is away and the message is redelivered; nothing is dropped |
| NODES | `mesh.node.>` | last per subject | one declaration per node, always the newest |
| BUILDS | `mesh.build.>` | work queue, explicit ack | at least once; a builder that dies mid-build has its message redelivered |
| EVENTS | `mesh.events.>` | limits (age, size), durable consumer per subscribing module | a subscriber that was down catches up; after `max-deliver` attempts the advisory feeds `mesh.events.dead` (its own small stream) |
Tool calls and heartbeats stay on core NATS: a lost heartbeat is the next heartbeat; a lost tool
call is a timeout the caller already handles.
Streams and consumers are objects the controller creates at genesis and asserts on start; a module
declares nothing about them. The controller is the only writer of stream definitions.
## 4. Accounts
[ADR 0043](../../02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) says a
module's account may publish only what it `emits` and consume only what it `consumes`. NATS
expresses this exactly, per subject, and better than a vhost could:
- **One NATS account for the mesh.** Accounts in NATS isolate subject spaces entirely; the mesh
is one space, so it is one account. The predecessor's compatibility broker is not on this bus at
all.
**One account is a choice with a cost, stated plainly on revision:** none of NATS's own
isolation is free here, because there is only the one subject space, for everyone. Two
consequences that a single-account mesh must therefore grant on purpose, not by omission:
- **A durable consumer needs permission to ack, or it never really consumes.** Acking a
JetStream delivery is a publish to that consumer's own ack-reply address
(`$JS.ACK.<stream>.<consumer>.>`), a different subject from anything the consumer subscribes.
A module's user is therefore granted publish on `$JS.ACK.EVENTS.<module>.>` as well as its
emits — scoped to the one consumer name the controller derives for that module, so a module
can ack only its own deliveries. Without this, first review found, every message it receives
would be redelivered forever: refused by the permission list it already has.
- **A reply inbox needs a subject nothing else can guess or enumerate.** With one account,
inbox privacy is the permission list or it is nothing — there is no second account backing
it up. So no user is ever granted a bare `_INBOX.>`. Each user's inbox subject is derived
from its own identity (`_INBOX.<module>.<node>.>`, or `_INBOX.person.<name>.>`), and its
permissions name only that one prefix, for the reply to any request it makes and nothing
wider. First review found the account note without this and read it as "any user may
subscribe any inbox" — which was accurate against the text as it stood.
- **One user per module per node**, as today, with publish permissions
`mesh.events.<module>.<event>` for each emit, `mesh.tools.<module>.>` to serve its tools, its
own ack-reply subject for each durable consumer it holds, and its own inbox prefix; subscribe
permissions for each consumed event's subject, its tool subjects, and that same inbox prefix.
Nothing else. A module that tries to publish outside its emits is refused by the server, not by
convention.
- **The controller's user** owns `mesh.control.>`, `mesh.node.>`, `mesh.build.>` and the streams.
**A host's user** may publish its own `mesh.control.<node>.>` and subscribe its own
`mesh.node.<node>.declare` — and nothing of any other node's.
- **A person's user** (§7) is a module-shaped user with permissions on the tool subjects it may
invoke, issued and revoked by the controller like any account.
**Accounts are configuration, not API calls.** The controller composes the server's user list and
permissions into a file the host declares. **How that file reaches the running server is §5's,
not this one's** — revision, first review: an earlier draft said "reloads" and cited a precedent
that does not apply to a container (see §5). No management API, no credential travelling through a
management call, and the [issue 102](../../04-ISSUES/102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md)
discipline from the first day: an address or a permission is read where it is used, never stored
with a port. Passwords are minted and sealed exactly as today; the file holds bcrypt hashes.
Alternative considered and not taken: the operator/JWT model (`nsc`), where accounts are signed
tokens resolved by the server. It is the right model for a multi-tenant NATS; the mesh is one
tenant, already has a sealing key and a controller that writes files, and would gain a second
signing hierarchy for nothing.
## 5. The broker as a module
`nats` is a catalogue module claiming the seat `mesh-broker`
([ADR 0079](../../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md): the seat
is the server, and the server changes). It declares one container (a single binary; JetStream on a
named volume), its listening ports — client, TLS, and the monitoring endpoint on loopback — and a
configuration file the controller composes (accounts, permissions, TLS, JetStream).
**How that file's changes reach the running server, corrected on revision.** First review: the
earlier draft named `reload-on` as the mechanism, citing the container runtime's own trust file as
precedent. `reload-on` is real, but it is a **service** field
([mesh-host declaration.go](https://git.novox.be/novox/mesh-host), `Service.ReloadOn` —
`docker.service` is reloaded via systemd, which is what the cited precedent actually does). A
**container** resource has no reload field at all — only `restart-on`, and a container's
`restart-on` is documented, exactly, to mean *recreate*. Declared as the earlier draft had it,
either the field is silently meaningless on a container resource or — if read as the nearest real
equivalent — every account, permission, or key change recreates the bus's own server: every
connection dropped, every in-flight JetStream ack lost, mid-flight the moment a module is added,
reassigned, or a person's access changes. For the one resource everything else depends on, that
is not an edge case; it is the common case.
**The fix asks nothing new of the host.** `nats-server` already reloads its own configuration
live on `SIGHUP` — accounts, permissions, everything in §4 — without dropping a connection; this
is the server's own documented capability, not something built for the mesh. So the composed
configuration file is mounted into a **directory** resource, not directly — a directory's contents
are not compared for change the way [issue 103](../../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md)'s
fix made a directly-mounted file's content, so a rewritten file inside it is not, on its own, a
reason to recreate the container. The image's own entrypoint watches that one file and sends
`nats-server` its own process `SIGHUP` when it changes — self-contained, inside the module, the
same place `modules/gitea/token.ts` keeps its own state rather than asking the host to model it.
The host's only job is what it already does for any directory resource: keep the file's content
current. Nothing is declared as `reload-on` or `restart-on` for this resource at all.
Its guard is the same rule as the AMQP broker's: the monitoring port is refused from anything but
the private network. It is raised at genesis like the store, adopted as a module in the same
phase. The predecessor's AMQP broker remains a module of its own, `lavinmq-compat`, with a single
purpose and a retirement condition: no client connected for a period the operator sets.
## 6. Joining: the enrolment handshake
Unchanged in shape, changed in transport. A node that has a token connects to the bus over TLS
with the **enrolment user** — a user that may publish `mesh.control.enrol` and subscribe its own
`_INBOX.enrol.<token-id>.>` and nothing else — publishes its request (the claim of the token, its
keys, its proof, its own reply subject as §2 now requires, and the found tunnel from
[ADR 0105](../../02-DECISIONS/0105-the-mesh-adopts-the-predecessors-tunnel-in-place.md)), and waits
on that inbox. The controller spends the token, records the node, composes the node's own user
into the server's configuration, and — reading the reply subject from the request's payload, never
from the transport `Reply` field the CONTROL consumer has already claimed for its own ack —
answers with the credentials sealed to the node's sealing key. The node reconnects as itself. The
enrolment user's permissions are what make a leaked token useless for anything but enrolling: it
cannot read a declaration or hear an event, and it cannot subscribe any inbox but the one its own
token derives.
## 7. A person's client
The operator asked for the mesh's tools from a workstation, and for it designed here rather than
bridged. It is three things:
1. **A person's account**: `operator issue <name>` on the controller creates a user whose
permissions are the tool subjects it may invoke — `mesh.tools.>` for an administrator, a list
for anyone else — and nothing on control, nodes or builds. It is issued, sealed to the person's
own key, and revoked, like a module's.
2. **A client that speaks the bus**: a small program on the workstation that connects as that user
over TLS, lists tools by asking the catalogue (`mesh.tools.mesh-catalog.catalog_tools`), and
turns each tool into a call — as an MCP server for an agent, and as a command line for a person.
It uses the sdk's `Broker` contract on the NATS runtime, so it is the same code path a module's
tools use, not a second protocol.
3. **Reachability**: the workstation reaches the bus over the private network once it is a node,
or over the predecessor's tunnel before that, on the bus's port; the guard and the openings
treat the bus as they do today.
Nothing is built of this before §10's bed passes; the MCP surface is a thin adapter over (2).
## 8. What a module sees
Nothing new. `publish` on an envelope becomes a publish on `mesh.events.<module>.<event>`;
`subscribe` with a pattern becomes a durable JetStream consumer on the matching subject filter;
`request`/`handle` become a NATS request and a queue-group subscription on
`mesh.tools.<module>.<tool>`. The envelope's shape ([ADR 0042](../../02-DECISIONS/0042-the-shape-of-an-event-on-the-wire.md))
is unchanged; it is the message body. A module built today runs on the new runtime without a
rebuild — that is the test of ADR 0039, and it is in §10.
## 9. Moving from the bus the mesh has
Per ADR 0106: built beside, cut over once, after the core.
1. The `nats` module, the controller's and host's link on NATS, the runtime's client — built and
proven in the lab (§10) while the migration continues on AMQP. Modules converted meanwhile
target the sdk contract and are untouched by this.
2. The cutover is one rollout, previewed: the controller assigns `nats` to the hub (raised beside
the AMQP broker on its own ports), composes every node's and module's account into it, then
rolls out the controller, every host and every runtime built for NATS. Each node's host connects
to the new bus as it comes up and reports; the controller confirms every node heard before it
stops listening on AMQP. The predecessor's clients never notice: their broker is the
compatibility module and stays.
3. The AMQP-side mesh accounts are removed from the compatibility broker; it keeps only the
predecessor's users. The bus's port settings follow ADR 0100 like any port.
4. The compatibility broker retires when its retirement condition holds.
What is not done: no dual-bus period for the mesh's own traffic, no bridge, no module rebuilt.
## 10. How it is checked
Two lab beds, both required green before any node's bus moves.
**The bus bed** — a mesh raised on NATS from genesis:
- a node enrols over TLS with a claimed token, and the enrolment user cannot read a declaration;
- a push composes; the store is stopped; the push is held (nak with delay), the store returns, the
push applies, nothing was lost or duplicated;
- a node that was away gets exactly the newest declaration, and a replayed older one is refused
by sequence;
- an upgrade rolls out to two nodes;
- a module's tool is invoked from another node and from a person's client, each with an account
that can invoke it, and refused from one that cannot;
- a module's account cannot publish outside its `emits` nor subscribe outside its `consumes` —
refused by the server;
- a module acks a delivery from its own durable consumer, and is refused acking another module's;
- a user subscribes another module's or person's inbox prefix and is refused by the server, not
by the client's own good behaviour;
- an event whose consumer keeps failing dead-letters after `max-deliver`;
- an enrolment request held by a `nak`-with-delay cycle still reaches the enrolling node's inbox
once the controller answers — proving the reply travels in the payload and not the transport
field a consumer's ack has already claimed;
- the `nats` container is not recreated when only its composed configuration file changes, and
a change to that file is live (a new user can connect, a revoked one cannot) within one
watcher-poll interval, without a restart;
- a module built before this design serves its tools unchanged on the new runtime.
**The cutover bed** — a mesh on AMQP with a predecessor stand-in on the compatibility broker moves
its bus in one rollout; every node reports on NATS afterwards; the stand-in's client on AMQP is
still connected throughout.
Unit tests hold the controller to composing accounts from `emits`/`consumes` and nothing else, to
creating the four streams and asserting them idempotently, and to spending a token exactly once;
the host to connecting as the enrolment user with nothing but enrolment permissions; the runtime
to mapping the sdk contract onto subjects exactly as §8 says.
## 11. Open, for the review
**Closed by this revision** (first review, recorded in `MIGRATION-LOG.md`, 2026-09-24): the
`reload-on`/container mismatch (§5), the eaten reply subject on a CONTROL-stream message (§2, §6),
the missing ack permission (§4), and the un-scoped reply inbox under one account (§4). Each is
named where it was wrong, not silently fixed, so a reader comparing against the first version can
find what changed and why.
**Still open:**
- Whether EVENTS should be one stream or one per emitting module (retention per module vs. one
policy). One stream is proposed; the review may disagree.
- The heartbeat interval and the controller's "quiet" threshold on core NATS without persistence —
the same numbers as today are proposed.
- Whether the person's client is a catalogue module (runs on an enrolled workstation node) or a
standalone program (runs anywhere with credentials). Both, in that order, is proposed.
- Leaf nodes: a NATS leaf per machine would make every module's connection local and survive the
hub's restart. Deliberately out of scope; noted so it is not forgotten.
- **New, from this revision:** the `nats` image's own entrypoint now carries logic (watch a file,
signal a process) that no other module's container needed before. Is a one-file-watcher-and-
`SIGHUP` helper common enough across future modules with the same shape (a service that reloads
on `SIGHUP` but runs in a container) to belong in `mesh-sdk` rather than written once per module
that needs it? [ADR 0039](../../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md)'s test —
*does editing it recompile unrelated modules, and does it change often* — probably says no for
one instance; worth asking again if a second module needs the same shape.
@@ -2,7 +2,7 @@
status: resolved
opened: 2026-09-23
located-in: [mesh-host internal/apply]
fixed-by: mesh-host PR #22 — a container records the digest of every file it reads at creation, its env-files and files mounted into it directly, and is recreated when one changes; a mounted directory still needs restart-on
fixed-by: mesh-host PR #22 — a container records the digest of every file it reads at creation, its env-files and files mounted into it directly, and is recreated when one changes; a pre-upgrade label is accepted once, and the plan names the file. A mounted directory still needs restart-on.
amended-design:
---
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-24
located-in: [mesh-catalog modules/minio]
fixed-by:
fixed-by: mesh-catalog — the object-store module repinned to a maintained fork of the withdrawn server image, its runtime sidecar built from source rather than pulled, and its data moved off the predecessor's live directory. The standing condition this report names is not closed by it — see What was done.
amended-design:
---
@@ -109,6 +109,28 @@ a registry the mesh does not control — **by tag or by digest, it makes no diff
dependency with no guarantee behind it, and the mesh currently learns it has lost one only by
trying to use it.
[Issue 064](../064-a-mesh-build-cannot-fetch-a-modules-external-dependencies/00-report.md) is the
nearest precedent, and it does not cover this. That issue asked whether the mesh's build
environment can **reach** a declared vendor image — a network-policy question, answered by
requiring the image be declared as a build input — and it assumed that an image, once declared,
stays fetchable. Withdrawal is the case the assumption does not cover: no network policy and no
declaration makes a deleted repository resolvable, so a module can satisfy 064 in full and still
be unbuildable on a node that holds nothing.
## What was done
The module was repinned to a maintained fork of the server image, published to a registry that
still serves it; its runtime sidecar is now built from source rather than pulled; and its data was
moved off the predecessor's live directory. The object store runs on the control-node from that
pin, and a node holding nothing can obtain it again.
That answers the instance and none of the three points above. The mesh still cannot say which of
its other pinned third-party images are still obtainable, and it would still learn of a withdrawal
only when a node without the image tried to deploy. The replacement question — S3 the protocol
rather than this product — is carried by
[research 015](../../01-RESEARCH/015-the-object-store-after-minio/00-overview.md); the detection
question is carried by nothing, and is the first of the open questions below.
## Open questions
- Should the mesh **hold** the images it depends on — mirroring third-party images into its own
@@ -0,0 +1,86 @@
---
status: open
opened: 2026-09-24
located-in: [mesh-controller module.json, mesh-host internal/apply]
fixed-by:
amended-design:
---
# 114 — Should the controller run as a container, or as a process the host supervises directly?
## What was observed
On the control-node, 2026-09-24, over a long session of operating the mesh through
`mesh-controller`'s CLI (build, push, plan, status, module moved). Every mutating step reached the
binary the same way: `docker exec mesh-controller /mesh-controller <command>` — because
`mesh-controller`'s own manifest declares its one resource as:
```json
{ "id": "server", "type": "container", "name": "mesh-controller", "network": "host", "args": ["serve"] }
```
Two things about that declaration are worth naming together, because neither is a problem on its
own and the combination is what raises the question:
- **`network: host`.** The controller does not use container network isolation, which is the
property a `container` resource type usually buys over a `process` one. It runs with the node's
own network namespace either way.
- **It is the mesh's single point of coordination.** [`03-DESIGN/01-to-be/06-the-controller.md`](../../03-DESIGN/01-to-be/06-the-controller.md)
is explicit: "one node runs it, and nothing takes over" — no election, no quorum, no failover;
recovery is restore, not failover.
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) gives the host — the one thing tier 0 requires
to be a real system daemon — exactly this reasoning for refusing to run in a container: *"installing
the container runtime is a step of the bootstrap, so a host inside a container would need the thing
it exists to install."* The controller is one tier up and does not install the runtime, but it
shares the profile that argument turns on: something the rest of the mesh's operation depends on,
sharing fate with a runtime that is not itself.
## Why it matters beyond this instance
Practically, tonight: every controller interaction was raw shell into a container (`docker exec`),
not a first-class surface — no logs command beyond `docker logs`, no `systemctl status`, and a
session permission classifier that (correctly) treats arbitrary shell into a container as needing
sign-off every time, unlike an ordinary supervised process. That friction is a symptom, not the
issue itself.
The actual question is whether `type: container` is buying the controller anything here besides
image-based delivery and a restart policy — both of which [ADR 0005](../../02-DECISIONS/0005-the-node-host.md)'s
launcher pattern already describes as buildable directly into the host's own supervision (restart on
exit, count consecutive failures, roll back after too many, halt after that), for the host's own
unit. If the controller were declared `type: process` instead — still built and versioned through
the same delivery pipeline, just executed on the node and supervised by the host the way the host
supervises itself — it would stop sharing fate with the container runtime's health (restarts,
upgrades, disk pressure evicting containers) for the one piece of software whose absence the rest of
the mesh is designed to tolerate but nothing is designed to *want*.
This is squarely a question, not a claim that today's shape is wrong: [ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)
already tolerates the controller being down by construction (nodes reconcile from their own
last-applied state), which may make the container-runtime coupling moot in practice. Nobody has
checked.
## Open questions
- Does `mesh-host`'s `process` resource type already support the restart/failure-counting semantics
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) describes for the host's own launcher well
enough for something this central — or would this need host-side work first?
- With `network: host` already in use, what does `type: container` provide the controller today that
`type: process` would not?
- Is there a real circularity risk — the controller's own health depending on the container runtime
it (indirectly, via the host) manages — or does "one node runs it, nothing takes over" already make
a controller outage tolerable regardless of which resource type it is?
- If the answer is "keep it a container," what does that answer, precisely, that this issue asked —
so the next person who notices the same asymmetry finds it answered rather than open again?
## The general case
[Issue 117](../117-a-modules-own-code-is-a-container-and-a-process/00-report.md) is the same
question asked of every module rather than of the controller: a module's own code is a `container`
in [ADR 0047](../../02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)
and a `process` in the to-be design, and no record moves it. Its
[diagnosis](../117-a-modules-own-code-is-a-container-and-a-process/01-diagnosis.md) answers the
first open question above: the host's `process` shape is built, applied and tested, including the
restart and run-to-completion semantics — so this would not need host-side work first.
The two do not collapse into one. The controller is not a code-carrying sidecar, and `network: host`
is what makes the asymmetry visible here and nowhere else.
@@ -0,0 +1,101 @@
---
status: located
opened: 2026-09-25
located-in: [hq, mesh-catalog modules/showcase, mesh-sdk src/tools/index.ts, mesh-tools]
fixed-by:
amended-design:
---
# 117 — A module's own code is a container in one record and a process in another
## What was observed
Asked what the "sidecar" is — the second container a code-carrying module runs beside its
service — and whether a supervised process would do instead. Reading the records to answer it,
the repository answers both ways, and nothing reconciles them.
| record | status | what runs a module's own code |
|---|---|---|
| [ADR 0047](../../02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) | **accepted**, 2026-09-04 | "a **container**, the tool runtime carrying that module's compiled code" — one module, one process, one account; events and tools in that same process, "not a second one to scope and seal" |
| [`01-to-be/18-building-a-module.md`](../../03-DESIGN/01-to-be/18-building-a-module.md) | proposed, 2026-09-21 | a resource type table in which `container` is "an image" and **`process`** is "**its own code**, in three modes", whose default mode is "a unit restarted when it exits", supervised by the machine |
| [`01-to-be/20-writing-a-module.md`](../../03-DESIGN/01-to-be/20-writing-a-module.md) | proposed, 2026-09-21 | one module declaring **four** `process` resources — events, tools, provisioner, a scheduled ingest — each with its own `run` argv, and the sentence "it is why these are `process` rather than four containers" |
Three disagreements, not one:
1. **Container or unit.** ADR 0047 chose a container and said why: a node-wide runtime loading
every module's code could not hold a per-module account, so the runtime is per-module. The
design docs choose a supervised unit running an argv and give no reason, because they do not
record that they are choosing.
2. **One process or several.** ADR 0047's "one module, one process, one account" is the whole
content of its second and third sections. The worked guide declares four for one module and
presents four as the point.
3. **Whether the record was consulted at all.** Neither design doc names ADR 0047 in
`decisions:`. No record supersedes or extends it on this. **The string `process` as a resource
type appears in no decision record** — the shape exists only in two `proposed` design docs.
Meanwhile the thing as built is the container. [ADR 0029](../../02-DECISIONS/0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md)
records that "anything that is a service plus a sidecar currently has to publish a port to talk
to itself," which is one of the things the host's `network` shape was added for.
[Issue 113's diagnosis](../113-the-object-stores-images-were-withdrawn-upstream/01-diagnosis.md)
found a catalogue module declaring "two container resources," the second a runtime sidecar
"pinned at an all-zeros digest, meaning nothing was ever published for it."
[Issue 095](../095-a-module-assigned-after-genesis-has-no-broker-account/00-report.md) is a
sidecar crash-looping on a credential while its service served correctly.
[ADR 0093](../../02-DECISIONS/0093-a-fixture-that-runs-a-modules-runtime-carries-its-name.md)
records that a bed wanting "a sidecar without its server raises the server."
### And the word is in no glossary
"Sidecar" appears sixteen times across five records — two decisions and three issues. It is
absent from [`00-META/glossary.md`](../../00-META/glossary.md), and absent from every document
under [`03-DESIGN/`](../../03-DESIGN/), in both layers. ADR 0047, which creates the thing, never
uses the word; it says "runtime process" and "runtime container". The glossary's own rule is that
"a new name for an existing thing lands here first, in the same change that introduces it in
code," and the page exists because "the terms kept drifting in conversation." A reader asking
what the sidecar is has nowhere in the design layer to look, which is how this was found.
## Why it matters beyond this instance
- **A module author reading the current guide writes a `process`; the catalogue as built declares
a `container`.** [`20-writing-a-module.md`](../../03-DESIGN/01-to-be/20-writing-a-module.md) is
a worked guide with a manifest in it. Whichever of the two is wrong, somebody follows it.
- **The cost of the container shape is paid in four places and totalled in none.** A published
image per code-carrying module, a network so a module can reach itself, a bed that cannot run a
runtime without raising the server it manages, and a credential failure that presents as the
module's own bug. Each record argues its own piece is worth paying. No record puts them beside
the alternative.
- **Both shapes carry a cost the other does not, and neither is written down.** A container
carries its own interpreter; a `process` declaring `run: ["node", "index.js"]` needs an
interpreter present on the machine, which is the machine dependency the statically linked host
([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)) exists to avoid. And `run` is an argv,
where [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) refuses `action` because the link may
not carry a command — a refusal [`18-building-a-module.md`](../../03-DESIGN/01-to-be/18-building-a-module.md)
restates on the same page that it introduces `process`.
- **This is the repository's own named failure mode, in its own records.** `cycle.py` enforces
that a to-be doc names *at least one* decision. Both docs do, so both pass, while introducing a
resource type no decision records and contradicting an accepted one. The rule is "no design
without a decision"; the check is "no design without *a* decision." An unenforced rule is
indistinguishable from a wrong one, and these two documents are what that gap looks like when
something walks through it.
## Open questions
- Which is the decision — container or supervised unit? If the design docs are right, ADR 0047
needs superseding rather than quietly outliving. If ADR 0047 is right, two proposed documents
and a worked manifest describe a resource type that does not exist.
- Is one account per module satisfied by a per-module *unit* as well as a per-module *container*?
ADR 0047's argument rules out a node-wide runtime sharing one account. It does not appear to
rule out a unit holding one scoped credential, and nothing has said so either way.
- If several processes for one module are right, what holds the accounts? ADR 0047 refused "a
second one to scope and seal" for events beside tools. Four processes are four somethings.
- How does a `process` get its interpreter, and does declaring one reintroduce the machine
dependency the host is built to avoid?
- Is `run` an argv the link may carry, given `action` is refused for being one? If the answer is
that a `process` reconciles and an `action` does not, that distinction is not written down.
- What is the thing called, and where does the design layer describe it? Whichever shape wins, no
document in either layer currently says a code-carrying module runs a second thing beside its
service.
- **How would this have been caught?** A decision and a design doc disagreeing on a resource type
is mechanically checkable: the resource types a design doc names are a closed set, and every
member of it either appears in a decision or does not. Whether that check is worth writing is
part of this issue, not settled by it.
@@ -0,0 +1,201 @@
# Diagnosis — 117
## Which trees were searched, 2026-09-25
Named first, because [issue 113](../113-the-object-stores-images-were-withdrawn-upstream/01-diagnosis.md)
is the record of reporting absence in one repository as absence in the mesh.
| Searched | At |
|---|---|
| `mesh-host`, `mesh-catalog`, `mesh-tools`, `mesh-sdk`, `mesh-controller` | `main`, fresh shallow clones |
| `hq` | `main`, and the two branches named under finding 7 |
**Not searched:** the private migration repository; the open pull requests on the catalogue and
the controller; any branch of a code repository other than `main`. A statement below about "the
catalogue" is a statement about its `main`.
## The report's central question is answered: the shape exists
`mesh-host` `internal/declaration/declaration.go` defines `TypeProcess Type = "process"`.
`internal/apply/process.go` applies it — it writes the unit, writes the timer for a scheduled one,
and gates what follows a run-once one. It has tests of its own in both packages. The resource
carries a bundle `source` with a `digest`, a `run` argv, `env` and `env-file`, a `user`,
`restart-on`, and the `run-once` and `schedule` modifiers.
So the report's alternative — "if ADR 0047 is right, two proposed documents and a worked manifest
describe a resource type that does not exist" — is **disproven**. It exists, it is implemented, it
is tested, and the host's vocabulary is now **twelve** shapes rather than the nine
[ADR 0029](../../02-DECISIONS/0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md)
counted.
## The enforcement ADR 0029 asked for is intact, and it recorded this gap rather than closing it
ADR 0029 said "the vocabulary is nine, and the count moves with a record. The test that asserts it
names this one." That test exists — `internal/declaration/declaration_test.go` asserts the count is
twelve and fails with the reason rather than a number. Above the assertion, a paragraph per
addition names what made it one:
| shape | the test names |
|---|---|
| `network`, ninth | ADR 0029 |
| `access`, tenth | ADR 0051 |
| **the eleventh** | **`03-DESIGN/01-to-be/18-building-a-module.md`** — a design document, `status: proposed` |
| `opening`, twelfth | ADR 0100 |
The eleventh is this one. The test still calls it `daemon`, the code calls it `TypeProcess`, and
its paragraph is the only one that names a design document where the others name a decision.
Independently: in `declaration.go`, `TypeProcess` is the **only** shape in the vocabulary whose doc
comment cites no ADR — `network` cites 0029, `access` 0051, `opening` 0100, `user` and the refusal
of `action` cite 0005.
**So ADR 0029's mechanism worked exactly as designed and was not enough.** It requires every
addition to name something. It does not require that something to be a decision, and the one
addition that named a proposed design document instead is the one this issue is about.
### A correction to this trail, recorded because it was one grep from being a finding
The first search here was for `len(Vocabulary())` and found nothing, and the working conclusion for
two steps was that no count assertion existed any more — which would have been written up as "the
mechanism ADR 0029 relied on is gone." It is not gone. The test binds the slice to a local variable
first, so the assertion reads `len(speaks) != 12`. The claim was wrong, it was caught by reading the
file rather than by grepping it, and the shape of the error is the same one issue 113 recorded: a
negative search result read as a fact about the world.
## The argument the report asked for already exists, in a test comment
The report asked why a container rather than a supervised process, and said the reasoning was not
written down. It is — in `declaration_test.go`, as the eleventh shape's paragraph:
> Running code of one's own meant a `container` and therefore an image; running a script meant a
> `service` and a unit somebody else had to install. One intent — run this and keep it running —
> expressed two unrelated ways, with the hosting chosen before anything could be declared. […] It
> is a full-host shape rather than a portable one: it needs a process supervisor to install into.
> It does NOT need a container runtime, which is the point — only software that genuinely needs
> isolation asks for a container.
That is a decision's Context and Consequences, in a Go comment, in another repository. Nothing in
`02-DECISIONS/` contains it. `TypeProcess`'s own doc comment adds the rest — that three modes beat
three kinds, and that a first draft added a `daemon` for the long-running case alone.
## The catalogue is containers, and the one exception is the reference module
71 modules on `main`. Counting the `type` of every declared resource:
| `container` | `process` |
|---|---|
| 115 | **3** |
All three `process` resources are in **one** module: `showcase` — the module
[`20-writing-a-module.md`](../../03-DESIGN/01-to-be/20-writing-a-module.md) is a worked guide for.
### And in that module, the tools do not run
`showcase` declares its migrate, server and reporting steps as `process`. Its fourth resource, the
one for tools, is a **`container`** — and its image is the module's `helper` artifact, which the
same manifest declares as `kind: upstream` from a bare distribution base. Its command is
`sleep infinity`. It mounts the broker credential and sets the variable naming it, and runs nothing.
Meanwhile the module's `code` bundle declares six entrypoints. Three are run by the three `process`
resources. The tools entrypoint and the provisioner entrypoint are **run by no resource in the
manifest.**
Two consequences worth stating separately:
- **The worked guide does not match the module it documents.** The guide shows four `process`
resources, the fourth being `{"id": "tools", "type": "process"}`. The module has three and a
container.
- **This is the condition ADR 0047 was written to end, in a new shape.** That record's Context says
the conversion "produced tools and events that, as it stands, never execute," and its first
Consequence is that they become runnable. In the reference module they do not execute again —
not for want of a runtime this time, but because nothing declares one that runs them.
## The harness has no per-module boundary, and nothing refuses a second module
This is where ADR 0047's isolation argument is load-bearing, so it was checked rather than assumed.
- `mesh-sdk` `src/tools/index.ts`: `serveTools` iterates `collectTools()` over a module-level
registration array and serves **every registered module's** tools over the **one** `broker` it
was handed.
- `mesh-tools` `src/main.ts`: the modules to load come from one variable as a **comma-separated
list**, and the runtime sets its module and node identity from the **single** credential.
- `mesh-sdk` `src/events/index.ts`: an emitted event's `x-source` is stamped from that single
module identity.
Put together: load two modules into one runtime and everything the second emits is attributed to
the first, because there is one credential and the identity comes from it. That is precisely the
failure ADR 0047 predicted — "able to emit as any of them" — reached by a different route, since
the credential is correct and there is only one of it for two modules. **Nothing in either
repository refuses the second module**, and no test asserts that a runtime serves one.
### Ruled out, in fairness to the implementation
- **The serving key conforms.** ADR 0047 replaced a single `tools.invoke` dispatch with a per-tool
key, and the SDK does that: a tool is served on `<module>.<tool>` with the account scoped
`serve.<module>.*`. The superseded `tools.invoke` survives only in **prose** — the doc comment
directly above the conforming code, and the `mesh-tools` README, which also describes the runtime
as per-node. The code is ahead of its own documentation.
- **The credential shape conforms.** The sealed per-module credential file is preferred in code, and
the plain URL is documented as the bootstrap case before a module has an account — not the
ordinary path.
So the account is the right shape and the key is the right shape. It is the **process boundary**
that is declared nowhere and enforced by nothing.
## An unmerged report already asks the narrow version of this
Branch `issue/113-controller-container-or-process`, one commit, 2026-09-24, adds a report titled
**"Should the controller run as a container, or as a process the host supervises directly?"** with
`located-in: [mesh-controller module.json, mesh-host internal/apply]`. Its observation is that the
controller is declared a `container` with `network: host` — so container network isolation, the
property that resource type usually buys, is not in use — and it asks what `type: container` buys
that `type: process` would not.
It was unmerged and numbered 113, which is taken. A sibling branch,
`issue/113-record-the-repin-and-fold-114`, is why `114` was free.
**That report and this one are the instance and the general condition**, and they do not conflict:
it asks about one module that is not a code-carrying sidecar at all, and reaches the same question
from the opposite end. So it lands in this change as
[issue 114](../114-should-the-controller-be-a-container-or-a-process/00-report.md), its commit and
authorship intact, with a section pointing here — rather than being folded in and losing the
`network: host` observation, which is its own and is not reproduced above.
This diagnosis answers its first open question. The host's `process` shape does support what
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) describes for the host's own launcher — the
unit, the timer, restart, and run-to-completion gating are implemented and tested — so that report
does not need host-side work before it can be decided.
## What is located, and what is not
**Located — and it is not a code defect.** The implementation and the design layer agree with each
other; the **decision record is what is missing**, and the accepted record that occupies its place
says the other thing. ADR 0047 is `accepted`, cited by the module protocol, and unsuperseded, while
the host it describes has had a purpose-built shape for a module's own code since the eleventh
vocabulary entry.
| Owner | What is theirs |
|---|---|
| `hq` | the missing record for the `process` shape; ADR 0047 left standing; the worked guide that does not match the module |
| `mesh-catalog modules/showcase` | tools and provisioner entrypoints that no resource runs; a tools container that sleeps |
| `mesh-sdk src/tools/index.ts` | several modules served over one credential, unrefused and untested; a doc comment describing a superseded dispatch |
| `mesh-tools` | a README describing a per-node multi-module runtime the code no longer prefers |
**Not located, and deliberately open:** whether `process` or `container` is *right* for a module's
own code. This diagnosis establishes that the question was answered in practice and never recorded
— not which answer is correct. The arguments on both sides now exist in writing; they exist in a
test comment and a proposed design document, and one of them contradicts an accepted decision.
## What would close it
1. A decision record for the `process` shape, carrying the argument currently in
`declaration_test.go`, and saying what becomes of ADR 0047 — superseded in whole, or in the part
that names a container.
2. `18-building-a-module.md` and `20-writing-a-module.md` naming that record in `decisions:`, and
the worked manifest agreeing with the module.
3. The eleventh shape's paragraph in the vocabulary test naming a decision, like the other three.
4. **How the rule is checked, since a rule states how it is checked:** every shape in the host's
vocabulary names a decision, asserted where the count is already asserted — which turns "no
design without a decision" into something stronger than "no design without *a* decision" for
the one vocabulary where each entry is a security decision.
5. Whether a runtime may serve more than one module answered either way, and asserted — a refusal
if not, a test that two modules' events keep their own source if so.
@@ -0,0 +1,62 @@
---
status: located
opened: 2026-09-26
located-in: [mesh-sdk src/provisioner, mesh-catalog modules/redis]
fixed-by:
amended-design:
---
# 120 — A provisioner remembers what it did, not what is there
## What was observed
The provisioner harness every provider is built on keeps, in memory, a hash of what it last applied
for each consumer: the login, the password and the values. On each pass it skips a consumer whose
hash has not changed. It never asks the backend whether what it made is still there.
The cache module shows what that allows. Its server is configured with a password and a data
directory, and **no ACL file**. So the per-consumer ACL users its provisioner creates exist only in
the server's memory. The server and the provisioner run in separate containers:
1. the provisioner creates an ACL user for each consumer, and records it as applied;
2. the server restarts, for an upgrade or a crash, and comes back with no consumer users;
3. the provisioner, still running, sees nothing changed in what it receives, and does nothing;
4. every consumer of the cache fails to authenticate, and **nothing reports it**. The provisioner's
log is quiet, and the mesh's status is green.
The consumers recover only when the provisioner itself restarts, because its memory is then empty.
Rotating the cache's administrative password happens to cover it, because that file is mounted into
the provisioner too and recreates it. Nothing else does.
Evidence, from the catalogue's and the SDK's main branches: the harness's reconcile loop (`applied`,
keyed by login, compared by hash before `create`), and the cache module's rendered configuration,
which names no ACL file. Found during research 016, how a credential can be rotated, proposed
alongside to-be 27.
## Why it matters beyond this instance
The cache is the case where the backend forgets on its own. The same gap opens whenever a backend
loses what was provisioned while the provisioner keeps running: a store restored from a backup taken
before a consumer was added, a login removed by hand, a server recreated on an empty data directory.
In each of them the provisioner reports that everything is applied, because it compares against its
own memory and not against the backend.
The harness's other half has the same shape. A consumer's contribution that disappears while the
provisioner is down is never removed, because only logins the running process applied are candidates
for removal. What the mesh wants and what the backend holds can drift in both directions, and the
harness sees neither.
This is the design permitting a silent failure. *A provider makes what its consumers require true*
is stated in [to-be 13](../../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md), and nothing
checks it after the first pass.
## Open questions
- Should the harness check each consumer's credential against the backend on every pass, or
periodically, instead of trusting its memory? Most adapters' `create` is already idempotent, so the
cheapest fix may be to drop the hash short-cut and apply every pass. What does that cost for a
provider that recreates an access key on every create, as the object store does?
- Should the cache keep its users in an ACL file, so a restart does not lose them? That fixes this
instance and leaves the gap for the others.
- Where does the record of what was applied live, if not in memory? ADR 0114, still
proposed, puts rotation state with the vault. The same place may answer this.