Re-home this session's new ADRs (0039-0049) and issues (032-037) onto the consolidated scheme; flip issue 003; port repos.md sdk line + feature-branches playbook (07); regenerate index

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
This commit is contained in:
2026-09-05 12:24:07 +02:00
parent 546caa31c4
commit e269f9a185
22 changed files with 1643 additions and 5 deletions
+1
View File
@@ -50,6 +50,7 @@ the expensive half.
| [04](04-build-handoff.md) | Build handoff | A design is ready to be built | | [04](04-build-handoff.md) | Build handoff | A design is ready to be built |
| [05](05-constitution-sync.md) | Constitution sync | `how-we-build.md` changed a rule the mesh enforces | | [05](05-constitution-sync.md) | Constitution sync | `how-we-build.md` changed a rule the mesh enforces |
| [06](06-writing-a-module.md) | Writing a module | Something that runs today must run on the mesh | | [06](06-writing-a-module.md) | Writing a module | Something that runs today must run on the mesh |
| [07](07-feature-branches.md) | Feature branches across repos | Work that changes code, in one repo or several at once |
## Status lives in frontmatter ## Status lives in frontmatter
+57
View File
@@ -0,0 +1,57 @@
# Playbook 07 — Feature branches across repos
**Trigger.** Work that changes code — in one code repo or in several at once (`mesh-sdk`,
`mesh-control`, `mesh-catalog`, `mesh-host`, `mesh-lab`, and `hq` when a decision rides along).
**Who runs it.** Anyone who writes code, engineers and agents alike. Agents follow it exactly —
it is the guard against the failure it was written for.
## The failure it prevents
A feature was worked as a branch-and-MR per *unit of thought* — one per decision, one per
stacked increment — and each MR was treated as finished when it was *opened*, not when it was
*merged*. Across repos the same feature took a different branch name in each. The MRs piled up
unmerged: one session left **sixteen** stacked intermediate MRs that had to be consolidated and
closed by hand. An MR is a review checkpoint, not a scratchpad.
## The rule
One feature is **one branch name**, **one worktree per repo**, **one MR per repo**, opened
**once, at the end**.
1. **Name the feature once.** `feat/<slug>`. The *same* branch name in every repo the feature
touches — never a different name per repo, never a fresh branch per increment within the
feature.
2. **Isolate each repo.** One git worktree per touched repo under `.work/<slug>/<repo>`, branched
off `main`:
```
git worktree add .work/<slug>/<repo> -b feat/<slug> origin/main
```
Parallel features never collide, and no shared checkout is edited.
3. **Commit as you go — locally.** Increments land on the one branch. Nothing is pushed and no
MR is opened mid-feature.
4. **Finish, then publish.** When the whole feature is done — every repo, tests green — push
every branch and open **one MR per touched repo**, together.
5. **Merge promptly, once approved.** Every merge into `main` is notified and approved
([ADR 0023](../../02-DECISIONS/0023-approval-is-the-checkpoint.md)); once it is, merge —
do not leave it sitting. The branch is deleted on merge.
6. **Leave nothing behind.** After the MRs merge, no `feat/<slug>` branch and no `.work/<slug>`
worktree survive.
## What this is not
- **Not a licence to batch unbounded work.** A feature is a *bounded* unit; if it sprawls for
days, end-of-feature bloat merely replaces per-increment bloat. Split it into features, each
its own branch and MR.
- **Not a second trunk.** Every repo branches off `main`. There is no longer an `initialization`
trunk.
## How it is checked
The end state is visible, and its absence is the smell:
- After a feature merges, `git branch -r | grep feat/<slug>` and `git worktree list` return
nothing for it. A surviving branch or worktree means step 6 was skipped.
- More than one open MR in a repo that share no feature name, or a stack of MRs none of which is
merged, is the failure this playbook exists to prevent — stop and consolidate before opening
more.
+1 -1
View File
@@ -31,7 +31,7 @@ target, not the present.
| `mesh-substrate` | 1 | the four pinned services, as declarations | | `mesh-substrate` | 1 | the four pinned services, as declarations |
| `mesh-control` | 2 | **exists.** The control plane and its contexts — one of seven built ([ADR 0006](../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)) | | `mesh-control` | 2 | **exists.** The control plane and its contexts — one of seven built ([ADR 0006](../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)) |
| `mesh-surfaces` | 3 | tools, web, cli | | `mesh-surfaces` | 3 | tools, web, cli |
| `mesh-sdk` | — | contracts shared across tiers | | `mesh-sdk` | — | the stable spine modules build against — the tool-serving harness, the messaging/event framework, the contracts and core primitives. Holds nothing per-module and nothing volatile ([ADR 0039](../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md)). |
| `mesh-lab` | — | **exists.** The lab — scenario lifecycle, networking, placement. Ships to nobody; runs on a workstation. | | `mesh-lab` | — | **exists.** The lab — scenario lifecycle, networking, placement. Ships to nobody; runs on a workstation. |
Tier 4's shape is open, and deliberately so: see ADR 0019 and Tier 4's shape is open, and deliberately so: see ADR 0019 and
@@ -0,0 +1,106 @@
---
topic: building it
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
---
# 39. What the SDK holds, and what it refuses
_Reconciliation note (2026-09-05): supersedes the earlier "repository structure" decision, which the consolidation folded; no standalone record remains to point at, so body references to it now point at the nearest surviving record, [ADR 0015](0015-applications-live-in-their-own-repository.md)._
## Context
The earlier "repository structure" decision (folded in consolidation; see the reconciliation note
above, and [ADR 0015](0015-applications-live-in-their-own-repository.md) as the nearest survivor)
named `mesh-sdk` "contracts shared across tiers:
types, not behaviour." That line is superseded here, because it draws the boundary in the wrong
place. The boundary that matters is not *types versus behaviour* — it is **how often the thing
changes**.
The current SDK is the cautionary tale, and its failure is precise. `hal/sdk` holds all the
code, including a per-module API client for every service (`clients/plex.ts`, `clients/gitea.ts`,
…) and a per-module tool implementation for each (`tools/plex.ts`, …). Every module depends on
the SDK, so **every edit to any of that per-module code rebuilds every module** — the cascade.
The SDK is under constant maintenance precisely because it became the place all the volatile
per-module logic accumulated.
The root cause is worth stating exactly, because the fix follows from it: the pressure was never
to share a client *between* modules. It was to share a client between one module's *own features*
— plex's tools, its health check and its hooks all wanted the same `PlexClient` — and the only
place to share code across a module's features was the global SDK. So **intra-module sharing
leaked out as inter-module coupling.**
## Decision
The SDK holds the **stable spine** that modules build against, and earns its place by rarely
changing. The test for membership is change-frequency, not kind.
### What it holds
- The **tool-serving harness** — the worker and registration mechanism, and the tool-definition
type. *How* a tool is declared and served is settled; it does not change when an individual
tool does.
- The **messaging and event framework** — the broker client, the event consumer, the envelope.
- The **contracts** — the manifest, declaration, provision and link shapes.
- **Core primitives** — sealing and crypto, semver, the shared resolution helpers.
These change rarely and deliberately. When one of them does change, a rebuild of everything is
the *correct* outcome, because the contract every module shares has genuinely changed.
### What it must not hold — the more important half
- **A module's API client.** A Plex client, a Gitea client, a MinIO client belong in their
module. They change when that service's API or the module's use of it changes, which is often,
and which has nothing to do with any other module.
- **A module's tool implementations.** Same reason, same place: in the module.
- **Anything volatile** — anything that changes when one service's features change.
The rule, stated so it can be applied without re-deriving it:
> If editing a thing recompiles unrelated modules **and** it changes often, it does not belong
> in the SDK.
Both conditions are load-bearing. A rare change that cascades is fine — that is a contract, and
the cascade is correct. A frequent change that stays local is fine — that is a module minding its
own business. Only **frequent *and* cascading** is the disease, and per-module clients and tools
are its carriers.
### Where per-module shared code lives instead
Code shared among a module's *own* features lives **in the module**. The default is the plainest
thing that works: an ordinary shared file the features import — `plex/client.ts`, imported by
`plex/tools/`. Within one module, features are files importing sibling files; no package
boundary, no ceremony.
A **module-local SDK** (a sub-package with its own version) is warranted only for the few modules
whose shared surface is large enough to version on its own. It is the exception, not the shape.
Either form gives the property the global SDK could not: editing a module's shared code rebuilds
**that module and nothing else**.
## Consequences
- The cascade becomes **structurally impossible for module logic**. There is no longer an edge
from one module's internals to another, so the only thing that can rebuild everything is a real
change to a shared contract in the SDK — which is rare, and when it happens, is right.
- The SDK is small and stable **by construction**, not by discipline. Its size is no longer a
thing anyone has to police.
- **Converting a module from the current system is partly a de-coupling, not just a move.** Its
client and its tools are pulled *out* of the shared SDK and *into* the module. A conversion
that copied `clients/plex.ts` into the SDK's replacement would rebuild the exact mistake.
- The host still does not import the SDK. It depends on nothing
([ADR 0005](0005-the-node-host.md)) and **mirrors** the contracts rather than
importing them, exactly as its apply-shapes table already does deliberately. The SDK is shared
by the tiers that *can* share code; the host is not one of them.
## References
- The earlier "repository structure" decision — named the repositories; its `mesh-sdk`
description ("types, not behaviour") is superseded by this record (folded in consolidation;
nearest survivor [ADR 0015](0015-applications-live-in-their-own-repository.md)).
- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — the decomposition this serves: code
belongs to the boundary that owns it.
- [ADR 0005](0005-the-node-host.md) — why the host mirrors the contracts instead of
importing the SDK.
+101
View File
@@ -0,0 +1,101 @@
---
topic: what runs on it
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
extends: 0009-modules-and-the-graph.md
---
# 40. What a module is
_Reconciliation note (2026-09-05): supersedes the earlier "grouped by domain" decision, which the consolidation folded into how-we-build.md; no standalone record remains to point at._
## Context
[ADR 0009](0009-modules-and-the-graph.md) settled that everything is a module, but never said what a
module *is* beyond "a directory the mesh processes." That gap let the catalogue's breadth read as a
smell: a module can carry a container, a built image, tools, a provisioner, migrations, health,
config, seat claims, requires and provides — so much that the unit seemed ill-defined.
The earlier "grouped by domain" decision (folded in consolidation; see the reconciliation note
above) tried to organise modules by domain, which is the wrong axis. This record states what a module is, drawn from the cases that
stress-tested it: the shell, i3-vs-sway, umami, and "database."
## Decision
**A module is one self-contained piece of software the mesh installs and manages** — everything
needed to make that one thing real and integrable: what runs, the seats it claims, what it provides
to other modules, what it requires from them, and what operates it.
The **software is the module's identity.** Capabilities, seats and provisioned resources are the
**relationships *between* modules**, not what a module is — and that is what binds a module into one
thing. umami is bound by *being umami*: its container runs umami, its provisioner creates umami sites,
its tools query umami, its `requires` gets umami a database. Every feature serves the one software.
### The three relationships
1. **Shared seat** — several modules fulfil a capability and coexist; one may be default. bash, zsh
and fish all join `shell`.
2. **Exclusive seat** — modules contend for a single slot; one holds it. i3 (needs x11) and sway
(needs wayland) contend for `display-session`.
3. **Provide / require** — a provider ships the **provisioner** that creates instances of the
resource it offers and returns sealed credentials; a consumer requires it and the mesh wires the
credential in. Symmetric: umami requires a database *and* provides analytics.
### Interfaces are mesh-owned; providers adapt to them
The mesh **defines the interface** for a capability — the provider-neutral contract of what a
consumer receives and how it integrates. Both sides conform: a provider's provisioner **adapts** its
software's real API to the mesh contract; a consumer depends on the **interface**, never on a
provider. Swap one provider for another and the consumer does not change.
### The naming rule — draw the interface at the consumer's real coupling
Name a `provides`/`requires` at the **widest boundary across which the consumer genuinely does not
care which implementation serves it**:
- Where the consumer's coupling is thin — an analytics embed snippet and dashboard, opaque to it —
the mesh defines a neutral interface (`analytics`) and providers (umami, amumi) adapt. Swappable
across vendors.
- Where the consumer **speaks a protocol** — a database's wire protocol and query dialect — the
interface *is* the protocol: `postgres-database`, `mssql-database`, `mongodb-database`. Swappable
only among protocol-compatible implementations, **never across**, because the application cannot
cross it either. "database" is not a capability; the protocol is.
- **Never false genericity.** A name must not promise a swap the contract cannot deliver
([research 005](../01-RESEARCH/005-domain-grouping/analysis.md)).
This is [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md)'s rule made general — "names the protocol, not
the product; a database names the engine because the app targets it" — with the reason stated: the
contract sits where the coupling is.
### What is not a module
- A **library** (built against, never deployed — [ADR 0039](0039-what-the-sdk-holds-and-refuses.md)).
- A **control-plane context** (the mesh itself — [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)).
A **swappable machine mechanism** (a firewall — ufw, nftables) *is* a module implementing a
capability. The host hardcodes no firewall, supervisor, package manager or runtime; it owns only the
generic apply primitives and platform detection, so it runs where none of those exist — an Android
phone has no ufw, systemd, pacman or Docker.
## Consequences
- **Supersedes the earlier "grouped by domain" decision** (folded in consolidation; see the
reconciliation note above). Modules are
organised by their relationships (seats, provisions), not grouped into domain folders.
- **Refines [ADR 0009](0009-modules-and-the-graph.md).** Everything the mesh runs and integrates is
a module — but a module is defined by the *software it delivers*, not by being a bucket of features.
- The target is a **self-fulfilling mesh**: declared wants bound to swappable modules, provisioners
wiring credentials, nothing hardcoded. The control plane's whole job is the binding.
- Converting a module from the old system includes pulling its per-module code out of the shared SDK
([ADR 0039](0039-what-the-sdk-holds-and-refuses.md)) and shipping its provisioner as an adapter to a
mesh interface — a de-coupling, not just a move.
## References
- [ADR 0009](0009-modules-and-the-graph.md) — everything is a module; this says what one is.
- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — contexts are the mesh, not modules.
- The earlier "grouped by domain" decision — superseded (folded in consolidation; see the note above).
- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — protocol-not-product, generalised here.
- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — per-module code lives in the module.
- [research 011](../01-RESEARCH/011-the-module-graph/00-overview.md) — the graph of these relationships.
@@ -0,0 +1,85 @@
---
topic: what runs on it
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
extends: 0040-what-a-module-is.md
---
# 41. Events are a relationship, the lighter sibling of provisioning
## Context
[ADR 0040](0040-what-a-module-is.md) names two relationships between modules — seats and
provide/require (provisioning). A third is latent in the mesh and worth making first-class: the
broker every node already runs ([ADR 0002](0002-nodes-communicate-over-a-broker.md)) can carry a
module's activity as **events**, which any other module reacts to. A logger that writes an audit
trail, a module that acts when another module acts, observability — all of it is one mechanism, and
today it is ambient rather than declared.
## Decision
**A module emits events and consumes events, and both are declared** — parallel to `provides` /
`requires`, so the mesh knows the event graph the same way it knows the provisioning graph.
### Events are provisioning's lighter sibling
| | provisioning | events |
|---|---|---|
| shape | **1:1**, a provider creates a resource *for* one consumer | **1:many**, a module emits, any number listen |
| credential | yes — sealed, per consumer | none — it is broadcast |
| machinery | a provisioner (the reconcile adapter) | nothing but the broker's topic routing |
| declared as | `provides` / `requires` | `emits` / `consumes` |
Because an event is broadcast and credential-free, there is no provisioner and no per-consumer
setup — only a subscription. That is why it is the *lighter* relationship, and why most
inter-module reaction should be an event, not a provision.
### An event carries what an audit needs
Every event carries its **type** (a dotted topic key, so listeners match by prefix), its **source**
module, the **node** it came from, and the **time**. A body follows. The metadata is not optional:
a reaction may only need the body, but an audit trail needs to know who did what, where and when,
and an event that cannot answer that is not auditable.
### The audit logger is just a consumer of everything
A logger that records the whole mesh's activity is **not a privileged component** — it is an
ordinary module that consumes `#` (every event) and writes them down. It holds no special access;
it only listens widely. That it falls out of the model with no new machinery is the check that the
model is right.
### `consumes` is validated like `requires`
A `consumes` for an event that **nothing** `emits` is a dangling edge, and the mesh refuses it
before deploy — the same rule that catches a `requires` for a resource nothing provides
([research 011](../01-RESEARCH/011-the-module-graph/00-overview.md)). A listener waiting for an
event that can never arrive is a silent failure, and this repository's whole discipline is against
silent failure.
### One runtime serves all three
The per-node module runtime that serves a module's tools also wires its `consumes` (subscribe,
dispatch to the handler) and lets its code `emit`. Tools are *invoked* (request/reply), resources
are *provisioned* (1:1, credentialed), events are *emitted and consumed* (1:many, broadcast) —
three relationships, one broker, one runtime, all declared on the manifest.
## Consequences
- The mesh gains a declared **event graph** alongside the provisioning graph — visible, validated,
reasoned over.
- **Reaction becomes the default coordination**: a module acts on another's event without either
knowing the other, and without a credentialed link. Coupling drops.
- An **audit trail** is a module, not a platform feature — and can be swapped, extended or run more
than once (a file logger and a queryable one) with no change to anything that emits.
- The runtime must dispatch a module's event handlers as well as its tools; that generalisation is
small (both arrive by importing the module's entrypoint) but it is real work.
## References
- [ADR 0002](0002-nodes-communicate-over-a-broker.md) — the broker events ride.
- [ADR 0040](0040-what-a-module-is.md) — the relationships this extends.
- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — `emit`/`on` are stable sdk surface; the
broker binding and the runtime are not.
- [research 011](../01-RESEARCH/011-the-module-graph/00-overview.md) — the graph these edges join.
@@ -0,0 +1,116 @@
---
topic: what runs on it
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
extends: 0041-events-are-a-relationship.md
---
# 42. The shape of an event on the wire
## Context
[ADR 0041](0041-events-are-a-relationship.md) made events a relationship — `emits`/`consumes`, the
graph, the audit logger. It did not say what an event *is* on the broker: the exchanges, the
routing keys, the headers, the queues and their configuration. That shape is a contract every
emitter and consumer conforms to, exactly as [ADR 0010](0010-delivery.md)
is for declarations — and it was being decided ad-hoc in code. This settles it, so the sdk and the
runtime implement one contract and a module never reinvents it.
## Decision
### Two exchanges, kept apart
- **`mesh.events`** — a durable topic exchange. Every event rides it: module, mesh and node.
- **`mesh.rpc`** — a durable topic exchange. Tool invocations (request/reply) ride it.
Kept separate because RPC is not an event: a `#` subscription on `mesh.events` is then a complete
audit of what happened, with none of the invocation traffic.
### The routing key is the event type, namespaced by origin
Dotted and hierarchical — `<origin>.<name>.<event…>` — with three reserved origins:
- `module.<module>.<event>` — `module.umami.site.created`
- `mesh.<context>.<event>` — `mesh.delivery.deployed`, `mesh.provisioning.granted`
- `node.<node>.<event>` — `node.anchor.joined`, `node.anchor.unreachable`
Topic matching gives a consumer `node.*.joined`, `module.umami.#`, or `#`. The origin roots are
reserved; everything after is the emitter's own namespace.
### Metadata in headers, payload in the body
An event's identity and provenance are AMQP **headers**, so a consumer — or the broker, or an
audit tool — reads who/when/what without parsing the body, and the body is only the domain payload.
**Required headers**
| header | meaning |
|---|---|
| `x-event-id` | a unique id — for dedup and audit (delivery is at-least-once, below) |
| `x-source` | the emitter: the module, context or node name |
| `x-node` | the node it was emitted from |
| `x-time` | emit time, RFC-3339 |
| `content-type` | `application/json` |
**Optional headers**
| header | meaning |
|---|---|
| `x-causation-id` | the event or command that caused this one — tracing |
| `x-schema` | a version of the body's shape, so a body evolves without silent misreads |
The routing key already carries the type; it is not duplicated as a header. An **unknown `x-`
header is ignored, not refused** — unlike a declaration, an event is observed by parties that need
not all understand every header, and refusing would couple every consumer to every emitter's
additions.
### Messages are persistent
Events are published persistent (delivery-mode 2). An audit trail that loses events on a broker
restart is not one, and the cost is disk the broker already spends on everything durable.
### Queues: one per consumer, durable, dead-lettered
- **A consumer's queue** is `<node>.<module>.events`, durable, bound to that module's consumed
patterns. Durable so a restart does not drop what arrived while it was down. **Manual ack** after
the handler succeeds — at-least-once.
- **Prefetch** bounds in-flight work (default 32) so one slow consumer does not pull the whole
backlog into memory.
- **A dead-letter exchange** `mesh.events.dead` receives a message rejected past a redelivery limit,
so a poison event is set aside for inspection rather than looping forever or vanishing silently.
- **The audit logger's queue** `<node>.audit-logger.events`, bound to `#`, is the same shape —
durable, persistent, dead-lettered — because completeness is its whole job.
- **RPC reply queues** are exclusive, auto-delete and server-named; **RPC serve queues**
`serve.<key>` are durable and shared, so several runtimes serving one tool key compete rather than
each answer.
### At-least-once, and consumers are idempotent
A handler may see an event twice — a redelivery after a crash between doing the work and acking.
Consumers must be idempotent, and `x-event-id` is what makes dedup possible. **Exactly-once is not
offered**: it is a promise no broker keeps honestly, and saying so is better than pretending.
## Consequences
- The event shape is a versioned, enforced contract, not conventions each module reinvents. The
sdk's `emit`/`on` and the runtime's AMQP binding implement it; a module never sees an exchange or
queue name.
- Metadata-in-headers means the body is exactly the domain payload, and a consumer that only wants
provenance never parses it.
- Adding a header or an origin root widens the contract and is reviewed as one — the discipline
[ADR 0010](0010-delivery.md) applies to the
declaration vocabulary.
- The sdk's first cut carried source/node/time in the *body*; this supersedes that — they move to
headers. That is code to align, in `mesh-sdk` (`emit`/`on`) and `mesh-tools` (the binding, queue
config, dead-letter).
## References
- [ADR 0041](0041-events-are-a-relationship.md) — events as a relationship; this is their wire shape.
- [ADR 0010](0010-delivery.md) — the precedent: a wire
contract, versioned, additions reviewed as security.
- [ADR 0002](0002-nodes-communicate-over-a-broker.md) — the broker.
- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — `emit`/`on` are stable sdk surface; the
binding, queue config and dead-letter are the runtime's, not the sdk's.
@@ -0,0 +1,101 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0041-events-are-a-relationship.md
---
# 43. A module's broker account is scoped by what it emits and consumes
## Context
[ADR 0041](0041-events-are-a-relationship.md) made events a relationship — `emits` and `consumes`
on the manifest. [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) gave them a wire shape — the
`mesh.events` exchange, the durable per-consumer queue, the reserved routing-key origins. Neither
said how a module *reaches* the broker: what account it holds, and what that account is allowed to
do.
As the code stands, there is no answer. The mesh can provision a **node** account (at enrolment)
and a **builder** account (scoped to the build queue), and it can *deliver* any module a sealed
own-secret at a declared path — but it has no way to provision a broker **account** for a general
module. A module that declares `own-secrets: {broker: …}` and nothing more receives thirty-two
random bytes, not a credential. So on the broker, `emits` and `consumes` are enforced by nothing: a
running module could bind any queue, consume any pattern, and publish under any origin, and the
manifest that says otherwise would be describing a boundary no code draws — the exact shape of fault
[04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) records, a scope
declared in manifests and read by nothing.
This settles it, so a module's place on the bus is a thing the broker enforces rather than a thing
the manifest merely claims.
## Decision
### A module gets a broker account when it is assigned, and its permissions are the manifest
When the mesh assigns a module to a node it provisions a broker account for that module on that node,
sealed to the node ([ADR 0004](0004-a-node-and-how-it-joins.md)) and delivered as the
module's `own-secrets` broker — `amqps://` with the mesh's fingerprint, the shape
[ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) already carries. The account's permissions are
derived from the manifest, and are exactly these:
- **What it consumes.** Read on `mesh.events`, and configure-and-read on its own queue
`<node>.<module>.events` bound to the patterns in `consumes`. It cannot bind or read another
module's queue. A module that consumes nothing gets no read on the events exchange at all.
- **What it emits.** Write to `mesh.events`, restricted to routing keys under its own origin,
`module.<name>.*`. It cannot publish as another module, and cannot publish under the reserved
`mesh.*` or `node.*` origins — those belong to the mesh and the host (ADR 0042). A module that
emits nothing gets no write.
- **Nothing else.** The events account reaches `mesh.events` and that module's own queue, and no
more. Tool serving and calling over `mesh.rpc` is a separate grant on the same principle — a
module serves the tool keys it declares and calls the ones it is bound to — and is scoped the same
way rather than folded in here.
### Consuming everything is a privilege, granted deliberately
`consumes: ["#"]` — the audit logger — is read across the whole bus: every module's events, the
mesh's, every node's. That is not a pattern like any other; it is the power to see everything, and
the account is where it becomes visible. The grant that lets one module read the entire bus is one
the mesh issues on purpose and can be audited — the answer to *who can read everything* is a row, not
a guess — rather than a breadth any manifest acquires by typing a single character. A `#` consume is
a reviewed grant, not a default one.
### The account is how the declaration is enforced
Because the account can do only what `emits` and `consumes` name, the broker itself refuses a module
that tries to consume a queue it did not declare or emit under an origin it does not own. That is what
makes an event relationship a rule and not a comment — the discipline that a stated rule says how it
is checked. A manifest that over-declares grants more than the module uses, which is visible and
reviewable; one that under-declares makes the module fail closed at the broker, which is the safe
direction to be wrong in.
## Consequences
- The control plane gains a **generic module broker-account**, derived from the manifest. The
builder stops being a special case: its access to the build queue becomes an ordinary expression of
what it consumes and serves, not a bespoke account method. One rule, and the builder is an instance
of it.
- The runtime reads its credential from a file (the broker own-secret), `amqps://` verified against
the mesh's fingerprint. The `guest` account is for raising the substrate, never for a module — a
module documented as holding its own credential and handed the broker's administrative one is worse
than one with no credential story at all.
- `emits` and `consumes` stop being advisory. They are the module's authority on the bus, so the
manifest is now a security boundary and is reviewed as one, the discipline
[ADR 0010](0010-delivery.md) applies to the declaration
vocabulary.
- *Who can read the whole bus* becomes an answerable question, because `#` is a grant and not an
accident.
## References
- [ADR 0041](0041-events-are-a-relationship.md) — events are a relationship; this scopes the account
by that relationship.
- [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) — the wire this account secures: the queue,
the origins, the `amqps` credential shape.
- [ADR 0004](0004-a-node-and-how-it-joins.md) — the link is the security boundary; a
module's account is sealed to its node the same way a node's is.
- [ADR 0010](0010-delivery.md) — a declaration is owned
and its additions reviewed; a module's broker permissions are that discipline applied to the bus.
- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — a scope declared
in manifests and enforced by no code: the fault this decision closes for events.
@@ -0,0 +1,94 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0027-a-provision-names-what-the-consumer-is-coupled-to.md
---
# 44. A public name is provisioned, not registered by hand
## Context
The mesh names and resolves its own machines internally: the overlay generates
`<service>.<node>.<suffix>` wildcards, dnsmasq answers them (`wildcard-resolution`), and the mesh
issues a certificate for each internal name. A service reachable at a *public* domain —
`plex.example.com`, not `plex.anchor.internal` — needs three things that machinery does not give it:
- a **public DNS record** at a registrar or DNS provider, so the name resolves on the internet;
- a **publicly-trusted certificate** for it, because the mesh's own authority is trusted by nobody
outside the mesh;
- and routing from that name to the module — which the reverse proxy already does: a module
`requires` the `route` capability and the proxy provides it, routing by the host it was asked for.
The routing exists. The public DNS record does not: the mesh has no way to make a name resolve on
the public internet, so today that is a step someone does by hand at a DNS provider, outside the
mesh, remembered nowhere. A public name is therefore the one part of reaching a service that the
declaration graph cannot grant or withdraw — which means it is created once and outlives whatever it
was for, the shape of drift this project exists to remove.
## Decision
### A public name is a capability, requested like any other
A module reachable at a public host declares `requires: ["public-dns"]` and contributes the hostname
it wants — beside `requires: ["route"]`, which exposes it through the proxy. The name is then
provisioned on declaration ([ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md)): created
when the module is assigned, removed when it is withdrawn, reconciled like every provision.
### The interface is neutral; the providers are the registrars
`public-dns` is drawn at the consumer's coupling: the consumer wants *a public name that resolves to
me*, and does not care whether Cloudflare, Route 53 or a registrar's own API puts the record there.
So the interface is neutral and the providers are provider-scoped — `cloudflare-dns`,
`route53-dns`, `porkbun-dns` — each implementing the one `public-dns` contract, the same way a
neutral database coupling is answered by `postgres-database` and `mssql-database`. A module names
`public-dns`; it never names a registrar.
### The record points at the mesh's public ingress, not at the node
What the name resolves to is the address the reverse proxy answers on, not the consuming machine's.
A public service is reachable only *through* the proxy — the proxy holds the `route` grant and routes
by host to the module — so the public name must resolve to the proxy. `public-dns` and `route` are
the two halves of one public exposure: the name, and what the name reaches.
### The record is a fact, not a secret
A DNS record is public by definition, so the grant returns the fully-qualified name and its TTL and
nothing sealed. The only secret is the provider's own API credential, which is the provider module's
own-secret and never leaves it — the module that wanted the name never sees it.
### Events
The provider emits `module.<provider>.record.created` and `module.<provider>.record.removed`
([ADR 0041](0041-events-are-a-relationship.md)), so *which names the mesh publishes, and where* is a
question answered from the event trail and the grants, not from a folder of records edited at a
provider.
### The public certificate is the proxy's, and is named here only to pair it
A public name without a publicly-trusted certificate is reachable and not trusted — the same pairing
the internal name and the mesh-issued certificate already have. Obtaining that certificate (ACME
against the now-resolving public name) is the reverse proxy's to do, and its mechanism is its own
decision; it is named here so the pairing is not forgotten, not resolved here.
## Consequences
- A public name is created and torn down with the module, so it cannot outlive it, and the mesh can
say which public names it publishes without anyone reading a registrar's dashboard.
- Adding a registrar is adding a provider that answers `public-dns`; the modules that want names do
not change.
- Public exposure of a service is a trio of separate, declared, enforced relationships: the firewall
opens the proxy's public port ([ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)),
`route` routes the host to the module, and `public-dns` makes the host resolve.
## References
- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — a capability is provisioned on
declaration; a public name is one.
- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the firewall, the other
half of the reachability question this was asked with.
- [ADR 0041](0041-events-are-a-relationship.md) — the provider's record events.
- [ADR 0040](0040-what-a-module-is.md) — a provider and its interface; the neutral-interface,
scoped-provider naming this follows.
@@ -0,0 +1,93 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0005-the-node-host.md
---
# 45. A machine's firewall is the sum of what its modules listen on
## Context
The reverse proxy is a *provider*: a module `requires` the `route` capability and a running proxy
provides it, routing traffic by name and reaching back to the consumer. A fair question follows —
is the firewall the same shape? Should a module *register* a port with a firewall provider the way
it requests a route?
It should not, and the difference is the point. A reverse proxy is a service another component
performs; a firewall is a property of the machine — a packet filter the host applies to itself.
Modelling it as a provider would invent a credential and a reach-back for something that has neither.
And the mesh already has the registration: a module declares `listens: [{ port, from }]` — the port
it accepts connections on, and from where. That *is* how a service says it wants a port open. What is
missing is not a model but enforcement. [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md)
records that a `scope:` key five manifests carry is read by no code: a manifest can appear to
restrict a port and restrict nothing — the exact fault
[how-we-build.md](../00-META/how-we-build.md) names, *an unenforced rule is indistinguishable from a
wrong one*, made worse because the declaration reads as a restriction.
## Decision
### The firewall is derived and host-applied, not a provider
A machine's firewall is the sum of what the modules assigned to it declare they listen on, computed
by the host and applied as one of its owned resources ([ADR 0005](0005-the-node-host.md):
the host applies, it does not decide; [ADR 0010](0010-delivery.md):
the declaration is owned resources). It is not a capability, not a per-consumer grant — opening a
port is a declarative fact about a machine, so it is computed and applied, not requested and
credentialed.
### `from` is the whole of public-versus-internal
The distinction the question is really about lives in `from`:
- `listens: [{ port: 5432, from: mesh }]` — open to the private overlay only.
- `listens: [{ port: 443, from: anywhere }]` — open to the public internet.
A module registers a port on the firewall by listening on it and saying from where. There is no
separate firewall capability, because the firewall is not a thing that reaches back or holds a
secret; it is the machine's own filter over the ports its modules named.
### The host enforces it both ways, and unknown keys are refused
A port a module listens on is opened to exactly the scope it named; a port nothing declares is
closed. And a key the firewall does not read — the `scope:` of issue 003 — is refused at the
manifest, not accepted and ignored, so a declaration that reads as a restriction is one. This is the
discipline [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) applied to the
broker account, applied here to the packet filter: the declaration is the enforcement, or it is a
comment.
### A public service is exposed through the proxy, not by opening its own port
Reaching the public internet is normally not `from: anywhere` on the service's own port. The service
listens `from: mesh` — only the proxy reaches it — and `requires: route`, so the sole machine with a
public opening is the one running the reverse proxy, and the service is exposed by name through it.
`from: anywhere` is the deliberate direct-exposure case, for a service that is its own front door.
## Consequences
- Issue 003 is closed: the firewall is computed from `listens` and enforced, so a declared scope is
real and an undeclared port is shut. Rejecting unknown manifest keys is the general fix, of which
the `scope:` key was one instance.
- The firewall and the reverse proxy stop being confused for one model: the firewall is the machine's
filter (host-derived from `listens.from`); `route` is a name-router (a provider); the public DNS
name is a third thing ([ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md)). A
public service uses all three.
- The modelling question is answered: a module registers a port by declaring `listens`, and reaches
the public internet by name through `route` + `public-dns` — never by the firewall being a
provider.
## References
- [ADR 0005](0005-the-node-host.md) — the host applies; the firewall is one of
the things it applies.
- [ADR 0010](0010-delivery.md) — the firewall is a derived
owned resource, not a grant.
- [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) — the same discipline:
a declaration is enforced, or it is a comment.
- [ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md) — the public name, the other
half of the reachability question this was asked with.
- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — the unenforced
`scope:` this closes.
@@ -0,0 +1,88 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0027-a-provision-names-what-the-consumer-is-coupled-to.md
---
# 46. A module's configuration is its assignment's, not its manifest's
## Context
A module is assigned to a node — `assign <node> <module>`, always to a machine; there is no
assignment to the mesh. "Mesh" is a *scope*, not a place: a `provides` or a `claim` scoped `mesh`
reaches the whole mesh, but the module still runs on a node. So the two kinds of thing a module can
carry are the manifest (what the module *is*) and, separately, what it should do *here* — which
differs by deployment and by node.
The mesh already has the second: **settings**. `settings set <module> [--node <node>]` — with a node
it is that machine's, without it the whole mesh's — layered over what the module declares and applied
at resolution, changeable without editing the module and without a rebuild. That is the surface a
meshboard would edit.
But settings today reach only a module's **config-file content** (a mergeable file the module owns).
Configuration that is not a file has been landing in the manifest instead, statically — a registrar's
zone and domain, the address public names point at, and, most sharply, `listens.from`. That last one
is the tell: whether a port is open to the private overlay or to the public internet is a
*per-node deployment choice* — the same database internal on one machine and public on another — and
a value fixed in the manifest is one value for every machine, so it cannot be. Static configuration in
the manifest is configuration in the wrong place: it cannot vary per node, and it cannot change
without a new module version.
## Decision
### The manifest is identity and defaults; the assignment's settings are the configuration
A module's manifest declares what it is — what it provides, requires and claims, the shape of its
resources — and, for anything configurable, a **default**. The values that make a running instance
*this* instance are settings, carried by the assignment: per-node, or mesh-wide when no node is named,
applied over the defaults at resolution. Change one and the next reconcile carries it; nothing is
edited on a machine and nothing is rebuilt.
### Settings drive the configurable fields the manifest marks, not only file content
Settings extend beyond a config file's content to the manifest fields a module declares settable —
foremost:
- **`listens.from`**: a module declares its safe default (`from: mesh`), and a per-node setting
raises or lowers it. postgres declares `listens: [{ port: 5432, from: mesh }]`; on the machine that
should expose it, a setting makes that port `from: anywhere`. Same module, different exposure, and
the firewall ([ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)) is computed
from the effective value, so the packet filter follows the setting.
- **A provider's own configuration**: a registrar's zone, domain and the ingress its names point at
([ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md)) are mesh-wide settings, not
manifest constants — one mesh's Cloudflare zone is not another's, and the module description is the
same for both.
### Unset is the default, and an unknown setting is refused
A field with no setting keeps the manifest's default, so a module runs correctly configured by nobody.
A setting that matches no settable field — like a config value that reaches no file today — is named,
not silently dropped, so a misspelled setting is found rather than believed (the discipline of
`UnusedSettings`, and of [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md):
a declaration is enforced or it is a comment).
## Consequences
- The postgres case works: one module, `from: mesh` by default, `from: anywhere` where a setting says
so — internal on ace, public on novox, changeable live.
- Provider modules stop carrying a mesh's specifics: `cloudflare-dns` describes *a Cloudflare
registrar*, and *which* zone and ingress is a setting, so the same module serves every mesh.
- Configuration becomes a thing a meshboard manages — set per node or mesh-wide, applied on the next
reconcile — rather than a manifest edit and a rebuild ([ADR 0011](0011-managed-files-are-generated-never-edited.md):
the way you change a managed thing is not by editing it).
- What a manifest may not do is grow a value that differs per machine; if it differs per machine it is
a setting, and the manifest holds only the default.
## References
- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — what is provisioned on
declaration; its per-instance values are the assignment's.
- [ADR 0011](0011-managed-files-are-generated-never-edited.md) — a managed thing is changed through
the mesh, not by editing it; settings are that, for configuration.
- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the firewall follows the
effective `listens.from`, so making `from` a setting makes exposure a setting.
- [ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md) — the provider whose zone and
ingress are settings, not manifest constants.
@@ -0,0 +1,87 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md
---
# 47. A module runs its code as its own process, with its own account
## Context
A module is one self-contained thing ([ADR 0040](0040-what-a-module-is.md)), and it gets a broker
account scoped to what it emits and consumes ([ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md)).
The catalogue now gives modules **tools** and **events** — real code, in the module ([ADR 0039](0039-what-the-sdk-holds-and-refuses.md)) —
but nothing has said what *runs* that code. The audit-logger showed one shape and was treated as an
exception: a container running the tool runtime carrying the module's compiled code, holding the
module's own scoped account. Every module with tools or events needs the same, and the tempting
alternative does not work.
**A node-wide runtime that loaded every assigned module's code cannot hold a per-module account.** It
would run under one account with the union of every module's permissions — able to emit as any of
them and read any of their queues — which is exactly the isolation ADR 0043 exists to draw. So the
runtime is per-module, not per-node, and treating the audit-logger as special left the other
modules' code with nothing to run it: the conversion produced tools and events that, as it stands,
never execute.
## Decision
### A module with tools or events runs a process of its own
A module that has tools or events runs a **runtime process** — a container, the tool runtime carrying
that module's compiled code — assigned and started like the module it is, holding the single broker
account the mesh scoped to it (ADR 0043). One module, one process, one account.
### It serves its tools, each on its own key
A tool is served on its own key (`serve.<tool>`), and a caller invokes a named tool. Only the module
that serves it answers, and the module's account is scoped to exactly its tool keys — so one module
cannot answer another's calls, the isolation ADR 0043 gives events extended to tools. This supersedes
a single `tools.invoke` endpoint that dispatched by name: that shape assumed one runtime for the
whole node, and per-module runtimes competing on one key would each be handed calls for tools they do
not have.
### It runs its events in the same process, under the same account
Emitting under the module's own origin and consuming its own queue ([ADR 0042](0042-the-shape-of-an-event-on-the-wire.md))
happen in that same process, with that same account — not a second one to scope and seal. A module's
tool code, its event code and, for a provider, its provisioner are the one module's code and run as
the one module's process.
### The runtime image is the tool runtime plus the module's code
Built from the module's source like any module image — the audit-logger's shape, made the rule, not
the exception. The module declares a `container` for it carrying `MESH_BROKER_FILE` (its sealed
credential, ADR 0043) and its compiled code. A module with **neither** tools nor events runs no such
process: a plain service module — the plex *server*, dnsmasq the resolver — is its service and files
and nothing more. A module that is both a service and code declares both containers: the service, and
the runtime beside it.
## Consequences
- The catalogue's tools and events become runnable: each tools-or-events module gains a runtime
container with its scoped credential, and the audit-logger stops being special. Until this, the
converted modules held code with nothing to execute it.
- A process, and a small image, per tools-or-events module. That is the cost of ADR 0043's isolation:
one account per module means one process per module. It is paid deliberately — a shared runtime is
cheaper and cannot be scoped, and a mesh where any module can emit as any other is not one worth the
saving.
- `serve.<tool>` per key replaces the single `tools.invoke` dispatch. The sdk's serving and a module's
account scope both come to name tools individually.
- **A provider's provisioner is a runtime process too.** It already runs as its own container; its
events (`bucket.created`, `database.provisioned`) belong to *that* process and need the same
credential. So a provisioner that emits carries `MESH_BROKER_FILE` and its scoped account like any
runtime — or it does not emit. (This is the fix for provisioners that emit today with no broker
bound: the emit is a runtime's, and the provisioner is a runtime.)
## References
- [ADR 0040](0040-what-a-module-is.md) — a module is one self-contained thing; its code runs as one
process.
- [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) — the scoped account
this process holds, and the isolation that makes it per-module.
- [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) — the events this process runs, and the
`serve.<key>` queue tools now use.
- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — the code lives in the module; this runs it.
@@ -0,0 +1,131 @@
---
topic: what runs on it
status: accepted
date: 2026-09-05
deciders: jochen
reconstructed: false
---
# 48. A provider creates the credential the mesh minted, and seals nothing
## Context
A provider module stands up a per-consumer resource — a database, a cache bucket, an object
store user — and the consumer must end up holding a credential that authenticates against it.
Building the module runtime (ADR 0047: a module runs its own code as its own process under its
own account), the provider's provisioner was run for the first time as a delivered thing, and
it did not work. It reads a seal key from the environment that nothing sets, and it seals every
credential it produces to that key with a symmetric passphrase.
Tracing the credential's path turned up something larger than a missing key. **The provisioner
harness the whole catalogue is built on describes a credential flow the mesh does not have, and
duplicates — incorrectly — one it does.**
What the sdk's `runProvisioner` does today:
- reads request files named `*.grant.json` — which nothing in the mesh writes;
- calls an adapter whose `create` **generates its own password** and returns it;
- seals that password with a symmetric key (`$MESH_SEAL_KEY`) and writes a `*.credential`
file — which nothing in the mesh reads, and no consumer ever unseals.
What the mesh already does, and has wired end to end:
- The control plane mints one password per (consumer, provider) pair (`Inventory.SecretFor` →
`secrets.Make`) and seals it to **both** node keys asymmetrically — a copy the consumer's
host can open and a copy the provider's host can open. No shared symmetric key exists
anywhere, on purpose: a key both ends hold is a key the mesh would have to distribute, which
is the same problem one level down, and the control plane's own code refuses it.
- The provider is handed, at the path its `receives` names, one contribution per consumer:
the **login to create** (`As`, derived by the mesh so the two ends agree by construction),
the consumer's address and requested values, and a **`Secret` file** holding that consumer's
password sealed to the provider and unsealed onto the machine by its host.
- The consumer is handed the *same* password, as plaintext its own host wrote by unsealing its
copy and substituting it into a config file. The consumer never unseals anything itself and
holds no key.
So the password a provider's provisioner invents is not even the password the consumer was
given: a consumer authenticating with the mesh's password against a resource the provisioner
created with its own would simply fail. The symmetric seal is not an incomplete feature to
finish delivering a key for. It is a second, contradictory credential model bolted beside the
real one, and it cannot be made to work without building the very thing the mesh was designed
not to have.
This is a decision and not a patch because the harness is the **provider contract**. Every
provider — the four that exist and the many a real mesh grows — is built on
`runProvisioner(resource, adapter)`. Whatever it says a provider is, they all inherit; and
changing it later is one migration per provider. It is cheaper and more honest to settle what
a provider is now.
## Decision
**A provider is handed the credential; it does not make one, does not seal one, and does not
hand one back.** The provisioner's only job is to make the mesh's grants true in its own
software.
Concretely, for the sdk harness and the adapter contract:
- The harness reconciles the **contributions the mesh delivers** to the provider's `receives`
path — the list of consumers, each with its login name (`As`), address, requested values,
and the path to its unsealed password (`Secret`). It does not read `*.grant.json` and it
does not write `*.credential`.
- For each consumer present, the harness reads the password from that consumer's `Secret` file
and calls the adapter to bring the resource into being under the given login. For each
consumer no longer present — the mesh drops it from the contributions file when its consumer
goes away — the harness calls the adapter to withdraw it.
- The adapter shrinks to the per-software half and nothing else. It is given the login, the
password, and the values, and it makes the resource exist or removes it. It generates no
password, derives no name, seals nothing, and returns no credential:
roughly `create({ as, password, values })` and `remove({ as })`, both returning nothing.
- `$MESH_SEAL_KEY`, the symmetric `seal()`/`writeSealedCredential` path, and the `*.grant.json`
/ `*.credential` files are removed from the provisioning path entirely. The credential
reaches the consumer through the mesh's own asymmetric channel, which already crosses node
boundaries and holds no shared secret.
Identity stays the mesh's to say. The login the provider creates is the name the mesh derived
and gave the consumer to present; the provider never invents a name, because a name the
consumer cannot learn is a name it cannot authenticate with.
## Consequences
- A provider module becomes smaller and unable to be wrong in this way: with no password to
generate and no key to seal to, the class of bug where the two ends hold different secrets
cannot be written. A provider added after this inherits the corrected contract and has no
seal to reintroduce.
- The four current providers (redis, postgres, minio, umami) each lose their `generatePassword`
+ seal code and gain a `create` that takes the password it is given. Their teardown becomes
"withdraw the login named `As`".
- The symmetric `seal()`/`unseal()` primitive loses its only caller and leaves — checked, not
assumed: nothing else in the sdk or the catalogue called it, so it is removed with the
provisioner it belonged to.
- **How this is verified:** redis is assigned as a provider in the lab, the contributions and
the unsealed password the mesh would deliver are put in its `receives` path, and a client
authenticates as that consumer with the mesh's password and gets PONG — where a provider that
invented its own password answers WRONGPASS — with `$MESH_SEAL_KEY` set nowhere and no
`.credential` file written. Proven: `provider-uses-mesh-credential` is green.
**What this does not cover — credential provisions, not data provisions.** This decision is about a
provision whose credential is a *secret the mesh mints* — a login and password (redis, postgres,
minio). A provider that instead *generates* the thing the consumer needs, and that thing is not a
secret — umami's `analytics`, where the consumer wants back a `siteId` umami assigned — does not fit,
because a contract that returns nothing has no way to hand that data back. The seal-key fault was
never umami's (it sealed no password; it returned a public id), so removing the seal does not break
it further, and it still reconciles its sites off the mesh's contributions. But delivering
provider-generated data back to a consumer is a *return path* the mesh does not have and this
decision does not build — a separate shape, left to a separate decision.
- Teardown beyond "remove the login" — data an object store leaves behind when a consumer
leaves — is named by each provider's adapter, not by the harness, and is out of scope here
except to say the contract must leave room for it.
## References
- [04-ISSUES/032](../04-ISSUES/032-provider-runtime-has-no-seal-key/00-report.md) — the
observation and the cross-repo trace this decision rests on.
- ADR 0047 (the module runtime) — what first ran a provider's provisioner as a delivered
process and exposed this; link to be filled when 0047 lands on the trunk.
- Control-plane mechanisms this relies on already existing: `mesh-control` —
`internal/inventory/secrets.go` (`SecretFor`, `SecretsFrom`), `internal/secrets/seal.go`
(`Make`, the two-blob asymmetric sealing), `cmd/mesh-control/plan.go` (`grantsFor`, the
`Grant.Sealed = ForProvider` delivery), `internal/catalogue/declaration.go` (the `receives`
contribution: `As`, `At`, `Values`, `Secret`).
- The path being removed: `mesh-sdk` — `src/provisioner/index.ts` (`runProvisioner`, `sealKey`,
`writeSealedCredential`) and the symmetric `src/primitives/index.ts` `seal()`/`unseal()`.
@@ -0,0 +1,130 @@
---
topic: what runs on it
status: accepted
date: 2026-09-05
deciders: jochen
reconstructed: false
---
# 49. A consumer's identity is bounded by the tightest backend that must accept it
## Context
The mesh says who a consumer is, once, and hands the same name to the provider (to create) and the
consumer (to present), so the two ends agree by construction rather than by two conventions (the
principle behind `ConsumerIdentity`, 04-ISSUES/023). The name is `mesh_<node>_<module>`, cleaned to
lower-case letters, digits and underscore.
Proving the provider contract per backend (ADR 0048) turned up 04-ISSUES/034: redis and postgres
create that name verbatim, but **minio refuses it** — an S3 access key is capped at 20 characters,
and `mesh_anchor_bucketuser` is 22. The provisioner then retries for ever, per consumer, and the
consumer holding that same too-long name could never present it either.
Two things about the existing derivation decide most of this:
- **The charset is already right.** `[^a-z0-9_]` is deliberately conservative, and its own comment
says it reaches "a PostgreSQL role, a MinIO access key, an LDAP uid and a Keycloak client without
quoting." That much is true.
- **The length is wrong.** `CheckIdentity` refuses names over `identityLimit = 63`, commented as
"the shortest identifier limit among the systems these names reach: PostgreSQL's". It is not the
shortest — S3's 20 is shorter — so the guard that was meant to catch exactly this lets it through,
and the failure lands at provision time as a silent retry instead of at assignment as a refusal.
So this is a small wrong constant with a real cost attached: whatever bound we set, `mesh_` (5) plus
a node name plus `_` plus a module name has to fit inside it.
## The options
**A — Bound the identity by the true minimum, and refuse early.** Lower `identityLimit` to the real
shortest (20, S3's), so `CheckIdentity` refuses an over-long name *at assignment* with a clear
message, the way it already refuses over-63 names. The derivation does not change; long names are
simply rejected before anything is provisioned.
- *For:* smallest change; keeps "the mesh says the identity once, verbatim" intact; the failure
moves from a per-consumer provision-time retry to an up-front, legible refusal — which is what
`CheckIdentity` exists to do.
- *Against:* a hard budget. `mesh_` + node + `_` + module ≤ 20 means node + module ≤ 14 characters.
`anchor` + `bucketuser` (16) is already over. It pushes the constraint onto how machines and
modules are named, which is a real limitation on legible names.
**B — Keep the readable name when it fits, compact it when it does not.** Below the bound, the name
is `mesh_<node>_<module>` as today; over it, the mesh substitutes a deterministic short form (e.g.
`mesh_` + a truncated hash of node+module) — still one derivation, so both ends still agree.
- *For:* no naming constraint; short backends always satisfied; the common case stays legible.
- *Against:* some identities become opaque, and a provisioner tracing "whose login is this" loses
the answer for exactly the consumers that overflowed. The mesh now owns a fallback format and its
collision properties (a truncated hash is not free of collisions at 15 characters).
**C — Let each interface declare its identifier bounds, and derive within the tightest a consumer
reaches.** `s3-bucket` states `identifier: { max: 20 }`; `postgres-database` states 63; the mesh
derives a name that fits the **minimum** bound across the providers a given consumer is granted.
- *For:* the most precise — each provision gets exactly the room it has, and a database consumer
keeps long legible names while an S3 consumer gets a short one; the constraint lives where the
fact does (on the interface).
- *Against:* the most work, and a consumer of two interfaces with different bounds must satisfy the
smaller — so its name shortens for both, reintroducing B's opacity in a narrower case. It also
means one consumer can hold **different** identities per provision, which the "said once" model
currently forbids.
**D — Let the provider generate a backend-valid identity and hand it back (rejected).** minio mints
its own access key and returns it to the consumer. This is the data-provision return path this era
keeps meeting — but it directly contradicts 023 and ADR 0048: the identity would no longer be the
mesh's single derivation the two ends share, it would be a value one side invents and the other must
be told. Listed for completeness; not recommended.
**E — A module (and a node) may declare a short slug; the identity is built from it.** The identity
becomes `mesh_<node-slug|node-name>_<module-slug|module-name>`: where a slug is declared it is used,
otherwise the cleaned name. A slug is a deliberately short, operator-chosen identifier — `kc` for
keycloak, `wkstn` for a workstation. It is optional: short names (`anchor`, `redis`) need none.
- *For:* this is the escape hatch B wanted to be, without the opacity. The name stays legible — a
provisioner can read `mesh_wkstn_kc` and know who is asking — because a person chose it, not a
hash function. And it makes an early refusal *palatable*: if even the slug-built identity overflows,
the refusal points at the slug, a field made for exactly this, rather than at the machine's name.
Both ends still derive it from one declared thing, so they agree by construction.
- *Against:* a new optional manifest field, and someone must pick the slug — but only for names that
would otherwise overflow, and picking a short legible identifier is a better job than being handed
a hash.
## What implementing A revealed
A was tried first. At `identityLimit = 20`, the readable budget is `mesh_` (5) + node + `_` + module
≤ 20, i.e. **node + module ≤ 14 characters** — far tighter than it looked. The catalogue's own
existing tests use `workstation`+`keycloak` (25), which compacts to `mesh_dbbc02f8dde34d3`; common
mesh names (`home-server`, `the-build-node`, `workstation`) blow the budget with any module. So B's
compact fallback would fire for the *common* case, not the rare overflow — which inverts A+B: most
identities would be opaque hashes. A alone (hard refusal at 20) would refuse most realistic names.
This is what moved the recommendation to E: the problem is not the limit, it is that the *readable
name* is the wrong source when it is long, and a slug is a better source than either a hash or a ban.
## Recommendation
**E, over a per-consumer bound (start with the global minimum, 20).** Build the identity from an
optional slug, keep it when it fits, and refuse at assignment with "declare or shorten `<module>`'s
slug" when it does not — no hash, no lost legibility, and the fix is a first-class field. Set the
bound to the true minimum (20) now; it needs no per-interface machinery to unblock S3, and a module
that consumes S3 simply declares a short slug. Graduate to **C** (per-interface bounds) later if it
turns out that non-S3 consumers are paying for S3's limit often enough to mind — E and C compose:
slugs are the mechanism, per-interface bounds refine where the ceiling sits. **B is dropped**: a
declared slug is a strictly better escape hatch than an opaque hash. **D stays rejected.**
## Consequences (of E)
- A module manifest gains an optional `slug`; a node may carry one too. `ConsumerIdentity` prefers
the slug over the cleaned name for each half. `identityLimit` becomes 20 (the true minimum), and
`CheckIdentity` refuses at `module add` / assignment — now with a message naming the slug to set.
- The common case stays legible; only names that overflow the budget need a slug, and what they get
is a name a person chose, not a hash.
- Existing modules/nodes whose names overflow declare a slug once — a migration cost paid as a clear
refusal with an obvious remedy, not a silent hash or a silent truncation.
- minio (04-ISSUES/034) is unblocked: an S3 consumer declares a short slug and its access key fits.
- **How it is checked:** the minio grant e2e — a consumer whose (slugged) identity fits reaches its
bucket with the credential the mesh delivered — plus unit tests that a slug is preferred, that an
un-sluggable over-long identity is refused (naming the slug), and that two consumers never collide.
## References
- [04-ISSUES/034](../04-ISSUES/034-mesh-login-exceeds-s3-access-key-limit/00-report.md) — the
observation.
- ADR 0048 — a provider creates the credential the mesh minted; the identity it creates it under is
the one this decision bounds.
- `mesh-control` `internal/catalogue/identity.go` — `ConsumerIdentity`, `identityUnusable`,
`identityLimit`, `CheckIdentity` — where the constant and the check live.
+13
View File
@@ -107,6 +107,17 @@ python3 00-META/checks/index.py fail if stale
- **0026** — [The mesh has a session of its own, and it is the node session's mechanism](0026-the-mesh-has-a-session-of-its-own.md) - **0026** — [The mesh has a session of its own, and it is the node session's mechanism](0026-the-mesh-has-a-session-of-its-own.md)
- **0027** — [A provision names what the consumer is coupled to, not the role it plays](0027-a-provision-names-what-the-consumer-is-coupled-to.md) - **0027** — [A provision names what the consumer is coupled to, not the role it plays](0027-a-provision-names-what-the-consumer-is-coupled-to.md)
- **0035** — [One implementation, several surfaces, and what that costs](0035-one-implementation-several-surfaces.md) - **0035** — [One implementation, several surfaces, and what that costs](0035-one-implementation-several-surfaces.md)
- **0038** — [The mesh assigns the port, and a module does not care](0038-the-mesh-assigns-the-port.md) *(proposed)*
- **0040** — [What a module is](0040-what-a-module-is.md)
- **0041** — [Events are a relationship, the lighter sibling of provisioning](0041-events-are-a-relationship.md)
- **0042** — [The shape of an event on the wire](0042-the-shape-of-an-event-on-the-wire.md)
- **0043** — [A module's broker account is scoped by what it emits and consumes](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md)
- **0044** — [A public name is provisioned, not registered by hand](0044-a-public-name-is-provisioned-like-any-capability.md)
- **0045** — [A machine's firewall is the sum of what its modules listen on](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)
- **0046** — [A module's configuration is its assignment's, not its manifest's](0046-a-module-configuration-is-its-assignments-not-its-manifest.md)
- **0047** — [A module runs its code as its own process, with its own account](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)
- **0048** — [A provider creates the credential the mesh minted, and seals nothing](0048-a-provider-creates-the-credential-the-mesh-minted.md)
- **0049** — [A consumer's identity is bounded by the tightest backend that must accept it](0049-a-consumers-identity-fits-the-tightest-backend.md)
### How it is built ### How it is built
@@ -116,6 +127,8 @@ python3 00-META/checks/index.py fail if stale
- **0014** — [No workspace — each module is a standalone package consuming published dependencies](0014-no-npm-workspace.md) - **0014** — [No workspace — each module is a standalone package consuming published dependencies](0014-no-npm-workspace.md)
- **0015** — [Applications live in their own repository; the monorepo is for the mesh](0015-applications-live-in-their-own-repository.md) - **0015** — [Applications live in their own repository; the monorepo is for the mesh](0015-applications-live-in-their-own-repository.md)
- **0016** — [The lab](0016-the-lab.md) - **0016** — [The lab](0016-the-lab.md)
- **0037** — [Where a module lives](0037-where-a-module-lives.md) *(proposed)*
- **0039** — [What the SDK holds, and what it refuses](0039-what-the-sdk-holds-and-refuses.md)
### How it is checked ### How it is checked
@@ -1,9 +1,9 @@
--- ---
status: resolved status: fixed
opened: 2026-08-22 opened: 2026-08-22
located-in: [mesh-control] located-in: [mesh-control/internal/catalogue, mesh-catalog/modules/firewall]
fixed-by: mesh-control — a machine's filtering is computed from what it was assigned fixed-by: the manifest refuses unknown keys, `from` is the field that scopes a port and it is rendered to nftables, and the firewall module applies it
amended-design: 03-DESIGN/01-to-be/08-connectivity.md amended-design: 0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md
--- ---
# 003 — A firewall rule's `scope:` is read by no code # 003 — A firewall rule's `scope:` is read by no code
@@ -0,0 +1,135 @@
---
status: resolved
opened: 2026-09-04
located-in: [mesh-sdk, mesh-catalog]
fixed-by: mesh-sdk src/provisioner rework + redis/postgres/minio/umami adapters (ADR 0048)
amended-design: 0048-a-provider-creates-the-credential-the-mesh-minted.md
---
# A provider's provisioner seals with a key the mesh has no way to deliver — and does not need to
## What was observed
Building the vertical slice for the module runtime (the module runs its own code as its own
process under its own account), a **provider** module — one that stands up a per-consumer
resource and hands back a credential — was assigned to a node and run as a broker-bound
runtime. The runtime hosts the module's provisioner (the sdk's `runProvisioner`), and the
harness opens by reading a **seal key** from `$MESH_SEAL_KEY`, failing immediately without
one. Every credential it produces for a consumer is sealed to that key with the sdk's
symmetric `seal()` (AES-256-GCM, `mesh-sdk/src/primitives/index.ts`) before being written.
Nothing in the mesh sets `$MESH_SEAL_KEY`. It is read in exactly two places in the sdk and
set nowhere — no manifest, no control-plane code, no host code. So a provider runtime, as
delivered, aborts at start-up. The slice proved the mechanism only by setting a lab-local key
in the manifest by hand.
## What a trace of the credential path turned up
The seal key is not a missing delivery. **The whole symmetric-seal provisioner is orphaned,
and it duplicates — badly — a job the mesh already does.**
- `runProvisioner` reads request files named `*.grant.json`. **Nothing writes those.**
- It writes sealed credential files named `<consumer>.<resource>.credential`. **Nothing reads
those** — not the host, not the control plane. The host reports applied-resource digests
upward and never ships credentials; the control plane has no reference to that filename.
- No consumer ever calls the symmetric `unseal()`. Consumers receive **plaintext**.
Meanwhile the mesh already carries a provider→consumer credential across nodes, with **no
shared key anywhere**:
- The control plane mints the password once (`secrets.Make`) and seals it **twice,
asymmetrically** — `ForConsumer` to the consumer node's X25519 public key, `ForProvider` to
the provider node's (`mesh-control/internal/secrets/seal.go`, `mesh-host/internal/identity/
sealing.go`, NaCl box).
- Each host opens its own copy with its own private key on the machine; the plaintext exists
only for the length of one function call (`mesh-host/internal/apply/apply.go`, the
`${secret:name}` substitution — ADR 0024's "the host is the only thing that ever holds
both").
- `serves` carries no credential and says so; `receives`/`bound` tell each side *where* its
sealed secret is, never the value.
The two models also **contradict** each other. The sdk's `seal()` comment says the key is "a
per-node passphrase the host holds"; the host holds no such passphrase — it holds an X25519
private key, and the control plane's own code refuses a shared symmetric key on principle:
"a key both ends hold is a key the mesh would have to distribute, which is this problem again
one level down" (`secrets/seal.go`). A symmetric `MESH_SEAL_KEY` shared between a provider
node and a consumer node is exactly the thing the mesh was built not to have.
And the provisioner's model is wrong in a second way: its adapter **generates its own
password** (`generatePassword()`) and creates the resource with it — a different password from
the one the mesh mints and hands the consumer. Even with a seal key delivered, a consumer
would authenticate with the mesh's password against a resource created with the provisioner's.
## Why it matters beyond this instance
This is not a four-module problem. The provider contract lives in **one place** — the sdk's
`runProvisioner(resource, adapter)` harness — and every provider is built on it. Four exist
today (redis, postgres, minio, umami); a mesh of any size ends up with many. Whatever the
provisioner harness does, every present and future provider inherits, so the orphaned
symmetric seal is a fault stamped into the interface, not into four adapters. That also sets
the cost of getting it wrong: a contract N providers depend on is N migrations to change
later, which is the argument for settling it deliberately now rather than patching around it.
As written, each provider carries a provisioner that cannot start (no key), and that, if it
did, would create resources with a password it invented — a *different* password from the one
the mesh minted and handed the consumer — and seal them for a reader that does not exist. The
rule the design states, "a consumer receives a sealed credential and unseals it," is enforced
by nothing: no consumer unseals, and no shared key exists to unseal with.
## The mesh already does this — confirmed
The premise the fix rests on is not a hope; it is in the control plane today. For a served
interface, `Inventory.SecretFor` mints one password per (consumer, provider) pair via
`secrets.Make`, sealing it to **both** node keys — `ForConsumer` and `ForProvider`.
`SecretsFrom(provider)` is documented as "every credential a provider node was issued, so it
can be told what to create," and `grantsFor` (plan.go) hands the provider node one `Grant` per
consumer carrying `Sealed: ForProvider`. The provider receives, at the path its `receives`
names, one `Contribution` per consumer: the login to create (`As`, derived by the mesh so both
ends agree — 04-ISSUES/023), the consumer's address (`At`) and requested `Values`, and
`Secret`, the file holding that consumer's password sealed to this provider and unsealed by
its host. Everything the provisioner needs is delivered. It reads the wrong files
(`*.grant.json`, which nothing writes) and invents a password instead of reading the one in
`Secret`.
## The fix this points to
A **one-place contract change in the sdk harness**, plus re-pointing today's adapters at it —
not per-provider surgery, and inherited correctly by every provider after them:
- `runProvisioner` reconciles the mesh-delivered `receives` contributions (not `*.grant.json`):
for each consumer, create the resource under the login `As` with the password read from the
delivered `Secret` file, for its `Values`; withdraw the login when a consumer leaves the file.
- The adapter stops generating a password and stops returning a credential — it is handed the
name and the password and only makes the resource exist. Roughly `create({as, password,
values})` / `remove({as})`, no return.
- `sealKey`, `seal()`, `writeSealedCredential`, `MESH_SEAL_KEY`, and the `.credential` file
leave entirely; the consumer already receives its copy through the mesh's own channel.
This is proposed as ADR 0048, which defines the corrected provider contract, for ratification.
## Resolution
ADR 0048 was accepted and implemented on the branches this issue is fixed by:
- `mesh-sdk` `src/provisioner/index.ts` now reconciles the mesh's `receives` contributions and,
per consumer, reads the mesh-minted password from the file the host unsealed, calling the
adapter to create the resource under the mesh's login. `$MESH_SEAL_KEY`, the symmetric seal,
`writeSealedCredential`, and the `*.grant.json` / `*.credential` files are gone. The symmetric
`seal()`/`unseal()` primitive had no other caller and was removed.
- The four adapters (redis, postgres, minio, umami) were re-pointed at the new contract —
`create({ as, password, values })` / `remove({ as })`, returning nothing. minio's client gained
a secret-key argument so it sets the mesh's secret rather than generating one.
- Proven in the mesh-lab: `provider-uses-mesh-credential` is green — redis creates the consumer's
login with the password the mesh minted, a client authenticates as that consumer and gets PONG,
with no seal key set anywhere.
Two things were carved out deliberately, neither blocking:
- **Data provisions are a separate shape.** umami's `analytics` returns a `siteId` umami
*generates*, not a secret the mesh mints, and a contract that returns nothing cannot hand that
back. ADR 0048 is scoped to credential provisions and says so; the provider→consumer return
path for generated data is left to a separate decision. umami compiles and reconciles under the
new harness; only that return is unaddressed, and it never had the seal-key fault.
- **Teardown beyond "remove the login"** — an object store's leftover data — is each adapter's to
name (minio leaves a non-empty bucket for an operator rather than deleting a consumer's data),
not the harness's.
@@ -0,0 +1,67 @@
---
status: open
opened: 2026-09-04
located-in: []
fixed-by:
amended-design:
---
# Changing a module's settings does not restart its runtime — config is stale until recreated
## What was observed
Rolling the module runtime out to the catalogue (the runtime that serves a module's tools and
runs its events under the module's own account), each tools+events module receives its
configuration the way the design intends: a mergeable config file the module declares, into
which the assignment's settings are merged. The runtime container mounts that file and reads
it once at start-up, when it builds its API client.
The design for settings says a config file a module owns can be changed **without editing
it** — a person states an intention, the file is regenerated, and the change takes effect.
The decision that config is the assignment's, not the manifest's, is explicitly so that
configuration can be updated *on the fly* and managed from a dashboard.
For a runtime delivered as a **container**, that last part does not hold. When settings
change, the control plane re-renders the config file on the node — but the runtime container
is only ever recreated when its **spec** changes, and the spec is image, name, env, ports,
volumes and args. The *content* of a mounted file is not part of it. So the file on disk
updates and the process that already read it keeps the value it read at start-up. The new
configuration does not take effect until something changes the container's spec, or it is
recreated by hand.
A **service** resource has `restart-on`, which names the resources whose change forces a
restart — exactly this problem, already solved, for units. A **container** resource has no
equivalent field, and the apply path for containers never consults the set of resources that
changed this pass. So the one kind of resource that hosts a module's runtime is the kind that
cannot say "restart me when my config changes."
The effect is quiet, which is the worst part: setting a value appears to succeed (the file is
correct on disk), and the running tools keep answering with the old configuration, or keep
failing to load because the value that would fix them is present but unread.
## Why it matters beyond this instance
Every tools+events module converted to the runtime model now takes its URL and credentials
this way, so this is not one module's quirk — it is the config path for the whole catalogue.
The gap turns the headline promise of the settings design ("change it without editing it, on
the fly") into "change it, then recreate the container by hand," which is the manual step the
design existed to remove. And because the file is genuinely updated, nothing surfaces the
staleness; a dashboard that set the value would report success while the mesh kept doing the
old thing.
Config set **before** the runtime first starts (settings, then assign, then push) does work —
the file is right when the process reads it. So the gap is specifically about *updates* to an
already-running runtime, which is precisely the case the "on the fly" promise is about.
## Open questions
- Should a `container` gain `restart-on`, mirroring the service field, so a module can point
it at its config resource?
- Or should the apply path recreate a container when a file it mounts changed this pass —
making mounted-file content behave like part of the spec, without a new field to declare?
- Should the config file's content (or a hash of it) fold into the container spec, so an
ordinary spec-diff already catches it? That restarts on every change with no new mechanism,
at the cost of a spec that is no longer only the container's own declaration.
- Is a restart even the right primitive for a runtime that could instead watch its config
file and rebuild its clients in place — and if so, is that each module's job or the
runtime host's?
@@ -0,0 +1,77 @@
---
status: resolved
opened: 2026-09-05
located-in: [mesh-control, mesh-catalog]
fixed-by: ADR 0049 (a slug for the login) + a shorter minted secret (mesh-control)
amended-design: 0049-a-consumers-identity-fits-the-tightest-backend.md
---
# The mesh's derived login does not fit every backend's identity rules — S3 rejects it
## What was observed
Proving the provider/consumer contract per backend (ADR 0048), redis and postgres passed: a
consumer authenticated against the provider with the login the mesh derived and the password
the mesh minted. **minio failed**, and not on the credential — on the *name*:
```
mc: <ERROR> Unable to add a new service account. The access key is invalid.
(access key length should be between 3 and 20).
```
The mesh derives a consumer's login as `mesh_<node>_<module>` — here `mesh_anchor_bucketuser`,
22 characters. That is a valid postgres role and a valid redis ACL user, so those providers
create it verbatim. S3 access keys are capped at **20 characters**, so minio refuses to create
the service account under it, and the provisioner retries forever while the consumer, holding
that same too-long access key, could never present it either.
## Why it matters beyond this instance
ADR 0048 says a provider creates *exactly* the login the mesh derived, so that the two ends
agree by construction — the mesh hands the same name to the provider (to create) and the
consumer (to present). That only holds if the derived name is one every provider can accept.
It is not: the mesh's `as` is a single format with no knowledge of a backend's identity rules,
and S3's are stricter than a database's. Any provider whose backend constrains identifiers more
tightly than postgres — a length cap, a charset, a required prefix — inherits this, and the
failure lands at provision time, per consumer, as an infinite retry rather than a refusal at
assignment.
This also shows the seam is real, not cosmetic: `as` is doing two jobs — a stable per-consumer
identity the two ends must agree on, and a literal identifier a specific backend must accept —
and those are not always the same string.
## The shape of a fix (open, not decided)
- **Constrain the derivation** so `as` is broadly acceptable — short (≤ 20), a conservative
charset, deterministic. This keeps "the provider creates exactly what the mesh derived" true
everywhere, at the cost of a less legible name, and it is a mesh-wide identity change (every
provider that already created the longer name would see it change).
- **Let a provider map `as` to a backend-valid identifier** it derives the same way on create
and on the consumer's behalf — but the consumer is generic and cannot run minio's mapping, so
this only works if the mapped identifier is *delivered back* to the consumer. That is the
data-provision return path this era keeps meeting (umami's siteId, cloudflare's record) and
does not yet have.
- **Declare the constraint on the interface** (`s3-bucket` states its identifier bounds) and
have the mesh derive within them — the most honest, the most work.
## Open questions
- Is `as` meant to be human-legible, or is a short opaque token acceptable — i.e., can the
derivation simply be shortened without anyone minding?
- Do redis/postgres actually want the long name, or did it only survive because they are
permissive? If nothing needs it long, the cheap fix is to cap it.
- Does this fold into the same decision as the data-provision return path, or is it separate?
## Resolution
Accepted **ADR 0049** (option E): a module declares an optional short `slug`, and the mesh derives
`mesh_<node>_<slug|name>`, bounded by the tightest backend (an S3 access key's 20) and refused at
assignment — naming the slug as the remedy — when it still would not fit. The minio grant e2e proved
it: `bucketuser` declares `slug: bkt`, so its access key `mesh_anchor_bkt` (15) is accepted where
`mesh_anchor_bucketuser` (22) was refused.
Proving that surfaced a **second S3 length constraint on the same credential** — the secret. The
mesh minted a 43-character password (32 random bytes, base64url), and an S3 secret key is 8–40. Fixed
in `mesh-control` `internal/secrets/seal.go` by minting 30 bytes → exactly 40 characters (240 bits,
ample), which fits S3 and every other backend. Both halves of an S3 credential — the access key
(login) and the secret key (password) — now fit the tightest backend, by the same rule.
@@ -0,0 +1,48 @@
---
status: open
opened: 2026-09-02
located-in: []
fixed-by:
amended-design:
---
# 035 — Reconciling a seed file wipes what grew in it
## The symptom, as observed
Found by review of the catalogue examples (2026-09-02), not by an outage — the outage is the
part the design permits to be silent.
The cache module declares its access-control file as an ordinary file resource with fixed,
empty content. The program that consumes the file requires it to exist at startup, which is
why the manifest declares it at all. But the same file is the one the provisioner writes
consumer users into, and the one the running program persists ACL changes back to.
A declaration is complete for what the host owns, and the host reconciles what is declared
([ADR 0010](../../02-DECISIONS/0010-delivery.md)).
So every apply that revisits this resource restores the declared content — empty — behind the
running program. Every consumer credential granted since the last apply is removed, the apply
reports success, and nothing anywhere says a grant vanished.
## Why it matters beyond the instance
The manifest needed *the file to exist before first start*, and the only vocabulary available
was *the file has this content, forever*. Those are different intentions, and the gap between
them is generic: any resource that a module seeds and something else then legitimately mutates
— an ACL file, a bootstrap configuration a program rewrites, an htpasswd a provisioner appends
to — has the same two owners and the same silent loss on reconcile.
It is also the mirror image of the boundary ADR 0010 draws so carefully on the *removal* side:
the host never removes what it did not create, but it happily overwrites what it *did* create,
even when what grew inside since is somebody else's work the mesh asked for.
## Open questions
- Is the missing thing a create-once file semantic ("present with this content if absent,
untouched otherwise"), or is the real fault that two owners share one file — and the
provisioner, not the declaration, should own it entirely, with first-start ordering solved
some other way?
- ADR 0010 treats every added resource type as a security artefact. Does a create-once
semantic widen what a compromised control plane can express, or narrow it?
- Are there other seeded-then-mutated files already in the catalogue that this failure is
waiting inside?
@@ -0,0 +1,51 @@
---
status: open
opened: 2026-09-02
located-in: []
fixed-by:
amended-design:
---
# 036 — Six modules own what they must share
## The symptom, as observed
Found by review of the catalogue examples (2026-09-02). The media stack is several modules —
a library server, the acquisition managers, a download client and their satellites — and each
of them declares the same library and download directories as its own resources.
The resolver refuses two modules that declare one path on one node, with no exemption for
identical content and no merge. That rule is right in general: two owners of one path is the
class of fault this repository keeps recording. But sharing those directories on one machine
is the entire point of this stack — the download client and the managers must see the same
downloads, the library server must see the same libraries. So the set, as written, refuses
its own only sensible assignment.
No test co-resolves any two of them, which is why the manifests pass today. The first machine
to be assigned the stack together is where the refusal would have surfaced.
## Why it matters beyond the instance
The manifests can express *a directory I own* and nothing else, so a directory that is the
shared workspace of several modules was written six times as six private ones. The intention
— several modules, one filesystem contract between them — has no vocabulary, and this is not
a media-stack peculiarity: any pipeline of modules handing files to each other on one machine
(an ingest directory, a spool, a drop folder) hits the same wall.
It is also a fork in the design the catalogue has otherwise avoided: the fix could be a new
owning module the others depend on, a shared-resource concept in the manifest, or a statement
that co-located file handoff is not a thing the mesh supports and these modules are one
module. Each answer changes what a module *is*, which is why this is an issue and not a patch.
## Open questions
- Is the unit wrong — is a stack that must share a filesystem one module with several
containers, the way the mail module already is?
- If it stays several modules: does one of them own the directories and the rest require
them, and is *requiring a directory from a neighbour* a provision, a claim, or a third
thing?
- The duplicate-path rule protects against genuinely rivalrous owners. Whatever expresses
sharing must not weaken it for the cases where refusal is the right answer — what
distinguishes the two, machine-checkably?
- The mesh's own rule is that a rule states how it is checked: whichever shape is chosen,
what test co-resolves the stack so this class of refusal is caught before a machine is?
@@ -0,0 +1,57 @@
---
status: open
opened: 2026-09-05
located-in: []
fixed-by:
amended-design:
---
# 037 — A module cannot run its own code at a lifecycle phase
## The symptom, as observed
Found while converting the catalogue (2026-09-05), across several modules at once. A module can
declare *things that exist* — a directory, a file with fixed content, a network, a container — but
it cannot declare *a step that runs* at a defined point in its own lifecycle. Three converted
modules need exactly that and have nowhere to put it:
- **mosquitto.** Its Dynamic Security plugin will not start unless `dynamic-security.json` already
contains an admin client *before the broker's first start* — the broker loads the plugin at
boot. Seeding it is a run-once step that must happen after the file resource exists and before
the container starts. The vocabulary has no "before first start."
- **The database providers (postgres/mongodb/mssql).** First-boot seeding works today only because
the *image* happens to do it from an env var. Anything the mesh itself must run once against the
server — a schema migration, an extension enable, a health gate before the module is announced
ready — has no home.
- The seed-then-mutate family already recorded in [035](../035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md)
is the same shape seen from the *content* side; this is it seen from the *timing* side.
## Why it matters beyond the instance
This is not a defect in a module — it is a **capability the module system does not yet offer.** A
real class of modules needs to run their own code at points in the build/install/run lifecycle:
seed-before-start, migrate, post-start health-gate, pre-remove drain. The declarative resource
model deliberately describes *state*, not *steps*, and that is right for what it covers; the gap is
that some modules genuinely have a step.
**Prior art, and its warning.** An earlier mesh had exactly this as a feature: event-driven
**hooks** that ran custom code at phases of the build/publish/deploy pipeline. It was powerful and
it was **complex to set up and flaky** — which is the real content of this record. The need is not
in question; the cost of the obvious answer is. Whatever shape this takes must not reproduce that
fragility, or it will be worse than the gap.
## Open questions
- Is the right unit narrow — a **run-once / init resource** ("run this once, here, in the
lifecycle") — or general — a **per-phase lifecycle hook** on a module, and if so which phases
(build / publish / install / pre-start / post-start / pre-remove)?
- Where does a hook's code run — in the module's own runtime container under its scoped account
(ADR 0043/0047), so it inherits the same isolation as its tools and events? Or is some of it the
host's, before a container exists?
- How is a step made **idempotent and reconcilable** so a re-apply does not re-run it
destructively — the same discipline the resource model gets for free and a step does not?
- What is the smallest version that unblocks the three modules above without rebuilding the old
flaky hook engine? Is "seed-before-first-start" alone enough for now, with the general case
deferred?
- A rule states how it is checked: whatever shape is chosen, what lab scenario proves a hook runs
exactly once, at the right phase, and converges on re-apply?