Merge pull request 'Issue 009 (fixed) + Issue 008 (resolved via ADR 0053) — module-runtime config & provider contract' (#21) from worktree-issue-provider-seal-key into main

This commit was merged in pull request #21.
This commit is contained in:
2026-09-05 03:02:34 +02:00
17 changed files with 1423 additions and 12 deletions
+1 -1
View File
@@ -31,7 +31,7 @@ target, not the present.
| `mesh-substrate` | 1 | the four pinned services, as declarations | | `mesh-substrate` | 1 | the four pinned services, as declarations |
| `mesh-control` | 2 | the control plane and its contexts | | `mesh-control` | 2 | the control plane and its contexts |
| `mesh-surfaces` | 3 | tools, web, cli | | `mesh-surfaces` | 3 | tools, web, cli |
| `mesh-sdk` | — | contracts shared across tiers | | `mesh-sdk` | — | the stable spine modules build against — the tool-serving harness, the messaging/event framework, the contracts and core primitives. Holds nothing per-module and nothing volatile ([ADR 0044](../02-DECISIONS/0044-what-the-sdk-holds-and-refuses.md)). |
| `mesh-lab` | — | **exists.** The lab — scenario lifecycle, networking, placement. Ships to nobody; runs on a workstation. | | `mesh-lab` | — | **exists.** The lab — scenario lifecycle, networking, placement. Ships to nobody; runs on a workstation. |
Tier 4's shape is open, and deliberately so: see ADR 0030 and Tier 4's shape is open, and deliberately so: see ADR 0030 and
@@ -1,5 +1,5 @@
--- ---
status: proposed status: accepted
date: 2026-08-23 date: 2026-08-23
deciders: jochen deciders: jochen
reconstructed: false reconstructed: false
@@ -0,0 +1,101 @@
---
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
supersedes: 0030-the-repository-structure.md
---
# 44. What the SDK holds, and what it refuses
## Context
[ADR 0030](0030-the-repository-structure.md) named `mesh-sdk` "contracts shared across tiers:
types, not behaviour." That line is superseded here, because it draws the boundary in the wrong
place. The boundary that matters is not *types versus behaviour* — it is **how often the thing
changes**.
The current SDK is the cautionary tale, and its failure is precise. `hal/sdk` holds all the
code, including a per-module API client for every service (`clients/plex.ts`, `clients/gitea.ts`,
…) and a per-module tool implementation for each (`tools/plex.ts`, …). Every module depends on
the SDK, so **every edit to any of that per-module code rebuilds every module** — the cascade.
The SDK is under constant maintenance precisely because it became the place all the volatile
per-module logic accumulated.
The root cause is worth stating exactly, because the fix follows from it: the pressure was never
to share a client *between* modules. It was to share a client between one module's *own features*
— plex's tools, its health check and its hooks all wanted the same `PlexClient` — and the only
place to share code across a module's features was the global SDK. So **intra-module sharing
leaked out as inter-module coupling.**
## Decision
The SDK holds the **stable spine** that modules build against, and earns its place by rarely
changing. The test for membership is change-frequency, not kind.
### What it holds
- The **tool-serving harness** — the worker and registration mechanism, and the tool-definition
type. *How* a tool is declared and served is settled; it does not change when an individual
tool does.
- The **messaging and event framework** — the broker client, the event consumer, the envelope.
- The **contracts** — the manifest, declaration, provision and link shapes.
- **Core primitives** — sealing and crypto, semver, the shared resolution helpers.
These change rarely and deliberately. When one of them does change, a rebuild of everything is
the *correct* outcome, because the contract every module shares has genuinely changed.
### What it must not hold — the more important half
- **A module's API client.** A Plex client, a Gitea client, a MinIO client belong in their
module. They change when that service's API or the module's use of it changes, which is often,
and which has nothing to do with any other module.
- **A module's tool implementations.** Same reason, same place: in the module.
- **Anything volatile** — anything that changes when one service's features change.
The rule, stated so it can be applied without re-deriving it:
> If editing a thing recompiles unrelated modules **and** it changes often, it does not belong
> in the SDK.
Both conditions are load-bearing. A rare change that cascades is fine — that is a contract, and
the cascade is correct. A frequent change that stays local is fine — that is a module minding its
own business. Only **frequent *and* cascading** is the disease, and per-module clients and tools
are its carriers.
### Where per-module shared code lives instead
Code shared among a module's *own* features lives **in the module**. The default is the plainest
thing that works: an ordinary shared file the features import — `plex/client.ts`, imported by
`plex/tools/`. Within one module, features are files importing sibling files; no package
boundary, no ceremony.
A **module-local SDK** (a sub-package with its own version) is warranted only for the few modules
whose shared surface is large enough to version on its own. It is the exception, not the shape.
Either form gives the property the global SDK could not: editing a module's shared code rebuilds
**that module and nothing else**.
## Consequences
- The cascade becomes **structurally impossible for module logic**. There is no longer an edge
from one module's internals to another, so the only thing that can rebuild everything is a real
change to a shared contract in the SDK — which is rare, and when it happens, is right.
- The SDK is small and stable **by construction**, not by discipline. Its size is no longer a
thing anyone has to police.
- **Converting a module from the current system is partly a de-coupling, not just a move.** Its
client and its tools are pulled *out* of the shared SDK and *into* the module. A conversion
that copied `clients/plex.ts` into the SDK's replacement would rebuild the exact mistake.
- The host still does not import the SDK. It depends on nothing
([ADR 0041](0041-the-host-depends-on-nothing.md)) and **mirrors** the contracts rather than
importing them, exactly as its apply-shapes table already does deliberately. The SDK is shared
by the tiers that *can* share code; the host is not one of them.
## References
- [ADR 0030](0030-the-repository-structure.md) — named the repositories; its `mesh-sdk`
description ("types, not behaviour") is superseded by this record.
- [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) — the decomposition this serves: code
belongs to the boundary that owns it.
- [ADR 0041](0041-the-host-depends-on-nothing.md) — why the host mirrors the contracts instead of
importing the SDK.
+98
View File
@@ -0,0 +1,98 @@
---
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
supersedes: 0017-modules-outside-the-core-are-grouped-by-domain.md
extends: 0002-everything-is-a-module.md
---
# 45. What a module is
## Context
[ADR 0002](0002-everything-is-a-module.md) settled that everything is a module, but never said what a
module *is* beyond "a directory the mesh processes." That gap let the catalogue's breadth read as a
smell: a module can carry a container, a built image, tools, a provisioner, migrations, health,
config, seat claims, requires and provides — so much that the unit seemed ill-defined.
[ADR 0017](0017-modules-outside-the-core-are-grouped-by-domain.md) tried to organise modules by
domain, which is the wrong axis. This record states what a module is, drawn from the cases that
stress-tested it: the shell, i3-vs-sway, umami, and "database."
## Decision
**A module is one self-contained piece of software the mesh installs and manages** — everything
needed to make that one thing real and integrable: what runs, the seats it claims, what it provides
to other modules, what it requires from them, and what operates it.
The **software is the module's identity.** Capabilities, seats and provisioned resources are the
**relationships *between* modules**, not what a module is — and that is what binds a module into one
thing. umami is bound by *being umami*: its container runs umami, its provisioner creates umami sites,
its tools query umami, its `requires` gets umami a database. Every feature serves the one software.
### The three relationships
1. **Shared seat** — several modules fulfil a capability and coexist; one may be default. bash, zsh
and fish all join `shell`.
2. **Exclusive seat** — modules contend for a single slot; one holds it. i3 (needs x11) and sway
(needs wayland) contend for `display-session`.
3. **Provide / require** — a provider ships the **provisioner** that creates instances of the
resource it offers and returns sealed credentials; a consumer requires it and the mesh wires the
credential in. Symmetric: umami requires a database *and* provides analytics.
### Interfaces are mesh-owned; providers adapt to them
The mesh **defines the interface** for a capability — the provider-neutral contract of what a
consumer receives and how it integrates. Both sides conform: a provider's provisioner **adapts** its
software's real API to the mesh contract; a consumer depends on the **interface**, never on a
provider. Swap one provider for another and the consumer does not change.
### The naming rule — draw the interface at the consumer's real coupling
Name a `provides`/`requires` at the **widest boundary across which the consumer genuinely does not
care which implementation serves it**:
- Where the consumer's coupling is thin — an analytics embed snippet and dashboard, opaque to it —
the mesh defines a neutral interface (`analytics`) and providers (umami, amumi) adapt. Swappable
across vendors.
- Where the consumer **speaks a protocol** — a database's wire protocol and query dialect — the
interface *is* the protocol: `postgres-database`, `mssql-database`, `mongodb-database`. Swappable
only among protocol-compatible implementations, **never across**, because the application cannot
cross it either. "database" is not a capability; the protocol is.
- **Never false genericity.** A name must not promise a swap the contract cannot deliver
([research 005](../01-RESEARCH/005-domain-grouping/analysis.md)).
This is [ADR 0027](0027-the-product-is-novox-mesh.md)'s rule made general — "names the protocol, not
the product; a database names the engine because the app targets it" — with the reason stated: the
contract sits where the coupling is.
### What is not a module
- A **library** (built against, never deployed — [ADR 0044](0044-what-the-sdk-holds-and-refuses.md)).
- A **control-plane context** (the mesh itself — [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md)).
A **swappable machine mechanism** (a firewall — ufw, nftables) *is* a module implementing a
capability. The host hardcodes no firewall, supervisor, package manager or runtime; it owns only the
generic apply primitives and platform detection, so it runs where none of those exist — an Android
phone has no ufw, systemd, pacman or Docker.
## Consequences
- **Supersedes [ADR 0017](0017-modules-outside-the-core-are-grouped-by-domain.md).** Modules are
organised by their relationships (seats, provisions), not grouped into domain folders.
- **Refines [ADR 0002](0002-everything-is-a-module.md).** Everything the mesh runs and integrates is
a module — but a module is defined by the *software it delivers*, not by being a bucket of features.
- The target is a **self-fulfilling mesh**: declared wants bound to swappable modules, provisioners
wiring credentials, nothing hardcoded. The control plane's whole job is the binding.
- Converting a module from the old system includes pulling its per-module code out of the shared SDK
([ADR 0044](0044-what-the-sdk-holds-and-refuses.md)) and shipping its provisioner as an adapter to a
mesh interface — a de-coupling, not just a move.
## References
- [ADR 0002](0002-everything-is-a-module.md) — everything is a module; this says what one is.
- [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) — contexts are the mesh, not modules.
- [ADR 0017](0017-modules-outside-the-core-are-grouped-by-domain.md) — superseded.
- [ADR 0027](0027-the-product-is-novox-mesh.md) — protocol-not-product, generalised here.
- [ADR 0044](0044-what-the-sdk-holds-and-refuses.md) — per-module code lives in the module.
- [research 011](../01-RESEARCH/011-the-module-graph/00-overview.md) — the graph of these relationships.
@@ -0,0 +1,84 @@
---
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
extends: 0045-what-a-module-is.md
---
# 46. Events are a relationship, the lighter sibling of provisioning
## Context
[ADR 0045](0045-what-a-module-is.md) names two relationships between modules — seats and
provide/require (provisioning). A third is latent in the mesh and worth making first-class: the
broker every node already runs ([ADR 0001](0001-nodes-communicate-over-a-broker.md)) can carry a
module's activity as **events**, which any other module reacts to. A logger that writes an audit
trail, a module that acts when another module acts, observability — all of it is one mechanism, and
today it is ambient rather than declared.
## Decision
**A module emits events and consumes events, and both are declared** — parallel to `provides` /
`requires`, so the mesh knows the event graph the same way it knows the provisioning graph.
### Events are provisioning's lighter sibling
| | provisioning | events |
|---|---|---|
| shape | **1:1**, a provider creates a resource *for* one consumer | **1:many**, a module emits, any number listen |
| credential | yes — sealed, per consumer | none — it is broadcast |
| machinery | a provisioner (the reconcile adapter) | nothing but the broker's topic routing |
| declared as | `provides` / `requires` | `emits` / `consumes` |
Because an event is broadcast and credential-free, there is no provisioner and no per-consumer
setup — only a subscription. That is why it is the *lighter* relationship, and why most
inter-module reaction should be an event, not a provision.
### An event carries what an audit needs
Every event carries its **type** (a dotted topic key, so listeners match by prefix), its **source**
module, the **node** it came from, and the **time**. A body follows. The metadata is not optional:
a reaction may only need the body, but an audit trail needs to know who did what, where and when,
and an event that cannot answer that is not auditable.
### The audit logger is just a consumer of everything
A logger that records the whole mesh's activity is **not a privileged component** — it is an
ordinary module that consumes `#` (every event) and writes them down. It holds no special access;
it only listens widely. That it falls out of the model with no new machinery is the check that the
model is right.
### `consumes` is validated like `requires`
A `consumes` for an event that **nothing** `emits` is a dangling edge, and the mesh refuses it
before deploy — the same rule that catches a `requires` for a resource nothing provides
([research 011](../01-RESEARCH/011-the-module-graph/00-overview.md)). A listener waiting for an
event that can never arrive is a silent failure, and this repository's whole discipline is against
silent failure.
### One runtime serves all three
The per-node module runtime that serves a module's tools also wires its `consumes` (subscribe,
dispatch to the handler) and lets its code `emit`. Tools are *invoked* (request/reply), resources
are *provisioned* (1:1, credentialed), events are *emitted and consumed* (1:many, broadcast) —
three relationships, one broker, one runtime, all declared on the manifest.
## Consequences
- The mesh gains a declared **event graph** alongside the provisioning graph — visible, validated,
reasoned over.
- **Reaction becomes the default coordination**: a module acts on another's event without either
knowing the other, and without a credentialed link. Coupling drops.
- An **audit trail** is a module, not a platform feature — and can be swapped, extended or run more
than once (a file logger and a queryable one) with no change to anything that emits.
- The runtime must dispatch a module's event handlers as well as its tools; that generalisation is
small (both arrive by importing the module's entrypoint) but it is real work.
## References
- [ADR 0001](0001-nodes-communicate-over-a-broker.md) — the broker events ride.
- [ADR 0045](0045-what-a-module-is.md) — the relationships this extends.
- [ADR 0044](0044-what-the-sdk-holds-and-refuses.md) — `emit`/`on` are stable sdk surface; the
broker binding and the runtime are not.
- [research 011](../01-RESEARCH/011-the-module-graph/00-overview.md) — the graph these edges join.
@@ -0,0 +1,115 @@
---
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
extends: 0046-events-are-a-relationship.md
---
# 47. The shape of an event on the wire
## Context
[ADR 0046](0046-events-are-a-relationship.md) made events a relationship — `emits`/`consumes`, the
graph, the audit logger. It did not say what an event *is* on the broker: the exchanges, the
routing keys, the headers, the queues and their configuration. That shape is a contract every
emitter and consumer conforms to, exactly as [ADR 0043](0043-a-declaration-is-an-ordered-list-of-owned-resources.md)
is for declarations — and it was being decided ad-hoc in code. This settles it, so the sdk and the
runtime implement one contract and a module never reinvents it.
## Decision
### Two exchanges, kept apart
- **`mesh.events`** — a durable topic exchange. Every event rides it: module, mesh and node.
- **`mesh.rpc`** — a durable topic exchange. Tool invocations (request/reply) ride it.
Kept separate because RPC is not an event: a `#` subscription on `mesh.events` is then a complete
audit of what happened, with none of the invocation traffic.
### The routing key is the event type, namespaced by origin
Dotted and hierarchical — `<origin>.<name>.<event…>` — with three reserved origins:
- `module.<module>.<event>` — `module.umami.site.created`
- `mesh.<context>.<event>` — `mesh.delivery.deployed`, `mesh.provisioning.granted`
- `node.<node>.<event>` — `node.anchor.joined`, `node.anchor.unreachable`
Topic matching gives a consumer `node.*.joined`, `module.umami.#`, or `#`. The origin roots are
reserved; everything after is the emitter's own namespace.
### Metadata in headers, payload in the body
An event's identity and provenance are AMQP **headers**, so a consumer — or the broker, or an
audit tool — reads who/when/what without parsing the body, and the body is only the domain payload.
**Required headers**
| header | meaning |
|---|---|
| `x-event-id` | a unique id — for dedup and audit (delivery is at-least-once, below) |
| `x-source` | the emitter: the module, context or node name |
| `x-node` | the node it was emitted from |
| `x-time` | emit time, RFC-3339 |
| `content-type` | `application/json` |
**Optional headers**
| header | meaning |
|---|---|
| `x-causation-id` | the event or command that caused this one — tracing |
| `x-schema` | a version of the body's shape, so a body evolves without silent misreads |
The routing key already carries the type; it is not duplicated as a header. An **unknown `x-`
header is ignored, not refused** — unlike a declaration, an event is observed by parties that need
not all understand every header, and refusing would couple every consumer to every emitter's
additions.
### Messages are persistent
Events are published persistent (delivery-mode 2). An audit trail that loses events on a broker
restart is not one, and the cost is disk the broker already spends on everything durable.
### Queues: one per consumer, durable, dead-lettered
- **A consumer's queue** is `<node>.<module>.events`, durable, bound to that module's consumed
patterns. Durable so a restart does not drop what arrived while it was down. **Manual ack** after
the handler succeeds — at-least-once.
- **Prefetch** bounds in-flight work (default 32) so one slow consumer does not pull the whole
backlog into memory.
- **A dead-letter exchange** `mesh.events.dead` receives a message rejected past a redelivery limit,
so a poison event is set aside for inspection rather than looping forever or vanishing silently.
- **The audit logger's queue** `<node>.audit-logger.events`, bound to `#`, is the same shape —
durable, persistent, dead-lettered — because completeness is its whole job.
- **RPC reply queues** are exclusive, auto-delete and server-named; **RPC serve queues**
`serve.<key>` are durable and shared, so several runtimes serving one tool key compete rather than
each answer.
### At-least-once, and consumers are idempotent
A handler may see an event twice — a redelivery after a crash between doing the work and acking.
Consumers must be idempotent, and `x-event-id` is what makes dedup possible. **Exactly-once is not
offered**: it is a promise no broker keeps honestly, and saying so is better than pretending.
## Consequences
- The event shape is a versioned, enforced contract, not conventions each module reinvents. The
sdk's `emit`/`on` and the runtime's AMQP binding implement it; a module never sees an exchange or
queue name.
- Metadata-in-headers means the body is exactly the domain payload, and a consumer that only wants
provenance never parses it.
- Adding a header or an origin root widens the contract and is reviewed as one — the discipline
[ADR 0043](0043-a-declaration-is-an-ordered-list-of-owned-resources.md) applies to the
declaration vocabulary.
- The sdk's first cut carried source/node/time in the *body*; this supersedes that — they move to
headers. That is code to align, in `mesh-sdk` (`emit`/`on`) and `mesh-tools` (the binding, queue
config, dead-letter).
## References
- [ADR 0046](0046-events-are-a-relationship.md) — events as a relationship; this is their wire shape.
- [ADR 0043](0043-a-declaration-is-an-ordered-list-of-owned-resources.md) — the precedent: a wire
contract, versioned, additions reviewed as security.
- [ADR 0001](0001-nodes-communicate-over-a-broker.md) — the broker.
- [ADR 0044](0044-what-the-sdk-holds-and-refuses.md) — `emit`/`on` are stable sdk surface; the
binding, queue config and dead-letter are the runtime's, not the sdk's.
@@ -0,0 +1,100 @@
---
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0046-events-are-a-relationship.md
---
# 48. A module's broker account is scoped by what it emits and consumes
## Context
[ADR 0046](0046-events-are-a-relationship.md) made events a relationship — `emits` and `consumes`
on the manifest. [ADR 0047](0047-the-shape-of-an-event-on-the-wire.md) gave them a wire shape — the
`mesh.events` exchange, the durable per-consumer queue, the reserved routing-key origins. Neither
said how a module *reaches* the broker: what account it holds, and what that account is allowed to
do.
As the code stands, there is no answer. The mesh can provision a **node** account (at enrolment)
and a **builder** account (scoped to the build queue), and it can *deliver* any module a sealed
own-secret at a declared path — but it has no way to provision a broker **account** for a general
module. A module that declares `own-secrets: {broker: …}` and nothing more receives thirty-two
random bytes, not a credential. So on the broker, `emits` and `consumes` are enforced by nothing: a
running module could bind any queue, consume any pattern, and publish under any origin, and the
manifest that says otherwise would be describing a boundary no code draws — the exact shape of fault
[04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) records, a scope
declared in manifests and read by nothing.
This settles it, so a module's place on the bus is a thing the broker enforces rather than a thing
the manifest merely claims.
## Decision
### A module gets a broker account when it is assigned, and its permissions are the manifest
When the mesh assigns a module to a node it provisions a broker account for that module on that node,
sealed to the node ([ADR 0039](0039-the-link-is-the-security-boundary.md)) and delivered as the
module's `own-secrets` broker — `amqps://` with the mesh's fingerprint, the shape
[ADR 0047](0047-the-shape-of-an-event-on-the-wire.md) already carries. The account's permissions are
derived from the manifest, and are exactly these:
- **What it consumes.** Read on `mesh.events`, and configure-and-read on its own queue
`<node>.<module>.events` bound to the patterns in `consumes`. It cannot bind or read another
module's queue. A module that consumes nothing gets no read on the events exchange at all.
- **What it emits.** Write to `mesh.events`, restricted to routing keys under its own origin,
`module.<name>.*`. It cannot publish as another module, and cannot publish under the reserved
`mesh.*` or `node.*` origins — those belong to the mesh and the host (ADR 0047). A module that
emits nothing gets no write.
- **Nothing else.** The events account reaches `mesh.events` and that module's own queue, and no
more. Tool serving and calling over `mesh.rpc` is a separate grant on the same principle — a
module serves the tool keys it declares and calls the ones it is bound to — and is scoped the same
way rather than folded in here.
### Consuming everything is a privilege, granted deliberately
`consumes: ["#"]` — the audit logger — is read across the whole bus: every module's events, the
mesh's, every node's. That is not a pattern like any other; it is the power to see everything, and
the account is where it becomes visible. The grant that lets one module read the entire bus is one
the mesh issues on purpose and can be audited — the answer to *who can read everything* is a row, not
a guess — rather than a breadth any manifest acquires by typing a single character. A `#` consume is
a reviewed grant, not a default one.
### The account is how the declaration is enforced
Because the account can do only what `emits` and `consumes` name, the broker itself refuses a module
that tries to consume a queue it did not declare or emit under an origin it does not own. That is what
makes an event relationship a rule and not a comment — the discipline that a stated rule says how it
is checked. A manifest that over-declares grants more than the module uses, which is visible and
reviewable; one that under-declares makes the module fail closed at the broker, which is the safe
direction to be wrong in.
## Consequences
- The control plane gains a **generic module broker-account**, derived from the manifest. The
builder stops being a special case: its access to the build queue becomes an ordinary expression of
what it consumes and serves, not a bespoke account method. One rule, and the builder is an instance
of it.
- The runtime reads its credential from a file (the broker own-secret), `amqps://` verified against
the mesh's fingerprint. The `guest` account is for raising the substrate, never for a module — a
module documented as holding its own credential and handed the broker's administrative one is worse
than one with no credential story at all.
- `emits` and `consumes` stop being advisory. They are the module's authority on the bus, so the
manifest is now a security boundary and is reviewed as one, the discipline
[ADR 0043](0043-a-declaration-is-an-ordered-list-of-owned-resources.md) applies to the declaration
vocabulary.
- *Who can read the whole bus* becomes an answerable question, because `#` is a grant and not an
accident.
## References
- [ADR 0046](0046-events-are-a-relationship.md) — events are a relationship; this scopes the account
by that relationship.
- [ADR 0047](0047-the-shape-of-an-event-on-the-wire.md) — the wire this account secures: the queue,
the origins, the `amqps` credential shape.
- [ADR 0039](0039-the-link-is-the-security-boundary.md) — the link is the security boundary; a
module's account is sealed to its node the same way a node's is.
- [ADR 0043](0043-a-declaration-is-an-ordered-list-of-owned-resources.md) — a declaration is owned
and its additions reviewed; a module's broker permissions are that discipline applied to the bus.
- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — a scope declared
in manifests and enforced by no code: the fault this decision closes for events.
@@ -0,0 +1,93 @@
---
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0005-capabilities-are-provisioned-on-declaration.md
---
# 49. A public name is provisioned, not registered by hand
## Context
The mesh names and resolves its own machines internally: the overlay generates
`<service>.<node>.<suffix>` wildcards, dnsmasq answers them (`wildcard-resolution`), and the mesh
issues a certificate for each internal name. A service reachable at a *public* domain —
`plex.example.com`, not `plex.anchor.internal` — needs three things that machinery does not give it:
- a **public DNS record** at a registrar or DNS provider, so the name resolves on the internet;
- a **publicly-trusted certificate** for it, because the mesh's own authority is trusted by nobody
outside the mesh;
- and routing from that name to the module — which the reverse proxy already does: a module
`requires` the `route` capability and the proxy provides it, routing by the host it was asked for.
The routing exists. The public DNS record does not: the mesh has no way to make a name resolve on
the public internet, so today that is a step someone does by hand at a DNS provider, outside the
mesh, remembered nowhere. A public name is therefore the one part of reaching a service that the
declaration graph cannot grant or withdraw — which means it is created once and outlives whatever it
was for, the shape of drift this project exists to remove.
## Decision
### A public name is a capability, requested like any other
A module reachable at a public host declares `requires: ["public-dns"]` and contributes the hostname
it wants — beside `requires: ["route"]`, which exposes it through the proxy. The name is then
provisioned on declaration ([ADR 0005](0005-capabilities-are-provisioned-on-declaration.md)): created
when the module is assigned, removed when it is withdrawn, reconciled like every provision.
### The interface is neutral; the providers are the registrars
`public-dns` is drawn at the consumer's coupling: the consumer wants *a public name that resolves to
me*, and does not care whether Cloudflare, Route 53 or a registrar's own API puts the record there.
So the interface is neutral and the providers are provider-scoped — `cloudflare-dns`,
`route53-dns`, `porkbun-dns` — each implementing the one `public-dns` contract, the same way a
neutral database coupling is answered by `postgres-database` and `mssql-database`. A module names
`public-dns`; it never names a registrar.
### The record points at the mesh's public ingress, not at the node
What the name resolves to is the address the reverse proxy answers on, not the consuming machine's.
A public service is reachable only *through* the proxy — the proxy holds the `route` grant and routes
by host to the module — so the public name must resolve to the proxy. `public-dns` and `route` are
the two halves of one public exposure: the name, and what the name reaches.
### The record is a fact, not a secret
A DNS record is public by definition, so the grant returns the fully-qualified name and its TTL and
nothing sealed. The only secret is the provider's own API credential, which is the provider module's
own-secret and never leaves it — the module that wanted the name never sees it.
### Events
The provider emits `module.<provider>.record.created` and `module.<provider>.record.removed`
([ADR 0046](0046-events-are-a-relationship.md)), so *which names the mesh publishes, and where* is a
question answered from the event trail and the grants, not from a folder of records edited at a
provider.
### The public certificate is the proxy's, and is named here only to pair it
A public name without a publicly-trusted certificate is reachable and not trusted — the same pairing
the internal name and the mesh-issued certificate already have. Obtaining that certificate (ACME
against the now-resolving public name) is the reverse proxy's to do, and its mechanism is its own
decision; it is named here so the pairing is not forgotten, not resolved here.
## Consequences
- A public name is created and torn down with the module, so it cannot outlive it, and the mesh can
say which public names it publishes without anyone reading a registrar's dashboard.
- Adding a registrar is adding a provider that answers `public-dns`; the modules that want names do
not change.
- Public exposure of a service is a trio of separate, declared, enforced relationships: the firewall
opens the proxy's public port ([ADR 0050](0050-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)),
`route` routes the host to the module, and `public-dns` makes the host resolve.
## References
- [ADR 0005](0005-capabilities-are-provisioned-on-declaration.md) — a capability is provisioned on
declaration; a public name is one.
- [ADR 0050](0050-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the firewall, the other
half of the reachability question this was asked with.
- [ADR 0046](0046-events-are-a-relationship.md) — the provider's record events.
- [ADR 0045](0045-what-a-module-is.md) — a provider and its interface; the neutral-interface,
scoped-provider naming this follows.
@@ -0,0 +1,92 @@
---
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0037-the-host-applies-it-does-not-decide.md
---
# 50. A machine's firewall is the sum of what its modules listen on
## Context
The reverse proxy is a *provider*: a module `requires` the `route` capability and a running proxy
provides it, routing traffic by name and reaching back to the consumer. A fair question follows —
is the firewall the same shape? Should a module *register* a port with a firewall provider the way
it requests a route?
It should not, and the difference is the point. A reverse proxy is a service another component
performs; a firewall is a property of the machine — a packet filter the host applies to itself.
Modelling it as a provider would invent a credential and a reach-back for something that has neither.
And the mesh already has the registration: a module declares `listens: [{ port, from }]` — the port
it accepts connections on, and from where. That *is* how a service says it wants a port open. What is
missing is not a model but enforcement. [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md)
records that a `scope:` key five manifests carry is read by no code: a manifest can appear to
restrict a port and restrict nothing — the exact fault
[how-we-build.md](../00-META/how-we-build.md) names, *an unenforced rule is indistinguishable from a
wrong one*, made worse because the declaration reads as a restriction.
## Decision
### The firewall is derived and host-applied, not a provider
A machine's firewall is the sum of what the modules assigned to it declare they listen on, computed
by the host and applied as one of its owned resources ([ADR 0037](0037-the-host-applies-it-does-not-decide.md):
the host applies, it does not decide; [ADR 0043](0043-a-declaration-is-an-ordered-list-of-owned-resources.md):
the declaration is owned resources). It is not a capability, not a per-consumer grant — opening a
port is a declarative fact about a machine, so it is computed and applied, not requested and
credentialed.
### `from` is the whole of public-versus-internal
The distinction the question is really about lives in `from`:
- `listens: [{ port: 5432, from: mesh }]` — open to the private overlay only.
- `listens: [{ port: 443, from: anywhere }]` — open to the public internet.
A module registers a port on the firewall by listening on it and saying from where. There is no
separate firewall capability, because the firewall is not a thing that reaches back or holds a
secret; it is the machine's own filter over the ports its modules named.
### The host enforces it both ways, and unknown keys are refused
A port a module listens on is opened to exactly the scope it named; a port nothing declares is
closed. And a key the firewall does not read — the `scope:` of issue 003 — is refused at the
manifest, not accepted and ignored, so a declaration that reads as a restriction is one. This is the
discipline [ADR 0048](0048-a-module-broker-account-is-scoped-by-emits-and-consumes.md) applied to the
broker account, applied here to the packet filter: the declaration is the enforcement, or it is a
comment.
### A public service is exposed through the proxy, not by opening its own port
Reaching the public internet is normally not `from: anywhere` on the service's own port. The service
listens `from: mesh` — only the proxy reaches it — and `requires: route`, so the sole machine with a
public opening is the one running the reverse proxy, and the service is exposed by name through it.
`from: anywhere` is the deliberate direct-exposure case, for a service that is its own front door.
## Consequences
- Issue 003 is closed: the firewall is computed from `listens` and enforced, so a declared scope is
real and an undeclared port is shut. Rejecting unknown manifest keys is the general fix, of which
the `scope:` key was one instance.
- The firewall and the reverse proxy stop being confused for one model: the firewall is the machine's
filter (host-derived from `listens.from`); `route` is a name-router (a provider); the public DNS
name is a third thing ([ADR 0049](0049-a-public-name-is-provisioned-like-any-capability.md)). A
public service uses all three.
- The modelling question is answered: a module registers a port by declaring `listens`, and reaches
the public internet by name through `route` + `public-dns` — never by the firewall being a
provider.
## References
- [ADR 0037](0037-the-host-applies-it-does-not-decide.md) — the host applies; the firewall is one of
the things it applies.
- [ADR 0043](0043-a-declaration-is-an-ordered-list-of-owned-resources.md) — the firewall is a derived
owned resource, not a grant.
- [ADR 0048](0048-a-module-broker-account-is-scoped-by-emits-and-consumes.md) — the same discipline:
a declaration is enforced, or it is a comment.
- [ADR 0049](0049-a-public-name-is-provisioned-like-any-capability.md) — the public name, the other
half of the reachability question this was asked with.
- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — the unenforced
`scope:` this closes.
@@ -0,0 +1,87 @@
---
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0005-capabilities-are-provisioned-on-declaration.md
---
# 51. A module's configuration is its assignment's, not its manifest's
## Context
A module is assigned to a node — `assign <node> <module>`, always to a machine; there is no
assignment to the mesh. "Mesh" is a *scope*, not a place: a `provides` or a `claim` scoped `mesh`
reaches the whole mesh, but the module still runs on a node. So the two kinds of thing a module can
carry are the manifest (what the module *is*) and, separately, what it should do *here* — which
differs by deployment and by node.
The mesh already has the second: **settings**. `settings set <module> [--node <node>]` — with a node
it is that machine's, without it the whole mesh's — layered over what the module declares and applied
at resolution, changeable without editing the module and without a rebuild. That is the surface a
meshboard would edit.
But settings today reach only a module's **config-file content** (a mergeable file the module owns).
Configuration that is not a file has been landing in the manifest instead, statically — a registrar's
zone and domain, the address public names point at, and, most sharply, `listens.from`. That last one
is the tell: whether a port is open to the private overlay or to the public internet is a
*per-node deployment choice* — the same database internal on one machine and public on another — and
a value fixed in the manifest is one value for every machine, so it cannot be. Static configuration in
the manifest is configuration in the wrong place: it cannot vary per node, and it cannot change
without a new module version.
## Decision
### The manifest is identity and defaults; the assignment's settings are the configuration
A module's manifest declares what it is — what it provides, requires and claims, the shape of its
resources — and, for anything configurable, a **default**. The values that make a running instance
*this* instance are settings, carried by the assignment: per-node, or mesh-wide when no node is named,
applied over the defaults at resolution. Change one and the next reconcile carries it; nothing is
edited on a machine and nothing is rebuilt.
### Settings drive the configurable fields the manifest marks, not only file content
Settings extend beyond a config file's content to the manifest fields a module declares settable —
foremost:
- **`listens.from`**: a module declares its safe default (`from: mesh`), and a per-node setting
raises or lowers it. postgres declares `listens: [{ port: 5432, from: mesh }]`; on the machine that
should expose it, a setting makes that port `from: anywhere`. Same module, different exposure, and
the firewall ([ADR 0050](0050-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)) is computed
from the effective value, so the packet filter follows the setting.
- **A provider's own configuration**: a registrar's zone, domain and the ingress its names point at
([ADR 0049](0049-a-public-name-is-provisioned-like-any-capability.md)) are mesh-wide settings, not
manifest constants — one mesh's Cloudflare zone is not another's, and the module description is the
same for both.
### Unset is the default, and an unknown setting is refused
A field with no setting keeps the manifest's default, so a module runs correctly configured by nobody.
A setting that matches no settable field — like a config value that reaches no file today — is named,
not silently dropped, so a misspelled setting is found rather than believed (the discipline of
`UnusedSettings`, and of [ADR 0048](0048-a-module-broker-account-is-scoped-by-emits-and-consumes.md):
a declaration is enforced or it is a comment).
## Consequences
- The postgres case works: one module, `from: mesh` by default, `from: anywhere` where a setting says
so — internal on ace, public on novox, changeable live.
- Provider modules stop carrying a mesh's specifics: `cloudflare-dns` describes *a Cloudflare
registrar*, and *which* zone and ingress is a setting, so the same module serves every mesh.
- Configuration becomes a thing a meshboard manages — set per node or mesh-wide, applied on the next
reconcile — rather than a manifest edit and a rebuild ([ADR 0004](0004-managed-files-are-generated-never-edited.md):
the way you change a managed thing is not by editing it).
- What a manifest may not do is grow a value that differs per machine; if it differs per machine it is
a setting, and the manifest holds only the default.
## References
- [ADR 0005](0005-capabilities-are-provisioned-on-declaration.md) — what is provisioned on
declaration; its per-instance values are the assignment's.
- [ADR 0004](0004-managed-files-are-generated-never-edited.md) — a managed thing is changed through
the mesh, not by editing it; settings are that, for configuration.
- [ADR 0050](0050-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the firewall follows the
effective `listens.from`, so making `from` a setting makes exposure a setting.
- [ADR 0049](0049-a-public-name-is-provisioned-like-any-capability.md) — the provider whose zone and
ingress are settings, not manifest constants.
@@ -0,0 +1,86 @@
---
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0048-a-module-broker-account-is-scoped-by-emits-and-consumes.md
---
# 52. A module runs its code as its own process, with its own account
## Context
A module is one self-contained thing ([ADR 0045](0045-what-a-module-is.md)), and it gets a broker
account scoped to what it emits and consumes ([ADR 0048](0048-a-module-broker-account-is-scoped-by-emits-and-consumes.md)).
The catalogue now gives modules **tools** and **events** — real code, in the module ([ADR 0044](0044-what-the-sdk-holds-and-refuses.md)) —
but nothing has said what *runs* that code. The audit-logger showed one shape and was treated as an
exception: a container running the tool runtime carrying the module's compiled code, holding the
module's own scoped account. Every module with tools or events needs the same, and the tempting
alternative does not work.
**A node-wide runtime that loaded every assigned module's code cannot hold a per-module account.** It
would run under one account with the union of every module's permissions — able to emit as any of
them and read any of their queues — which is exactly the isolation ADR 0048 exists to draw. So the
runtime is per-module, not per-node, and treating the audit-logger as special left the other
modules' code with nothing to run it: the conversion produced tools and events that, as it stands,
never execute.
## Decision
### A module with tools or events runs a process of its own
A module that has tools or events runs a **runtime process** — a container, the tool runtime carrying
that module's compiled code — assigned and started like the module it is, holding the single broker
account the mesh scoped to it (ADR 0048). One module, one process, one account.
### It serves its tools, each on its own key
A tool is served on its own key (`serve.<tool>`), and a caller invokes a named tool. Only the module
that serves it answers, and the module's account is scoped to exactly its tool keys — so one module
cannot answer another's calls, the isolation ADR 0048 gives events extended to tools. This supersedes
a single `tools.invoke` endpoint that dispatched by name: that shape assumed one runtime for the
whole node, and per-module runtimes competing on one key would each be handed calls for tools they do
not have.
### It runs its events in the same process, under the same account
Emitting under the module's own origin and consuming its own queue ([ADR 0047](0047-the-shape-of-an-event-on-the-wire.md))
happen in that same process, with that same account — not a second one to scope and seal. A module's
tool code, its event code and, for a provider, its provisioner are the one module's code and run as
the one module's process.
### The runtime image is the tool runtime plus the module's code
Built from the module's source like any module image — the audit-logger's shape, made the rule, not
the exception. The module declares a `container` for it carrying `MESH_BROKER_FILE` (its sealed
credential, ADR 0048) and its compiled code. A module with **neither** tools nor events runs no such
process: a plain service module — the plex *server*, dnsmasq the resolver — is its service and files
and nothing more. A module that is both a service and code declares both containers: the service, and
the runtime beside it.
## Consequences
- The catalogue's tools and events become runnable: each tools-or-events module gains a runtime
container with its scoped credential, and the audit-logger stops being special. Until this, the
converted modules held code with nothing to execute it.
- A process, and a small image, per tools-or-events module. That is the cost of ADR 0048's isolation:
one account per module means one process per module. It is paid deliberately — a shared runtime is
cheaper and cannot be scoped, and a mesh where any module can emit as any other is not one worth the
saving.
- `serve.<tool>` per key replaces the single `tools.invoke` dispatch. The sdk's serving and a module's
account scope both come to name tools individually.
- **A provider's provisioner is a runtime process too.** It already runs as its own container; its
events (`bucket.created`, `database.provisioned`) belong to *that* process and need the same
credential. So a provisioner that emits carries `MESH_BROKER_FILE` and its scoped account like any
runtime — or it does not emit. (This is the fix for provisioners that emit today with no broker
bound: the emit is a runtime's, and the provisioner is a runtime.)
## References
- [ADR 0045](0045-what-a-module-is.md) — a module is one self-contained thing; its code runs as one
process.
- [ADR 0048](0048-a-module-broker-account-is-scoped-by-emits-and-consumes.md) — the scoped account
this process holds, and the isolation that makes it per-module.
- [ADR 0047](0047-the-shape-of-an-event-on-the-wire.md) — the events this process runs, and the
`serve.<key>` queue tools now use.
- [ADR 0044](0044-what-the-sdk-holds-and-refuses.md) — the code lives in the module; this runs it.
@@ -0,0 +1,130 @@
---
status: accepted
date: 2026-09-05
deciders: jochen
reconstructed: false
---
# 53. A provider creates the credential the mesh minted, and seals nothing
## Context
A provider module stands up a per-consumer resource — a database, a cache bucket, an object
store user — and the consumer must end up holding a credential that authenticates against it.
Building the module runtime (ADR 0052: a module runs its own code as its own process under its
own account), the provider's provisioner was run for the first time as a delivered thing, and
it did not work. It reads a seal key from the environment that nothing sets, and it seals every
credential it produces to that key with a symmetric passphrase.
Tracing the credential's path turned up something larger than a missing key. **The provisioner
harness the whole catalogue is built on describes a credential flow the mesh does not have, and
duplicates — incorrectly — one it does.**
What the sdk's `runProvisioner` does today:
- reads request files named `*.grant.json` — which nothing in the mesh writes;
- calls an adapter whose `create` **generates its own password** and returns it;
- seals that password with a symmetric key (`$MESH_SEAL_KEY`) and writes a `*.credential`
file — which nothing in the mesh reads, and no consumer ever unseals.
What the mesh already does, and has wired end to end:
- The control plane mints one password per (consumer, provider) pair (`Inventory.SecretFor` →
`secrets.Make`) and seals it to **both** node keys asymmetrically — a copy the consumer's
host can open and a copy the provider's host can open. No shared symmetric key exists
anywhere, on purpose: a key both ends hold is a key the mesh would have to distribute, which
is the same problem one level down, and the control plane's own code refuses it.
- The provider is handed, at the path its `receives` names, one contribution per consumer:
the **login to create** (`As`, derived by the mesh so the two ends agree by construction),
the consumer's address and requested values, and a **`Secret` file** holding that consumer's
password sealed to the provider and unsealed onto the machine by its host.
- The consumer is handed the *same* password, as plaintext its own host wrote by unsealing its
copy and substituting it into a config file. The consumer never unseals anything itself and
holds no key.
So the password a provider's provisioner invents is not even the password the consumer was
given: a consumer authenticating with the mesh's password against a resource the provisioner
created with its own would simply fail. The symmetric seal is not an incomplete feature to
finish delivering a key for. It is a second, contradictory credential model bolted beside the
real one, and it cannot be made to work without building the very thing the mesh was designed
not to have.
This is a decision and not a patch because the harness is the **provider contract**. Every
provider — the four that exist and the many a real mesh grows — is built on
`runProvisioner(resource, adapter)`. Whatever it says a provider is, they all inherit; and
changing it later is one migration per provider. It is cheaper and more honest to settle what
a provider is now.
## Decision
**A provider is handed the credential; it does not make one, does not seal one, and does not
hand one back.** The provisioner's only job is to make the mesh's grants true in its own
software.
Concretely, for the sdk harness and the adapter contract:
- The harness reconciles the **contributions the mesh delivers** to the provider's `receives`
path — the list of consumers, each with its login name (`As`), address, requested values,
and the path to its unsealed password (`Secret`). It does not read `*.grant.json` and it
does not write `*.credential`.
- For each consumer present, the harness reads the password from that consumer's `Secret` file
and calls the adapter to bring the resource into being under the given login. For each
consumer no longer present — the mesh drops it from the contributions file when its consumer
goes away — the harness calls the adapter to withdraw it.
- The adapter shrinks to the per-software half and nothing else. It is given the login, the
password, and the values, and it makes the resource exist or removes it. It generates no
password, derives no name, seals nothing, and returns no credential:
roughly `create({ as, password, values })` and `remove({ as })`, both returning nothing.
- `$MESH_SEAL_KEY`, the symmetric `seal()`/`writeSealedCredential` path, and the `*.grant.json`
/ `*.credential` files are removed from the provisioning path entirely. The credential
reaches the consumer through the mesh's own asymmetric channel, which already crosses node
boundaries and holds no shared secret.
Identity stays the mesh's to say. The login the provider creates is the name the mesh derived
and gave the consumer to present; the provider never invents a name, because a name the
consumer cannot learn is a name it cannot authenticate with.
## Consequences
- A provider module becomes smaller and unable to be wrong in this way: with no password to
generate and no key to seal to, the class of bug where the two ends hold different secrets
cannot be written. A provider added after this inherits the corrected contract and has no
seal to reintroduce.
- The four current providers (redis, postgres, minio, umami) each lose their `generatePassword`
+ seal code and gain a `create` that takes the password it is given. Their teardown becomes
"withdraw the login named `As`".
- The symmetric `seal()`/`unseal()` primitive loses its only caller and leaves — checked, not
assumed: nothing else in the sdk or the catalogue called it, so it is removed with the
provisioner it belonged to.
- **How this is verified:** redis is assigned as a provider in the lab, the contributions and
the unsealed password the mesh would deliver are put in its `receives` path, and a client
authenticates as that consumer with the mesh's password and gets PONG — where a provider that
invented its own password answers WRONGPASS — with `$MESH_SEAL_KEY` set nowhere and no
`.credential` file written. Proven: `provider-uses-mesh-credential` is green.
**What this does not cover — credential provisions, not data provisions.** This decision is about a
provision whose credential is a *secret the mesh mints* — a login and password (redis, postgres,
minio). A provider that instead *generates* the thing the consumer needs, and that thing is not a
secret — umami's `analytics`, where the consumer wants back a `siteId` umami assigned — does not fit,
because a contract that returns nothing has no way to hand that data back. The seal-key fault was
never umami's (it sealed no password; it returned a public id), so removing the seal does not break
it further, and it still reconciles its sites off the mesh's contributions. But delivering
provider-generated data back to a consumer is a *return path* the mesh does not have and this
decision does not build — a separate shape, left to a separate decision.
- Teardown beyond "remove the login" — data an object store leaves behind when a consumer
leaves — is named by each provider's adapter, not by the harness, and is out of scope here
except to say the contract must leave room for it.
## References
- [04-ISSUES/008](../04-ISSUES/008-provider-runtime-has-no-seal-key/00-report.md) — the
observation and the cross-repo trace this decision rests on.
- ADR 0052 (the module runtime) — what first ran a provider's provisioner as a delivered
process and exposed this; link to be filled when 0052 lands on the trunk.
- Control-plane mechanisms this relies on already existing: `mesh-control` —
`internal/inventory/secrets.go` (`SecretFor`, `SecretsFrom`), `internal/secrets/seal.go`
(`Make`, the two-blob asymmetric sealing), `cmd/mesh-control/plan.go` (`grantsFor`, the
`Grant.Sealed = ForProvider` delivery), `internal/catalogue/declaration.go` (the `receives`
contribution: `As`, `At`, `Values`, `Secret`).
- The path being removed: `mesh-sdk` — `src/provisioner/index.ts` (`runProvisioner`, `sealKey`,
`writeSealedCredential`) and the symmetric `src/primitives/index.ts` `seal()`/`unseal()`.
@@ -0,0 +1,129 @@
---
status: accepted
date: 2026-09-05
deciders: jochen
reconstructed: false
---
# 54. A consumer's identity is bounded by the tightest backend that must accept it
## Context
The mesh says who a consumer is, once, and hands the same name to the provider (to create) and the
consumer (to present), so the two ends agree by construction rather than by two conventions (the
principle behind `ConsumerIdentity`, 04-ISSUES/023). The name is `mesh_<node>_<module>`, cleaned to
lower-case letters, digits and underscore.
Proving the provider contract per backend (ADR 0053) turned up 04-ISSUES/010: redis and postgres
create that name verbatim, but **minio refuses it** — an S3 access key is capped at 20 characters,
and `mesh_anchor_bucketuser` is 22. The provisioner then retries for ever, per consumer, and the
consumer holding that same too-long name could never present it either.
Two things about the existing derivation decide most of this:
- **The charset is already right.** `[^a-z0-9_]` is deliberately conservative, and its own comment
says it reaches "a PostgreSQL role, a MinIO access key, an LDAP uid and a Keycloak client without
quoting." That much is true.
- **The length is wrong.** `CheckIdentity` refuses names over `identityLimit = 63`, commented as
"the shortest identifier limit among the systems these names reach: PostgreSQL's". It is not the
shortest — S3's 20 is shorter — so the guard that was meant to catch exactly this lets it through,
and the failure lands at provision time as a silent retry instead of at assignment as a refusal.
So this is a small wrong constant with a real cost attached: whatever bound we set, `mesh_` (5) plus
a node name plus `_` plus a module name has to fit inside it.
## The options
**A — Bound the identity by the true minimum, and refuse early.** Lower `identityLimit` to the real
shortest (20, S3's), so `CheckIdentity` refuses an over-long name *at assignment* with a clear
message, the way it already refuses over-63 names. The derivation does not change; long names are
simply rejected before anything is provisioned.
- *For:* smallest change; keeps "the mesh says the identity once, verbatim" intact; the failure
moves from a per-consumer provision-time retry to an up-front, legible refusal — which is what
`CheckIdentity` exists to do.
- *Against:* a hard budget. `mesh_` + node + `_` + module ≤ 20 means node + module ≤ 14 characters.
`anchor` + `bucketuser` (16) is already over. It pushes the constraint onto how machines and
modules are named, which is a real limitation on legible names.
**B — Keep the readable name when it fits, compact it when it does not.** Below the bound, the name
is `mesh_<node>_<module>` as today; over it, the mesh substitutes a deterministic short form (e.g.
`mesh_` + a truncated hash of node+module) — still one derivation, so both ends still agree.
- *For:* no naming constraint; short backends always satisfied; the common case stays legible.
- *Against:* some identities become opaque, and a provisioner tracing "whose login is this" loses
the answer for exactly the consumers that overflowed. The mesh now owns a fallback format and its
collision properties (a truncated hash is not free of collisions at 15 characters).
**C — Let each interface declare its identifier bounds, and derive within the tightest a consumer
reaches.** `s3-bucket` states `identifier: { max: 20 }`; `postgres-database` states 63; the mesh
derives a name that fits the **minimum** bound across the providers a given consumer is granted.
- *For:* the most precise — each provision gets exactly the room it has, and a database consumer
keeps long legible names while an S3 consumer gets a short one; the constraint lives where the
fact does (on the interface).
- *Against:* the most work, and a consumer of two interfaces with different bounds must satisfy the
smaller — so its name shortens for both, reintroducing B's opacity in a narrower case. It also
means one consumer can hold **different** identities per provision, which the "said once" model
currently forbids.
**D — Let the provider generate a backend-valid identity and hand it back (rejected).** minio mints
its own access key and returns it to the consumer. This is the data-provision return path this era
keeps meeting — but it directly contradicts 023 and ADR 0053: the identity would no longer be the
mesh's single derivation the two ends share, it would be a value one side invents and the other must
be told. Listed for completeness; not recommended.
**E — A module (and a node) may declare a short slug; the identity is built from it.** The identity
becomes `mesh_<node-slug|node-name>_<module-slug|module-name>`: where a slug is declared it is used,
otherwise the cleaned name. A slug is a deliberately short, operator-chosen identifier — `kc` for
keycloak, `wkstn` for a workstation. It is optional: short names (`anchor`, `redis`) need none.
- *For:* this is the escape hatch B wanted to be, without the opacity. The name stays legible — a
provisioner can read `mesh_wkstn_kc` and know who is asking — because a person chose it, not a
hash function. And it makes an early refusal *palatable*: if even the slug-built identity overflows,
the refusal points at the slug, a field made for exactly this, rather than at the machine's name.
Both ends still derive it from one declared thing, so they agree by construction.
- *Against:* a new optional manifest field, and someone must pick the slug — but only for names that
would otherwise overflow, and picking a short legible identifier is a better job than being handed
a hash.
## What implementing A revealed
A was tried first. At `identityLimit = 20`, the readable budget is `mesh_` (5) + node + `_` + module
≤ 20, i.e. **node + module ≤ 14 characters** — far tighter than it looked. The catalogue's own
existing tests use `workstation`+`keycloak` (25), which compacts to `mesh_dbbc02f8dde34d3`; common
mesh names (`home-server`, `the-build-node`, `workstation`) blow the budget with any module. So B's
compact fallback would fire for the *common* case, not the rare overflow — which inverts A+B: most
identities would be opaque hashes. A alone (hard refusal at 20) would refuse most realistic names.
This is what moved the recommendation to E: the problem is not the limit, it is that the *readable
name* is the wrong source when it is long, and a slug is a better source than either a hash or a ban.
## Recommendation
**E, over a per-consumer bound (start with the global minimum, 20).** Build the identity from an
optional slug, keep it when it fits, and refuse at assignment with "declare or shorten `<module>`'s
slug" when it does not — no hash, no lost legibility, and the fix is a first-class field. Set the
bound to the true minimum (20) now; it needs no per-interface machinery to unblock S3, and a module
that consumes S3 simply declares a short slug. Graduate to **C** (per-interface bounds) later if it
turns out that non-S3 consumers are paying for S3's limit often enough to mind — E and C compose:
slugs are the mechanism, per-interface bounds refine where the ceiling sits. **B is dropped**: a
declared slug is a strictly better escape hatch than an opaque hash. **D stays rejected.**
## Consequences (of E)
- A module manifest gains an optional `slug`; a node may carry one too. `ConsumerIdentity` prefers
the slug over the cleaned name for each half. `identityLimit` becomes 20 (the true minimum), and
`CheckIdentity` refuses at `module add` / assignment — now with a message naming the slug to set.
- The common case stays legible; only names that overflow the budget need a slug, and what they get
is a name a person chose, not a hash.
- Existing modules/nodes whose names overflow declare a slug once — a migration cost paid as a clear
refusal with an obvious remedy, not a silent hash or a silent truncation.
- minio (04-ISSUES/010) is unblocked: an S3 consumer declares a short slug and its access key fits.
- **How it is checked:** the minio grant e2e — a consumer whose (slugged) identity fits reaches its
bucket with the credential the mesh delivered — plus unit tests that a slug is preferred, that an
un-sluggable over-long identity is refused (naming the slug), and that two consumers never collide.
## References
- [04-ISSUES/010](../04-ISSUES/010-mesh-login-exceeds-s3-access-key-limit/00-report.md) — the
observation.
- ADR 0053 — a provider creates the credential the mesh minted; the identity it creates it under is
the one this decision bounds.
- `mesh-control` `internal/catalogue/identity.go` — `ConsumerIdentity`, `identityUnusable`,
`identityLimit`, `CheckIdentity` — where the constant and the check live.
@@ -1,9 +1,9 @@
--- ---
status: open status: fixed
opened: 2026-08-22 opened: 2026-08-22
located-in: [] located-in: [mesh-control/internal/catalogue, mesh-catalog/modules/firewall]
fixed-by: fixed-by: the manifest refuses unknown keys, `from` is the field that scopes a port and it is rendered to nftables, and the firewall module applies it
amended-design: amended-design: 0050-a-machine-firewall-is-the-sum-of-what-it-listens-on.md
--- ---
# 003 — A firewall rule's `scope:` is read by no code # 003 — A firewall rule's `scope:` is read by no code
@@ -31,10 +31,27 @@ any check.
- Five manifests carry the key. Zero code paths consume it. - Five manifests carry the key. Zero code paths consume it.
- Recorded as an observation on 2026-08-22. - Recorded as an observation on 2026-08-22.
## Open questions ## Resolution
- Should the manifest reject unknown keys outright? That is the general fix; this is one Both open questions are answered, and the chain from a declared scope to a packet actually dropped
instance of it. is closed — recorded as [ADR 0050](../../02-DECISIONS/0050-a-machine-firewall-is-the-sum-of-what-it-listens-on.md).
- Were the five declarations intended to restrict something that is currently open? Each needs
checking against what the node actually exposes — the declaration cannot be trusted either - **Unknown keys are refused, not accepted.** `ParseManifest` decodes with
way. `DisallowUnknownFields`, so a `scope:` key the firewall type does not have is now rejected at the
manifest — the general fix, of which this was one instance. A key that reads as a restriction can
no longer be one nothing enforces.
- **The field that scopes a port is `from`, and it is read.** A `listens` entry names its source —
`mesh`, `anywhere` or `machine` — and the control plane renders the union of every module's
`listens` into a node's whole nftables rule set (`AsNftables`), default-drop with an accept scoped
to exactly the source each port named. The five `scope:` declarations were the wrong spelling of
that intent; `from` is the right one, and it is enforced.
- **A module applies it.** The rendered rule set is written to the node (the `filtering` resource),
and the `firewall` module (mesh-catalog) loads it — the last link, without which the rules were
computed and never dropped a packet.
## Original open questions
- Should the manifest reject unknown keys outright? — **Yes; it does now** (`DisallowUnknownFields`).
- Were the five declarations intended to restrict something that is currently open? — They meant to
scope a port and used a key nothing read; expressed through `from`, that intent is now enforced.
Any manifest still carrying `scope:` is refused at parse, so it is found rather than believed.
@@ -0,0 +1,135 @@
---
status: resolved
opened: 2026-09-04
located-in: [mesh-sdk, mesh-catalog]
fixed-by: mesh-sdk src/provisioner rework + redis/postgres/minio/umami adapters (ADR 0053)
amended-design: 0053-a-provider-creates-the-credential-the-mesh-minted.md
---
# A provider's provisioner seals with a key the mesh has no way to deliver — and does not need to
## What was observed
Building the vertical slice for the module runtime (the module runs its own code as its own
process under its own account), a **provider** module — one that stands up a per-consumer
resource and hands back a credential — was assigned to a node and run as a broker-bound
runtime. The runtime hosts the module's provisioner (the sdk's `runProvisioner`), and the
harness opens by reading a **seal key** from `$MESH_SEAL_KEY`, failing immediately without
one. Every credential it produces for a consumer is sealed to that key with the sdk's
symmetric `seal()` (AES-256-GCM, `mesh-sdk/src/primitives/index.ts`) before being written.
Nothing in the mesh sets `$MESH_SEAL_KEY`. It is read in exactly two places in the sdk and
set nowhere — no manifest, no control-plane code, no host code. So a provider runtime, as
delivered, aborts at start-up. The slice proved the mechanism only by setting a lab-local key
in the manifest by hand.
## What a trace of the credential path turned up
The seal key is not a missing delivery. **The whole symmetric-seal provisioner is orphaned,
and it duplicates — badly — a job the mesh already does.**
- `runProvisioner` reads request files named `*.grant.json`. **Nothing writes those.**
- It writes sealed credential files named `<consumer>.<resource>.credential`. **Nothing reads
those** — not the host, not the control plane. The host reports applied-resource digests
upward and never ships credentials; the control plane has no reference to that filename.
- No consumer ever calls the symmetric `unseal()`. Consumers receive **plaintext**.
Meanwhile the mesh already carries a provider→consumer credential across nodes, with **no
shared key anywhere**:
- The control plane mints the password once (`secrets.Make`) and seals it **twice,
asymmetrically** — `ForConsumer` to the consumer node's X25519 public key, `ForProvider` to
the provider node's (`mesh-control/internal/secrets/seal.go`, `mesh-host/internal/identity/
sealing.go`, NaCl box).
- Each host opens its own copy with its own private key on the machine; the plaintext exists
only for the length of one function call (`mesh-host/internal/apply/apply.go`, the
`${secret:name}` substitution — ADR 0024's "the host is the only thing that ever holds
both").
- `serves` carries no credential and says so; `receives`/`bound` tell each side *where* its
sealed secret is, never the value.
The two models also **contradict** each other. The sdk's `seal()` comment says the key is "a
per-node passphrase the host holds"; the host holds no such passphrase — it holds an X25519
private key, and the control plane's own code refuses a shared symmetric key on principle:
"a key both ends hold is a key the mesh would have to distribute, which is this problem again
one level down" (`secrets/seal.go`). A symmetric `MESH_SEAL_KEY` shared between a provider
node and a consumer node is exactly the thing the mesh was built not to have.
And the provisioner's model is wrong in a second way: its adapter **generates its own
password** (`generatePassword()`) and creates the resource with it — a different password from
the one the mesh mints and hands the consumer. Even with a seal key delivered, a consumer
would authenticate with the mesh's password against a resource created with the provisioner's.
## Why it matters beyond this instance
This is not a four-module problem. The provider contract lives in **one place** — the sdk's
`runProvisioner(resource, adapter)` harness — and every provider is built on it. Four exist
today (redis, postgres, minio, umami); a mesh of any size ends up with many. Whatever the
provisioner harness does, every present and future provider inherits, so the orphaned
symmetric seal is a fault stamped into the interface, not into four adapters. That also sets
the cost of getting it wrong: a contract N providers depend on is N migrations to change
later, which is the argument for settling it deliberately now rather than patching around it.
As written, each provider carries a provisioner that cannot start (no key), and that, if it
did, would create resources with a password it invented — a *different* password from the one
the mesh minted and handed the consumer — and seal them for a reader that does not exist. The
rule the design states, "a consumer receives a sealed credential and unseals it," is enforced
by nothing: no consumer unseals, and no shared key exists to unseal with.
## The mesh already does this — confirmed
The premise the fix rests on is not a hope; it is in the control plane today. For a served
interface, `Inventory.SecretFor` mints one password per (consumer, provider) pair via
`secrets.Make`, sealing it to **both** node keys — `ForConsumer` and `ForProvider`.
`SecretsFrom(provider)` is documented as "every credential a provider node was issued, so it
can be told what to create," and `grantsFor` (plan.go) hands the provider node one `Grant` per
consumer carrying `Sealed: ForProvider`. The provider receives, at the path its `receives`
names, one `Contribution` per consumer: the login to create (`As`, derived by the mesh so both
ends agree — 04-ISSUES/023), the consumer's address (`At`) and requested `Values`, and
`Secret`, the file holding that consumer's password sealed to this provider and unsealed by
its host. Everything the provisioner needs is delivered. It reads the wrong files
(`*.grant.json`, which nothing writes) and invents a password instead of reading the one in
`Secret`.
## The fix this points to
A **one-place contract change in the sdk harness**, plus re-pointing today's adapters at it —
not per-provider surgery, and inherited correctly by every provider after them:
- `runProvisioner` reconciles the mesh-delivered `receives` contributions (not `*.grant.json`):
for each consumer, create the resource under the login `As` with the password read from the
delivered `Secret` file, for its `Values`; withdraw the login when a consumer leaves the file.
- The adapter stops generating a password and stops returning a credential — it is handed the
name and the password and only makes the resource exist. Roughly `create({as, password,
values})` / `remove({as})`, no return.
- `sealKey`, `seal()`, `writeSealedCredential`, `MESH_SEAL_KEY`, and the `.credential` file
leave entirely; the consumer already receives its copy through the mesh's own channel.
This is proposed as ADR 0053, which defines the corrected provider contract, for ratification.
## Resolution
ADR 0053 was accepted and implemented on the branches this issue is fixed by:
- `mesh-sdk` `src/provisioner/index.ts` now reconciles the mesh's `receives` contributions and,
per consumer, reads the mesh-minted password from the file the host unsealed, calling the
adapter to create the resource under the mesh's login. `$MESH_SEAL_KEY`, the symmetric seal,
`writeSealedCredential`, and the `*.grant.json` / `*.credential` files are gone. The symmetric
`seal()`/`unseal()` primitive had no other caller and was removed.
- The four adapters (redis, postgres, minio, umami) were re-pointed at the new contract —
`create({ as, password, values })` / `remove({ as })`, returning nothing. minio's client gained
a secret-key argument so it sets the mesh's secret rather than generating one.
- Proven in the mesh-lab: `provider-uses-mesh-credential` is green — redis creates the consumer's
login with the password the mesh minted, a client authenticates as that consumer and gets PONG,
with no seal key set anywhere.
Two things were carved out deliberately, neither blocking:
- **Data provisions are a separate shape.** umami's `analytics` returns a `siteId` umami
*generates*, not a secret the mesh mints, and a contract that returns nothing cannot hand that
back. ADR 0053 is scoped to credential provisions and says so; the provider→consumer return
path for generated data is left to a separate decision. umami compiles and reconciles under the
new harness; only that return is unaddressed, and it never had the seal-key fault.
- **Teardown beyond "remove the login"** — an object store's leftover data — is each adapter's to
name (minio leaves a non-empty bucket for an operator rather than deleting a consumer's data),
not the harness's.
@@ -0,0 +1,67 @@
---
status: open
opened: 2026-09-04
located-in: []
fixed-by:
amended-design:
---
# Changing a module's settings does not restart its runtime — config is stale until recreated
## What was observed
Rolling the module runtime out to the catalogue (the runtime that serves a module's tools and
runs its events under the module's own account), each tools+events module receives its
configuration the way the design intends: a mergeable config file the module declares, into
which the assignment's settings are merged. The runtime container mounts that file and reads
it once at start-up, when it builds its API client.
The design for settings says a config file a module owns can be changed **without editing
it** — a person states an intention, the file is regenerated, and the change takes effect.
The decision that config is the assignment's, not the manifest's, is explicitly so that
configuration can be updated *on the fly* and managed from a dashboard.
For a runtime delivered as a **container**, that last part does not hold. When settings
change, the control plane re-renders the config file on the node — but the runtime container
is only ever recreated when its **spec** changes, and the spec is image, name, env, ports,
volumes and args. The *content* of a mounted file is not part of it. So the file on disk
updates and the process that already read it keeps the value it read at start-up. The new
configuration does not take effect until something changes the container's spec, or it is
recreated by hand.
A **service** resource has `restart-on`, which names the resources whose change forces a
restart — exactly this problem, already solved, for units. A **container** resource has no
equivalent field, and the apply path for containers never consults the set of resources that
changed this pass. So the one kind of resource that hosts a module's runtime is the kind that
cannot say "restart me when my config changes."
The effect is quiet, which is the worst part: setting a value appears to succeed (the file is
correct on disk), and the running tools keep answering with the old configuration, or keep
failing to load because the value that would fix them is present but unread.
## Why it matters beyond this instance
Every tools+events module converted to the runtime model now takes its URL and credentials
this way, so this is not one module's quirk — it is the config path for the whole catalogue.
The gap turns the headline promise of the settings design ("change it without editing it, on
the fly") into "change it, then recreate the container by hand," which is the manual step the
design existed to remove. And because the file is genuinely updated, nothing surfaces the
staleness; a dashboard that set the value would report success while the mesh kept doing the
old thing.
Config set **before** the runtime first starts (settings, then assign, then push) does work —
the file is right when the process reads it. So the gap is specifically about *updates* to an
already-running runtime, which is precisely the case the "on the fly" promise is about.
## Open questions
- Should a `container` gain `restart-on`, mirroring the service field, so a module can point
it at its config resource?
- Or should the apply path recreate a container when a file it mounts changed this pass —
making mounted-file content behave like part of the spec, without a new field to declare?
- Should the config file's content (or a hash of it) fold into the container spec, so an
ordinary spec-diff already catches it? That restarts on every change with no new mechanism,
at the cost of a spec that is no longer only the container's own declaration.
- Is a restart even the right primitive for a runtime that could instead watch its config
file and rebuild its clients in place — and if so, is that each module's job or the
runtime host's?
@@ -0,0 +1,77 @@
---
status: resolved
opened: 2026-09-05
located-in: [mesh-control, mesh-catalog]
fixed-by: ADR 0054 (a slug for the login) + a shorter minted secret (mesh-control)
amended-design: 0054-a-consumers-identity-fits-the-tightest-backend.md
---
# The mesh's derived login does not fit every backend's identity rules — S3 rejects it
## What was observed
Proving the provider/consumer contract per backend (ADR 0053), redis and postgres passed: a
consumer authenticated against the provider with the login the mesh derived and the password
the mesh minted. **minio failed**, and not on the credential — on the *name*:
```
mc: <ERROR> Unable to add a new service account. The access key is invalid.
(access key length should be between 3 and 20).
```
The mesh derives a consumer's login as `mesh_<node>_<module>` — here `mesh_anchor_bucketuser`,
22 characters. That is a valid postgres role and a valid redis ACL user, so those providers
create it verbatim. S3 access keys are capped at **20 characters**, so minio refuses to create
the service account under it, and the provisioner retries forever while the consumer, holding
that same too-long access key, could never present it either.
## Why it matters beyond this instance
ADR 0053 says a provider creates *exactly* the login the mesh derived, so that the two ends
agree by construction — the mesh hands the same name to the provider (to create) and the
consumer (to present). That only holds if the derived name is one every provider can accept.
It is not: the mesh's `as` is a single format with no knowledge of a backend's identity rules,
and S3's are stricter than a database's. Any provider whose backend constrains identifiers more
tightly than postgres — a length cap, a charset, a required prefix — inherits this, and the
failure lands at provision time, per consumer, as an infinite retry rather than a refusal at
assignment.
This also shows the seam is real, not cosmetic: `as` is doing two jobs — a stable per-consumer
identity the two ends must agree on, and a literal identifier a specific backend must accept —
and those are not always the same string.
## The shape of a fix (open, not decided)
- **Constrain the derivation** so `as` is broadly acceptable — short (≤ 20), a conservative
charset, deterministic. This keeps "the provider creates exactly what the mesh derived" true
everywhere, at the cost of a less legible name, and it is a mesh-wide identity change (every
provider that already created the longer name would see it change).
- **Let a provider map `as` to a backend-valid identifier** it derives the same way on create
and on the consumer's behalf — but the consumer is generic and cannot run minio's mapping, so
this only works if the mapped identifier is *delivered back* to the consumer. That is the
data-provision return path this era keeps meeting (umami's siteId, cloudflare's record) and
does not yet have.
- **Declare the constraint on the interface** (`s3-bucket` states its identifier bounds) and
have the mesh derive within them — the most honest, the most work.
## Open questions
- Is `as` meant to be human-legible, or is a short opaque token acceptable — i.e., can the
derivation simply be shortened without anyone minding?
- Do redis/postgres actually want the long name, or did it only survive because they are
permissive? If nothing needs it long, the cheap fix is to cap it.
- Does this fold into the same decision as the data-provision return path, or is it separate?
## Resolution
Accepted **ADR 0054** (option E): a module declares an optional short `slug`, and the mesh derives
`mesh_<node>_<slug|name>`, bounded by the tightest backend (an S3 access key's 20) and refused at
assignment — naming the slug as the remedy — when it still would not fit. The minio grant e2e proved
it: `bucketuser` declares `slug: bkt`, so its access key `mesh_anchor_bkt` (15) is accepted where
`mesh_anchor_bucketuser` (22) was refused.
Proving that surfaced a **second S3 length constraint on the same credential** — the secret. The
mesh minted a 43-character password (32 random bytes, base64url), and an S3 secret key is 8–40. Fixed
in `mesh-control` `internal/secrets/seal.go` by minting 30 bytes → exactly 40 characters (240 bits,
ample), which fits S3 and every other backend. Both halves of an S3 credential — the access key
(login) and the secret key (password) — now fit the tightest backend, by the same rule.