diff --git a/00-META/process/00-overview.md b/00-META/process/00-overview.md index 0926f21..ba54ecf 100644 --- a/00-META/process/00-overview.md +++ b/00-META/process/00-overview.md @@ -50,6 +50,7 @@ the expensive half. | [04](04-build-handoff.md) | Build handoff | A design is ready to be built | | [05](05-constitution-sync.md) | Constitution sync | `how-we-build.md` changed a rule the mesh enforces | | [06](06-writing-a-module.md) | Writing a module | Something that runs today must run on the mesh | +| [07](07-feature-branches.md) | Feature branches across repos | Work that changes code, in one repo or several at once | ## Status lives in frontmatter diff --git a/00-META/process/07-feature-branches.md b/00-META/process/07-feature-branches.md new file mode 100644 index 0000000..bb5f451 --- /dev/null +++ b/00-META/process/07-feature-branches.md @@ -0,0 +1,57 @@ +# Playbook 07 — Feature branches across repos + +**Trigger.** Work that changes code — in one code repo or in several at once (`mesh-sdk`, +`mesh-control`, `mesh-catalog`, `mesh-host`, `mesh-lab`, and `hq` when a decision rides along). + +**Who runs it.** Anyone who writes code, engineers and agents alike. Agents follow it exactly — +it is the guard against the failure it was written for. + +## The failure it prevents + +A feature was worked as a branch-and-MR per *unit of thought* — one per decision, one per +stacked increment — and each MR was treated as finished when it was *opened*, not when it was +*merged*. Across repos the same feature took a different branch name in each. The MRs piled up +unmerged: one session left **sixteen** stacked intermediate MRs that had to be consolidated and +closed by hand. An MR is a review checkpoint, not a scratchpad. + +## The rule + +One feature is **one branch name**, **one worktree per repo**, **one MR per repo**, opened +**once, at the end**. + +1. **Name the feature once.** `feat/`. The *same* branch name in every repo the feature + touches — never a different name per repo, never a fresh branch per increment within the + feature. +2. **Isolate each repo.** One git worktree per touched repo under `.work//`, branched + off `main`: + ``` + git worktree add .work// -b feat/ origin/main + ``` + Parallel features never collide, and no shared checkout is edited. +3. **Commit as you go — locally.** Increments land on the one branch. Nothing is pushed and no + MR is opened mid-feature. +4. **Finish, then publish.** When the whole feature is done — every repo, tests green — push + every branch and open **one MR per touched repo**, together. +5. **Merge promptly, once approved.** Every merge into `main` is notified and approved + ([ADR 0023](../../02-DECISIONS/0023-approval-is-the-checkpoint.md)); once it is, merge — + do not leave it sitting. The branch is deleted on merge. +6. **Leave nothing behind.** After the MRs merge, no `feat/` branch and no `.work/` + worktree survive. + +## What this is not + +- **Not a licence to batch unbounded work.** A feature is a *bounded* unit; if it sprawls for + days, end-of-feature bloat merely replaces per-increment bloat. Split it into features, each + its own branch and MR. +- **Not a second trunk.** Every repo branches off `main`. There is no longer an `initialization` + trunk. + +## How it is checked + +The end state is visible, and its absence is the smell: + +- After a feature merges, `git branch -r | grep feat/` and `git worktree list` return + nothing for it. A surviving branch or worktree means step 6 was skipped. +- More than one open MR in a repo that share no feature name, or a stack of MRs none of which is + merged, is the failure this playbook exists to prevent — stop and consolidate before opening + more. diff --git a/00-META/repos.md b/00-META/repos.md index 7112bd5..d8e9194 100644 --- a/00-META/repos.md +++ b/00-META/repos.md @@ -31,7 +31,7 @@ target, not the present. | `mesh-substrate` | 1 | the four pinned services, as declarations | | `mesh-control` | 2 | **exists.** The control plane and its contexts — one of seven built ([ADR 0006](../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)) | | `mesh-surfaces` | 3 | tools, web, cli | -| `mesh-sdk` | — | contracts shared across tiers | +| `mesh-sdk` | — | the stable spine modules build against — the tool-serving harness, the messaging/event framework, the contracts and core primitives. Holds nothing per-module and nothing volatile ([ADR 0039](../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md)). | | `mesh-lab` | — | **exists.** The lab — scenario lifecycle, networking, placement. Ships to nobody; runs on a workstation. | Tier 4's shape is open, and deliberately so: see ADR 0019 and diff --git a/02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md b/02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md new file mode 100644 index 0000000..4199b09 --- /dev/null +++ b/02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md @@ -0,0 +1,106 @@ +--- +topic: building it +status: accepted +date: 2026-09-03 +deciders: jochen +reconstructed: false +--- + +# 39. What the SDK holds, and what it refuses + +_Reconciliation note (2026-09-05): supersedes the earlier "repository structure" decision, which the consolidation folded; no standalone record remains to point at, so body references to it now point at the nearest surviving record, [ADR 0015](0015-applications-live-in-their-own-repository.md)._ + +## Context + +The earlier "repository structure" decision (folded in consolidation; see the reconciliation note +above, and [ADR 0015](0015-applications-live-in-their-own-repository.md) as the nearest survivor) +named `mesh-sdk` "contracts shared across tiers: +types, not behaviour." That line is superseded here, because it draws the boundary in the wrong +place. The boundary that matters is not *types versus behaviour* — it is **how often the thing +changes**. + +The current SDK is the cautionary tale, and its failure is precise. `hal/sdk` holds all the +code, including a per-module API client for every service (`clients/plex.ts`, `clients/gitea.ts`, +…) and a per-module tool implementation for each (`tools/plex.ts`, …). Every module depends on +the SDK, so **every edit to any of that per-module code rebuilds every module** — the cascade. +The SDK is under constant maintenance precisely because it became the place all the volatile +per-module logic accumulated. + +The root cause is worth stating exactly, because the fix follows from it: the pressure was never +to share a client *between* modules. It was to share a client between one module's *own features* +— plex's tools, its health check and its hooks all wanted the same `PlexClient` — and the only +place to share code across a module's features was the global SDK. So **intra-module sharing +leaked out as inter-module coupling.** + +## Decision + +The SDK holds the **stable spine** that modules build against, and earns its place by rarely +changing. The test for membership is change-frequency, not kind. + +### What it holds + +- The **tool-serving harness** — the worker and registration mechanism, and the tool-definition + type. *How* a tool is declared and served is settled; it does not change when an individual + tool does. +- The **messaging and event framework** — the broker client, the event consumer, the envelope. +- The **contracts** — the manifest, declaration, provision and link shapes. +- **Core primitives** — sealing and crypto, semver, the shared resolution helpers. + +These change rarely and deliberately. When one of them does change, a rebuild of everything is +the *correct* outcome, because the contract every module shares has genuinely changed. + +### What it must not hold — the more important half + +- **A module's API client.** A Plex client, a Gitea client, a MinIO client belong in their + module. They change when that service's API or the module's use of it changes, which is often, + and which has nothing to do with any other module. +- **A module's tool implementations.** Same reason, same place: in the module. +- **Anything volatile** — anything that changes when one service's features change. + +The rule, stated so it can be applied without re-deriving it: + +> If editing a thing recompiles unrelated modules **and** it changes often, it does not belong +> in the SDK. + +Both conditions are load-bearing. A rare change that cascades is fine — that is a contract, and +the cascade is correct. A frequent change that stays local is fine — that is a module minding its +own business. Only **frequent *and* cascading** is the disease, and per-module clients and tools +are its carriers. + +### Where per-module shared code lives instead + +Code shared among a module's *own* features lives **in the module**. The default is the plainest +thing that works: an ordinary shared file the features import — `plex/client.ts`, imported by +`plex/tools/`. Within one module, features are files importing sibling files; no package +boundary, no ceremony. + +A **module-local SDK** (a sub-package with its own version) is warranted only for the few modules +whose shared surface is large enough to version on its own. It is the exception, not the shape. + +Either form gives the property the global SDK could not: editing a module's shared code rebuilds +**that module and nothing else**. + +## Consequences + +- The cascade becomes **structurally impossible for module logic**. There is no longer an edge + from one module's internals to another, so the only thing that can rebuild everything is a real + change to a shared contract in the SDK — which is rare, and when it happens, is right. +- The SDK is small and stable **by construction**, not by discipline. Its size is no longer a + thing anyone has to police. +- **Converting a module from the current system is partly a de-coupling, not just a move.** Its + client and its tools are pulled *out* of the shared SDK and *into* the module. A conversion + that copied `clients/plex.ts` into the SDK's replacement would rebuild the exact mistake. +- The host still does not import the SDK. It depends on nothing + ([ADR 0005](0005-the-node-host.md)) and **mirrors** the contracts rather than + importing them, exactly as its apply-shapes table already does deliberately. The SDK is shared + by the tiers that *can* share code; the host is not one of them. + +## References + +- The earlier "repository structure" decision — named the repositories; its `mesh-sdk` + description ("types, not behaviour") is superseded by this record (folded in consolidation; + nearest survivor [ADR 0015](0015-applications-live-in-their-own-repository.md)). +- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — the decomposition this serves: code + belongs to the boundary that owns it. +- [ADR 0005](0005-the-node-host.md) — why the host mirrors the contracts instead of + importing the SDK. diff --git a/02-DECISIONS/0040-what-a-module-is.md b/02-DECISIONS/0040-what-a-module-is.md new file mode 100644 index 0000000..c94825d --- /dev/null +++ b/02-DECISIONS/0040-what-a-module-is.md @@ -0,0 +1,101 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-03 +deciders: jochen +reconstructed: false +extends: 0009-modules-and-the-graph.md +--- + +# 40. What a module is + +_Reconciliation note (2026-09-05): supersedes the earlier "grouped by domain" decision, which the consolidation folded into how-we-build.md; no standalone record remains to point at._ + +## Context + +[ADR 0009](0009-modules-and-the-graph.md) settled that everything is a module, but never said what a +module *is* beyond "a directory the mesh processes." That gap let the catalogue's breadth read as a +smell: a module can carry a container, a built image, tools, a provisioner, migrations, health, +config, seat claims, requires and provides — so much that the unit seemed ill-defined. +The earlier "grouped by domain" decision (folded in consolidation; see the reconciliation note +above) tried to organise modules by domain, which is the wrong axis. This record states what a module is, drawn from the cases that +stress-tested it: the shell, i3-vs-sway, umami, and "database." + +## Decision + +**A module is one self-contained piece of software the mesh installs and manages** — everything +needed to make that one thing real and integrable: what runs, the seats it claims, what it provides +to other modules, what it requires from them, and what operates it. + +The **software is the module's identity.** Capabilities, seats and provisioned resources are the +**relationships *between* modules**, not what a module is — and that is what binds a module into one +thing. umami is bound by *being umami*: its container runs umami, its provisioner creates umami sites, +its tools query umami, its `requires` gets umami a database. Every feature serves the one software. + +### The three relationships + +1. **Shared seat** — several modules fulfil a capability and coexist; one may be default. bash, zsh + and fish all join `shell`. +2. **Exclusive seat** — modules contend for a single slot; one holds it. i3 (needs x11) and sway + (needs wayland) contend for `display-session`. +3. **Provide / require** — a provider ships the **provisioner** that creates instances of the + resource it offers and returns sealed credentials; a consumer requires it and the mesh wires the + credential in. Symmetric: umami requires a database *and* provides analytics. + +### Interfaces are mesh-owned; providers adapt to them + +The mesh **defines the interface** for a capability — the provider-neutral contract of what a +consumer receives and how it integrates. Both sides conform: a provider's provisioner **adapts** its +software's real API to the mesh contract; a consumer depends on the **interface**, never on a +provider. Swap one provider for another and the consumer does not change. + +### The naming rule — draw the interface at the consumer's real coupling + +Name a `provides`/`requires` at the **widest boundary across which the consumer genuinely does not +care which implementation serves it**: + +- Where the consumer's coupling is thin — an analytics embed snippet and dashboard, opaque to it — + the mesh defines a neutral interface (`analytics`) and providers (umami, amumi) adapt. Swappable + across vendors. +- Where the consumer **speaks a protocol** — a database's wire protocol and query dialect — the + interface *is* the protocol: `postgres-database`, `mssql-database`, `mongodb-database`. Swappable + only among protocol-compatible implementations, **never across**, because the application cannot + cross it either. "database" is not a capability; the protocol is. +- **Never false genericity.** A name must not promise a swap the contract cannot deliver + ([research 005](../01-RESEARCH/005-domain-grouping/analysis.md)). + +This is [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md)'s rule made general — "names the protocol, not +the product; a database names the engine because the app targets it" — with the reason stated: the +contract sits where the coupling is. + +### What is not a module + +- A **library** (built against, never deployed — [ADR 0039](0039-what-the-sdk-holds-and-refuses.md)). +- A **control-plane context** (the mesh itself — [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)). + +A **swappable machine mechanism** (a firewall — ufw, nftables) *is* a module implementing a +capability. The host hardcodes no firewall, supervisor, package manager or runtime; it owns only the +generic apply primitives and platform detection, so it runs where none of those exist — an Android +phone has no ufw, systemd, pacman or Docker. + +## Consequences + +- **Supersedes the earlier "grouped by domain" decision** (folded in consolidation; see the + reconciliation note above). Modules are + organised by their relationships (seats, provisions), not grouped into domain folders. +- **Refines [ADR 0009](0009-modules-and-the-graph.md).** Everything the mesh runs and integrates is + a module — but a module is defined by the *software it delivers*, not by being a bucket of features. +- The target is a **self-fulfilling mesh**: declared wants bound to swappable modules, provisioners + wiring credentials, nothing hardcoded. The control plane's whole job is the binding. +- Converting a module from the old system includes pulling its per-module code out of the shared SDK + ([ADR 0039](0039-what-the-sdk-holds-and-refuses.md)) and shipping its provisioner as an adapter to a + mesh interface — a de-coupling, not just a move. + +## References + +- [ADR 0009](0009-modules-and-the-graph.md) — everything is a module; this says what one is. +- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — contexts are the mesh, not modules. +- The earlier "grouped by domain" decision — superseded (folded in consolidation; see the note above). +- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — protocol-not-product, generalised here. +- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — per-module code lives in the module. +- [research 011](../01-RESEARCH/011-the-module-graph/00-overview.md) — the graph of these relationships. diff --git a/02-DECISIONS/0041-events-are-a-relationship.md b/02-DECISIONS/0041-events-are-a-relationship.md new file mode 100644 index 0000000..a381307 --- /dev/null +++ b/02-DECISIONS/0041-events-are-a-relationship.md @@ -0,0 +1,85 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-03 +deciders: jochen +reconstructed: false +extends: 0040-what-a-module-is.md +--- + +# 41. Events are a relationship, the lighter sibling of provisioning + +## Context + +[ADR 0040](0040-what-a-module-is.md) names two relationships between modules — seats and +provide/require (provisioning). A third is latent in the mesh and worth making first-class: the +broker every node already runs ([ADR 0002](0002-nodes-communicate-over-a-broker.md)) can carry a +module's activity as **events**, which any other module reacts to. A logger that writes an audit +trail, a module that acts when another module acts, observability — all of it is one mechanism, and +today it is ambient rather than declared. + +## Decision + +**A module emits events and consumes events, and both are declared** — parallel to `provides` / +`requires`, so the mesh knows the event graph the same way it knows the provisioning graph. + +### Events are provisioning's lighter sibling + +| | provisioning | events | +|---|---|---| +| shape | **1:1**, a provider creates a resource *for* one consumer | **1:many**, a module emits, any number listen | +| credential | yes — sealed, per consumer | none — it is broadcast | +| machinery | a provisioner (the reconcile adapter) | nothing but the broker's topic routing | +| declared as | `provides` / `requires` | `emits` / `consumes` | + +Because an event is broadcast and credential-free, there is no provisioner and no per-consumer +setup — only a subscription. That is why it is the *lighter* relationship, and why most +inter-module reaction should be an event, not a provision. + +### An event carries what an audit needs + +Every event carries its **type** (a dotted topic key, so listeners match by prefix), its **source** +module, the **node** it came from, and the **time**. A body follows. The metadata is not optional: +a reaction may only need the body, but an audit trail needs to know who did what, where and when, +and an event that cannot answer that is not auditable. + +### The audit logger is just a consumer of everything + +A logger that records the whole mesh's activity is **not a privileged component** — it is an +ordinary module that consumes `#` (every event) and writes them down. It holds no special access; +it only listens widely. That it falls out of the model with no new machinery is the check that the +model is right. + +### `consumes` is validated like `requires` + +A `consumes` for an event that **nothing** `emits` is a dangling edge, and the mesh refuses it +before deploy — the same rule that catches a `requires` for a resource nothing provides +([research 011](../01-RESEARCH/011-the-module-graph/00-overview.md)). A listener waiting for an +event that can never arrive is a silent failure, and this repository's whole discipline is against +silent failure. + +### One runtime serves all three + +The per-node module runtime that serves a module's tools also wires its `consumes` (subscribe, +dispatch to the handler) and lets its code `emit`. Tools are *invoked* (request/reply), resources +are *provisioned* (1:1, credentialed), events are *emitted and consumed* (1:many, broadcast) — +three relationships, one broker, one runtime, all declared on the manifest. + +## Consequences + +- The mesh gains a declared **event graph** alongside the provisioning graph — visible, validated, + reasoned over. +- **Reaction becomes the default coordination**: a module acts on another's event without either + knowing the other, and without a credentialed link. Coupling drops. +- An **audit trail** is a module, not a platform feature — and can be swapped, extended or run more + than once (a file logger and a queryable one) with no change to anything that emits. +- The runtime must dispatch a module's event handlers as well as its tools; that generalisation is + small (both arrive by importing the module's entrypoint) but it is real work. + +## References + +- [ADR 0002](0002-nodes-communicate-over-a-broker.md) — the broker events ride. +- [ADR 0040](0040-what-a-module-is.md) — the relationships this extends. +- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — `emit`/`on` are stable sdk surface; the + broker binding and the runtime are not. +- [research 011](../01-RESEARCH/011-the-module-graph/00-overview.md) — the graph these edges join. diff --git a/02-DECISIONS/0042-the-shape-of-an-event-on-the-wire.md b/02-DECISIONS/0042-the-shape-of-an-event-on-the-wire.md new file mode 100644 index 0000000..0479407 --- /dev/null +++ b/02-DECISIONS/0042-the-shape-of-an-event-on-the-wire.md @@ -0,0 +1,116 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-03 +deciders: jochen +reconstructed: false +extends: 0041-events-are-a-relationship.md +--- + +# 42. The shape of an event on the wire + +## Context + +[ADR 0041](0041-events-are-a-relationship.md) made events a relationship — `emits`/`consumes`, the +graph, the audit logger. It did not say what an event *is* on the broker: the exchanges, the +routing keys, the headers, the queues and their configuration. That shape is a contract every +emitter and consumer conforms to, exactly as [ADR 0010](0010-delivery.md) +is for declarations — and it was being decided ad-hoc in code. This settles it, so the sdk and the +runtime implement one contract and a module never reinvents it. + +## Decision + +### Two exchanges, kept apart + +- **`mesh.events`** — a durable topic exchange. Every event rides it: module, mesh and node. +- **`mesh.rpc`** — a durable topic exchange. Tool invocations (request/reply) ride it. + +Kept separate because RPC is not an event: a `#` subscription on `mesh.events` is then a complete +audit of what happened, with none of the invocation traffic. + +### The routing key is the event type, namespaced by origin + +Dotted and hierarchical — `..` — with three reserved origins: + +- `module..` — `module.umami.site.created` +- `mesh..` — `mesh.delivery.deployed`, `mesh.provisioning.granted` +- `node..` — `node.anchor.joined`, `node.anchor.unreachable` + +Topic matching gives a consumer `node.*.joined`, `module.umami.#`, or `#`. The origin roots are +reserved; everything after is the emitter's own namespace. + +### Metadata in headers, payload in the body + +An event's identity and provenance are AMQP **headers**, so a consumer — or the broker, or an +audit tool — reads who/when/what without parsing the body, and the body is only the domain payload. + +**Required headers** + +| header | meaning | +|---|---| +| `x-event-id` | a unique id — for dedup and audit (delivery is at-least-once, below) | +| `x-source` | the emitter: the module, context or node name | +| `x-node` | the node it was emitted from | +| `x-time` | emit time, RFC-3339 | +| `content-type` | `application/json` | + +**Optional headers** + +| header | meaning | +|---|---| +| `x-causation-id` | the event or command that caused this one — tracing | +| `x-schema` | a version of the body's shape, so a body evolves without silent misreads | + +The routing key already carries the type; it is not duplicated as a header. An **unknown `x-` +header is ignored, not refused** — unlike a declaration, an event is observed by parties that need +not all understand every header, and refusing would couple every consumer to every emitter's +additions. + +### Messages are persistent + +Events are published persistent (delivery-mode 2). An audit trail that loses events on a broker +restart is not one, and the cost is disk the broker already spends on everything durable. + +### Queues: one per consumer, durable, dead-lettered + +- **A consumer's queue** is `..events`, durable, bound to that module's consumed + patterns. Durable so a restart does not drop what arrived while it was down. **Manual ack** after + the handler succeeds — at-least-once. +- **Prefetch** bounds in-flight work (default 32) so one slow consumer does not pull the whole + backlog into memory. +- **A dead-letter exchange** `mesh.events.dead` receives a message rejected past a redelivery limit, + so a poison event is set aside for inspection rather than looping forever or vanishing silently. +- **The audit logger's queue** `.audit-logger.events`, bound to `#`, is the same shape — + durable, persistent, dead-lettered — because completeness is its whole job. +- **RPC reply queues** are exclusive, auto-delete and server-named; **RPC serve queues** + `serve.` are durable and shared, so several runtimes serving one tool key compete rather than + each answer. + +### At-least-once, and consumers are idempotent + +A handler may see an event twice — a redelivery after a crash between doing the work and acking. +Consumers must be idempotent, and `x-event-id` is what makes dedup possible. **Exactly-once is not +offered**: it is a promise no broker keeps honestly, and saying so is better than pretending. + +## Consequences + +- The event shape is a versioned, enforced contract, not conventions each module reinvents. The + sdk's `emit`/`on` and the runtime's AMQP binding implement it; a module never sees an exchange or + queue name. +- Metadata-in-headers means the body is exactly the domain payload, and a consumer that only wants + provenance never parses it. +- Adding a header or an origin root widens the contract and is reviewed as one — the discipline + [ADR 0010](0010-delivery.md) applies to the + declaration vocabulary. +- The sdk's first cut carried source/node/time in the *body*; this supersedes that — they move to + headers. That is code to align, in `mesh-sdk` (`emit`/`on`) and `mesh-tools` (the binding, queue + config, dead-letter). + +## References + +- [ADR 0041](0041-events-are-a-relationship.md) — events as a relationship; this is their wire shape. +- [ADR 0010](0010-delivery.md) — the precedent: a wire + contract, versioned, additions reviewed as security. +- [ADR 0002](0002-nodes-communicate-over-a-broker.md) — the broker. +- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — `emit`/`on` are stable sdk surface; the + binding, queue config and dead-letter are the runtime's, not the sdk's. diff --git a/02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md b/02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md new file mode 100644 index 0000000..cefd111 --- /dev/null +++ b/02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md @@ -0,0 +1,101 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-04 +deciders: jochen +reconstructed: false +extends: 0041-events-are-a-relationship.md +--- + +# 43. A module's broker account is scoped by what it emits and consumes + +## Context + +[ADR 0041](0041-events-are-a-relationship.md) made events a relationship — `emits` and `consumes` +on the manifest. [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) gave them a wire shape — the +`mesh.events` exchange, the durable per-consumer queue, the reserved routing-key origins. Neither +said how a module *reaches* the broker: what account it holds, and what that account is allowed to +do. + +As the code stands, there is no answer. The mesh can provision a **node** account (at enrolment) +and a **builder** account (scoped to the build queue), and it can *deliver* any module a sealed +own-secret at a declared path — but it has no way to provision a broker **account** for a general +module. A module that declares `own-secrets: {broker: …}` and nothing more receives thirty-two +random bytes, not a credential. So on the broker, `emits` and `consumes` are enforced by nothing: a +running module could bind any queue, consume any pattern, and publish under any origin, and the +manifest that says otherwise would be describing a boundary no code draws — the exact shape of fault +[04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) records, a scope +declared in manifests and read by nothing. + +This settles it, so a module's place on the bus is a thing the broker enforces rather than a thing +the manifest merely claims. + +## Decision + +### A module gets a broker account when it is assigned, and its permissions are the manifest + +When the mesh assigns a module to a node it provisions a broker account for that module on that node, +sealed to the node ([ADR 0004](0004-a-node-and-how-it-joins.md)) and delivered as the +module's `own-secrets` broker — `amqps://` with the mesh's fingerprint, the shape +[ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) already carries. The account's permissions are +derived from the manifest, and are exactly these: + +- **What it consumes.** Read on `mesh.events`, and configure-and-read on its own queue + `..events` bound to the patterns in `consumes`. It cannot bind or read another + module's queue. A module that consumes nothing gets no read on the events exchange at all. +- **What it emits.** Write to `mesh.events`, restricted to routing keys under its own origin, + `module..*`. It cannot publish as another module, and cannot publish under the reserved + `mesh.*` or `node.*` origins — those belong to the mesh and the host (ADR 0042). A module that + emits nothing gets no write. +- **Nothing else.** The events account reaches `mesh.events` and that module's own queue, and no + more. Tool serving and calling over `mesh.rpc` is a separate grant on the same principle — a + module serves the tool keys it declares and calls the ones it is bound to — and is scoped the same + way rather than folded in here. + +### Consuming everything is a privilege, granted deliberately + +`consumes: ["#"]` — the audit logger — is read across the whole bus: every module's events, the +mesh's, every node's. That is not a pattern like any other; it is the power to see everything, and +the account is where it becomes visible. The grant that lets one module read the entire bus is one +the mesh issues on purpose and can be audited — the answer to *who can read everything* is a row, not +a guess — rather than a breadth any manifest acquires by typing a single character. A `#` consume is +a reviewed grant, not a default one. + +### The account is how the declaration is enforced + +Because the account can do only what `emits` and `consumes` name, the broker itself refuses a module +that tries to consume a queue it did not declare or emit under an origin it does not own. That is what +makes an event relationship a rule and not a comment — the discipline that a stated rule says how it +is checked. A manifest that over-declares grants more than the module uses, which is visible and +reviewable; one that under-declares makes the module fail closed at the broker, which is the safe +direction to be wrong in. + +## Consequences + +- The control plane gains a **generic module broker-account**, derived from the manifest. The + builder stops being a special case: its access to the build queue becomes an ordinary expression of + what it consumes and serves, not a bespoke account method. One rule, and the builder is an instance + of it. +- The runtime reads its credential from a file (the broker own-secret), `amqps://` verified against + the mesh's fingerprint. The `guest` account is for raising the substrate, never for a module — a + module documented as holding its own credential and handed the broker's administrative one is worse + than one with no credential story at all. +- `emits` and `consumes` stop being advisory. They are the module's authority on the bus, so the + manifest is now a security boundary and is reviewed as one, the discipline + [ADR 0010](0010-delivery.md) applies to the declaration + vocabulary. +- *Who can read the whole bus* becomes an answerable question, because `#` is a grant and not an + accident. + +## References + +- [ADR 0041](0041-events-are-a-relationship.md) — events are a relationship; this scopes the account + by that relationship. +- [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) — the wire this account secures: the queue, + the origins, the `amqps` credential shape. +- [ADR 0004](0004-a-node-and-how-it-joins.md) — the link is the security boundary; a + module's account is sealed to its node the same way a node's is. +- [ADR 0010](0010-delivery.md) — a declaration is owned + and its additions reviewed; a module's broker permissions are that discipline applied to the bus. +- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — a scope declared + in manifests and enforced by no code: the fault this decision closes for events. diff --git a/02-DECISIONS/0044-a-public-name-is-provisioned-like-any-capability.md b/02-DECISIONS/0044-a-public-name-is-provisioned-like-any-capability.md new file mode 100644 index 0000000..d86e0a5 --- /dev/null +++ b/02-DECISIONS/0044-a-public-name-is-provisioned-like-any-capability.md @@ -0,0 +1,94 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-04 +deciders: jochen +reconstructed: false +extends: 0027-a-provision-names-what-the-consumer-is-coupled-to.md +--- + +# 44. A public name is provisioned, not registered by hand + +## Context + +The mesh names and resolves its own machines internally: the overlay generates +`..` wildcards, dnsmasq answers them (`wildcard-resolution`), and the mesh +issues a certificate for each internal name. A service reachable at a *public* domain — +`plex.example.com`, not `plex.anchor.internal` — needs three things that machinery does not give it: + +- a **public DNS record** at a registrar or DNS provider, so the name resolves on the internet; +- a **publicly-trusted certificate** for it, because the mesh's own authority is trusted by nobody + outside the mesh; +- and routing from that name to the module — which the reverse proxy already does: a module + `requires` the `route` capability and the proxy provides it, routing by the host it was asked for. + +The routing exists. The public DNS record does not: the mesh has no way to make a name resolve on +the public internet, so today that is a step someone does by hand at a DNS provider, outside the +mesh, remembered nowhere. A public name is therefore the one part of reaching a service that the +declaration graph cannot grant or withdraw — which means it is created once and outlives whatever it +was for, the shape of drift this project exists to remove. + +## Decision + +### A public name is a capability, requested like any other + +A module reachable at a public host declares `requires: ["public-dns"]` and contributes the hostname +it wants — beside `requires: ["route"]`, which exposes it through the proxy. The name is then +provisioned on declaration ([ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md)): created +when the module is assigned, removed when it is withdrawn, reconciled like every provision. + +### The interface is neutral; the providers are the registrars + +`public-dns` is drawn at the consumer's coupling: the consumer wants *a public name that resolves to +me*, and does not care whether Cloudflare, Route 53 or a registrar's own API puts the record there. +So the interface is neutral and the providers are provider-scoped — `cloudflare-dns`, +`route53-dns`, `porkbun-dns` — each implementing the one `public-dns` contract, the same way a +neutral database coupling is answered by `postgres-database` and `mssql-database`. A module names +`public-dns`; it never names a registrar. + +### The record points at the mesh's public ingress, not at the node + +What the name resolves to is the address the reverse proxy answers on, not the consuming machine's. +A public service is reachable only *through* the proxy — the proxy holds the `route` grant and routes +by host to the module — so the public name must resolve to the proxy. `public-dns` and `route` are +the two halves of one public exposure: the name, and what the name reaches. + +### The record is a fact, not a secret + +A DNS record is public by definition, so the grant returns the fully-qualified name and its TTL and +nothing sealed. The only secret is the provider's own API credential, which is the provider module's +own-secret and never leaves it — the module that wanted the name never sees it. + +### Events + +The provider emits `module..record.created` and `module..record.removed` +([ADR 0041](0041-events-are-a-relationship.md)), so *which names the mesh publishes, and where* is a +question answered from the event trail and the grants, not from a folder of records edited at a +provider. + +### The public certificate is the proxy's, and is named here only to pair it + +A public name without a publicly-trusted certificate is reachable and not trusted — the same pairing +the internal name and the mesh-issued certificate already have. Obtaining that certificate (ACME +against the now-resolving public name) is the reverse proxy's to do, and its mechanism is its own +decision; it is named here so the pairing is not forgotten, not resolved here. + +## Consequences + +- A public name is created and torn down with the module, so it cannot outlive it, and the mesh can + say which public names it publishes without anyone reading a registrar's dashboard. +- Adding a registrar is adding a provider that answers `public-dns`; the modules that want names do + not change. +- Public exposure of a service is a trio of separate, declared, enforced relationships: the firewall + opens the proxy's public port ([ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)), + `route` routes the host to the module, and `public-dns` makes the host resolve. + +## References + +- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — a capability is provisioned on + declaration; a public name is one. +- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the firewall, the other + half of the reachability question this was asked with. +- [ADR 0041](0041-events-are-a-relationship.md) — the provider's record events. +- [ADR 0040](0040-what-a-module-is.md) — a provider and its interface; the neutral-interface, + scoped-provider naming this follows. diff --git a/02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md b/02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md new file mode 100644 index 0000000..a5c61ac --- /dev/null +++ b/02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md @@ -0,0 +1,93 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-04 +deciders: jochen +reconstructed: false +extends: 0005-the-node-host.md +--- + +# 45. A machine's firewall is the sum of what its modules listen on + +## Context + +The reverse proxy is a *provider*: a module `requires` the `route` capability and a running proxy +provides it, routing traffic by name and reaching back to the consumer. A fair question follows — +is the firewall the same shape? Should a module *register* a port with a firewall provider the way +it requests a route? + +It should not, and the difference is the point. A reverse proxy is a service another component +performs; a firewall is a property of the machine — a packet filter the host applies to itself. +Modelling it as a provider would invent a credential and a reach-back for something that has neither. + +And the mesh already has the registration: a module declares `listens: [{ port, from }]` — the port +it accepts connections on, and from where. That *is* how a service says it wants a port open. What is +missing is not a model but enforcement. [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) +records that a `scope:` key five manifests carry is read by no code: a manifest can appear to +restrict a port and restrict nothing — the exact fault +[how-we-build.md](../00-META/how-we-build.md) names, *an unenforced rule is indistinguishable from a +wrong one*, made worse because the declaration reads as a restriction. + +## Decision + +### The firewall is derived and host-applied, not a provider + +A machine's firewall is the sum of what the modules assigned to it declare they listen on, computed +by the host and applied as one of its owned resources ([ADR 0005](0005-the-node-host.md): +the host applies, it does not decide; [ADR 0010](0010-delivery.md): +the declaration is owned resources). It is not a capability, not a per-consumer grant — opening a +port is a declarative fact about a machine, so it is computed and applied, not requested and +credentialed. + +### `from` is the whole of public-versus-internal + +The distinction the question is really about lives in `from`: + +- `listens: [{ port: 5432, from: mesh }]` — open to the private overlay only. +- `listens: [{ port: 443, from: anywhere }]` — open to the public internet. + +A module registers a port on the firewall by listening on it and saying from where. There is no +separate firewall capability, because the firewall is not a thing that reaches back or holds a +secret; it is the machine's own filter over the ports its modules named. + +### The host enforces it both ways, and unknown keys are refused + +A port a module listens on is opened to exactly the scope it named; a port nothing declares is +closed. And a key the firewall does not read — the `scope:` of issue 003 — is refused at the +manifest, not accepted and ignored, so a declaration that reads as a restriction is one. This is the +discipline [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) applied to the +broker account, applied here to the packet filter: the declaration is the enforcement, or it is a +comment. + +### A public service is exposed through the proxy, not by opening its own port + +Reaching the public internet is normally not `from: anywhere` on the service's own port. The service +listens `from: mesh` — only the proxy reaches it — and `requires: route`, so the sole machine with a +public opening is the one running the reverse proxy, and the service is exposed by name through it. +`from: anywhere` is the deliberate direct-exposure case, for a service that is its own front door. + +## Consequences + +- Issue 003 is closed: the firewall is computed from `listens` and enforced, so a declared scope is + real and an undeclared port is shut. Rejecting unknown manifest keys is the general fix, of which + the `scope:` key was one instance. +- The firewall and the reverse proxy stop being confused for one model: the firewall is the machine's + filter (host-derived from `listens.from`); `route` is a name-router (a provider); the public DNS + name is a third thing ([ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md)). A + public service uses all three. +- The modelling question is answered: a module registers a port by declaring `listens`, and reaches + the public internet by name through `route` + `public-dns` — never by the firewall being a + provider. + +## References + +- [ADR 0005](0005-the-node-host.md) — the host applies; the firewall is one of + the things it applies. +- [ADR 0010](0010-delivery.md) — the firewall is a derived + owned resource, not a grant. +- [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) — the same discipline: + a declaration is enforced, or it is a comment. +- [ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md) — the public name, the other + half of the reachability question this was asked with. +- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — the unenforced + `scope:` this closes. diff --git a/02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md b/02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md new file mode 100644 index 0000000..402c5ba --- /dev/null +++ b/02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md @@ -0,0 +1,88 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-04 +deciders: jochen +reconstructed: false +extends: 0027-a-provision-names-what-the-consumer-is-coupled-to.md +--- + +# 46. A module's configuration is its assignment's, not its manifest's + +## Context + +A module is assigned to a node — `assign `, always to a machine; there is no +assignment to the mesh. "Mesh" is a *scope*, not a place: a `provides` or a `claim` scoped `mesh` +reaches the whole mesh, but the module still runs on a node. So the two kinds of thing a module can +carry are the manifest (what the module *is*) and, separately, what it should do *here* — which +differs by deployment and by node. + +The mesh already has the second: **settings**. `settings set [--node ]` — with a node +it is that machine's, without it the whole mesh's — layered over what the module declares and applied +at resolution, changeable without editing the module and without a rebuild. That is the surface a +meshboard would edit. + +But settings today reach only a module's **config-file content** (a mergeable file the module owns). +Configuration that is not a file has been landing in the manifest instead, statically — a registrar's +zone and domain, the address public names point at, and, most sharply, `listens.from`. That last one +is the tell: whether a port is open to the private overlay or to the public internet is a +*per-node deployment choice* — the same database internal on one machine and public on another — and +a value fixed in the manifest is one value for every machine, so it cannot be. Static configuration in +the manifest is configuration in the wrong place: it cannot vary per node, and it cannot change +without a new module version. + +## Decision + +### The manifest is identity and defaults; the assignment's settings are the configuration + +A module's manifest declares what it is — what it provides, requires and claims, the shape of its +resources — and, for anything configurable, a **default**. The values that make a running instance +*this* instance are settings, carried by the assignment: per-node, or mesh-wide when no node is named, +applied over the defaults at resolution. Change one and the next reconcile carries it; nothing is +edited on a machine and nothing is rebuilt. + +### Settings drive the configurable fields the manifest marks, not only file content + +Settings extend beyond a config file's content to the manifest fields a module declares settable — +foremost: + +- **`listens.from`**: a module declares its safe default (`from: mesh`), and a per-node setting + raises or lowers it. postgres declares `listens: [{ port: 5432, from: mesh }]`; on the machine that + should expose it, a setting makes that port `from: anywhere`. Same module, different exposure, and + the firewall ([ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)) is computed + from the effective value, so the packet filter follows the setting. +- **A provider's own configuration**: a registrar's zone, domain and the ingress its names point at + ([ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md)) are mesh-wide settings, not + manifest constants — one mesh's Cloudflare zone is not another's, and the module description is the + same for both. + +### Unset is the default, and an unknown setting is refused + +A field with no setting keeps the manifest's default, so a module runs correctly configured by nobody. +A setting that matches no settable field — like a config value that reaches no file today — is named, +not silently dropped, so a misspelled setting is found rather than believed (the discipline of +`UnusedSettings`, and of [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md): +a declaration is enforced or it is a comment). + +## Consequences + +- The postgres case works: one module, `from: mesh` by default, `from: anywhere` where a setting says + so — internal on ace, public on novox, changeable live. +- Provider modules stop carrying a mesh's specifics: `cloudflare-dns` describes *a Cloudflare + registrar*, and *which* zone and ingress is a setting, so the same module serves every mesh. +- Configuration becomes a thing a meshboard manages — set per node or mesh-wide, applied on the next + reconcile — rather than a manifest edit and a rebuild ([ADR 0011](0011-managed-files-are-generated-never-edited.md): + the way you change a managed thing is not by editing it). +- What a manifest may not do is grow a value that differs per machine; if it differs per machine it is + a setting, and the manifest holds only the default. + +## References + +- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — what is provisioned on + declaration; its per-instance values are the assignment's. +- [ADR 0011](0011-managed-files-are-generated-never-edited.md) — a managed thing is changed through + the mesh, not by editing it; settings are that, for configuration. +- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the firewall follows the + effective `listens.from`, so making `from` a setting makes exposure a setting. +- [ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md) — the provider whose zone and + ingress are settings, not manifest constants. diff --git a/02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md b/02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md new file mode 100644 index 0000000..1ed3edd --- /dev/null +++ b/02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md @@ -0,0 +1,87 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-04 +deciders: jochen +reconstructed: false +extends: 0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md +--- + +# 47. A module runs its code as its own process, with its own account + +## Context + +A module is one self-contained thing ([ADR 0040](0040-what-a-module-is.md)), and it gets a broker +account scoped to what it emits and consumes ([ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md)). +The catalogue now gives modules **tools** and **events** — real code, in the module ([ADR 0039](0039-what-the-sdk-holds-and-refuses.md)) — +but nothing has said what *runs* that code. The audit-logger showed one shape and was treated as an +exception: a container running the tool runtime carrying the module's compiled code, holding the +module's own scoped account. Every module with tools or events needs the same, and the tempting +alternative does not work. + +**A node-wide runtime that loaded every assigned module's code cannot hold a per-module account.** It +would run under one account with the union of every module's permissions — able to emit as any of +them and read any of their queues — which is exactly the isolation ADR 0043 exists to draw. So the +runtime is per-module, not per-node, and treating the audit-logger as special left the other +modules' code with nothing to run it: the conversion produced tools and events that, as it stands, +never execute. + +## Decision + +### A module with tools or events runs a process of its own + +A module that has tools or events runs a **runtime process** — a container, the tool runtime carrying +that module's compiled code — assigned and started like the module it is, holding the single broker +account the mesh scoped to it (ADR 0043). One module, one process, one account. + +### It serves its tools, each on its own key + +A tool is served on its own key (`serve.`), and a caller invokes a named tool. Only the module +that serves it answers, and the module's account is scoped to exactly its tool keys — so one module +cannot answer another's calls, the isolation ADR 0043 gives events extended to tools. This supersedes +a single `tools.invoke` endpoint that dispatched by name: that shape assumed one runtime for the +whole node, and per-module runtimes competing on one key would each be handed calls for tools they do +not have. + +### It runs its events in the same process, under the same account + +Emitting under the module's own origin and consuming its own queue ([ADR 0042](0042-the-shape-of-an-event-on-the-wire.md)) +happen in that same process, with that same account — not a second one to scope and seal. A module's +tool code, its event code and, for a provider, its provisioner are the one module's code and run as +the one module's process. + +### The runtime image is the tool runtime plus the module's code + +Built from the module's source like any module image — the audit-logger's shape, made the rule, not +the exception. The module declares a `container` for it carrying `MESH_BROKER_FILE` (its sealed +credential, ADR 0043) and its compiled code. A module with **neither** tools nor events runs no such +process: a plain service module — the plex *server*, dnsmasq the resolver — is its service and files +and nothing more. A module that is both a service and code declares both containers: the service, and +the runtime beside it. + +## Consequences + +- The catalogue's tools and events become runnable: each tools-or-events module gains a runtime + container with its scoped credential, and the audit-logger stops being special. Until this, the + converted modules held code with nothing to execute it. +- A process, and a small image, per tools-or-events module. That is the cost of ADR 0043's isolation: + one account per module means one process per module. It is paid deliberately — a shared runtime is + cheaper and cannot be scoped, and a mesh where any module can emit as any other is not one worth the + saving. +- `serve.` per key replaces the single `tools.invoke` dispatch. The sdk's serving and a module's + account scope both come to name tools individually. +- **A provider's provisioner is a runtime process too.** It already runs as its own container; its + events (`bucket.created`, `database.provisioned`) belong to *that* process and need the same + credential. So a provisioner that emits carries `MESH_BROKER_FILE` and its scoped account like any + runtime — or it does not emit. (This is the fix for provisioners that emit today with no broker + bound: the emit is a runtime's, and the provisioner is a runtime.) + +## References + +- [ADR 0040](0040-what-a-module-is.md) — a module is one self-contained thing; its code runs as one + process. +- [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) — the scoped account + this process holds, and the isolation that makes it per-module. +- [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) — the events this process runs, and the + `serve.` queue tools now use. +- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — the code lives in the module; this runs it. diff --git a/02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md b/02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md new file mode 100644 index 0000000..e047950 --- /dev/null +++ b/02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md @@ -0,0 +1,131 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-05 +deciders: jochen +reconstructed: false +--- + +# 48. A provider creates the credential the mesh minted, and seals nothing + +## Context + +A provider module stands up a per-consumer resource — a database, a cache bucket, an object +store user — and the consumer must end up holding a credential that authenticates against it. +Building the module runtime (ADR 0047: a module runs its own code as its own process under its +own account), the provider's provisioner was run for the first time as a delivered thing, and +it did not work. It reads a seal key from the environment that nothing sets, and it seals every +credential it produces to that key with a symmetric passphrase. + +Tracing the credential's path turned up something larger than a missing key. **The provisioner +harness the whole catalogue is built on describes a credential flow the mesh does not have, and +duplicates — incorrectly — one it does.** + +What the sdk's `runProvisioner` does today: + +- reads request files named `*.grant.json` — which nothing in the mesh writes; +- calls an adapter whose `create` **generates its own password** and returns it; +- seals that password with a symmetric key (`$MESH_SEAL_KEY`) and writes a `*.credential` + file — which nothing in the mesh reads, and no consumer ever unseals. + +What the mesh already does, and has wired end to end: + +- The control plane mints one password per (consumer, provider) pair (`Inventory.SecretFor` → + `secrets.Make`) and seals it to **both** node keys asymmetrically — a copy the consumer's + host can open and a copy the provider's host can open. No shared symmetric key exists + anywhere, on purpose: a key both ends hold is a key the mesh would have to distribute, which + is the same problem one level down, and the control plane's own code refuses it. +- The provider is handed, at the path its `receives` names, one contribution per consumer: + the **login to create** (`As`, derived by the mesh so the two ends agree by construction), + the consumer's address and requested values, and a **`Secret` file** holding that consumer's + password sealed to the provider and unsealed onto the machine by its host. +- The consumer is handed the *same* password, as plaintext its own host wrote by unsealing its + copy and substituting it into a config file. The consumer never unseals anything itself and + holds no key. + +So the password a provider's provisioner invents is not even the password the consumer was +given: a consumer authenticating with the mesh's password against a resource the provisioner +created with its own would simply fail. The symmetric seal is not an incomplete feature to +finish delivering a key for. It is a second, contradictory credential model bolted beside the +real one, and it cannot be made to work without building the very thing the mesh was designed +not to have. + +This is a decision and not a patch because the harness is the **provider contract**. Every +provider — the four that exist and the many a real mesh grows — is built on +`runProvisioner(resource, adapter)`. Whatever it says a provider is, they all inherit; and +changing it later is one migration per provider. It is cheaper and more honest to settle what +a provider is now. + +## Decision + +**A provider is handed the credential; it does not make one, does not seal one, and does not +hand one back.** The provisioner's only job is to make the mesh's grants true in its own +software. + +Concretely, for the sdk harness and the adapter contract: + +- The harness reconciles the **contributions the mesh delivers** to the provider's `receives` + path — the list of consumers, each with its login name (`As`), address, requested values, + and the path to its unsealed password (`Secret`). It does not read `*.grant.json` and it + does not write `*.credential`. +- For each consumer present, the harness reads the password from that consumer's `Secret` file + and calls the adapter to bring the resource into being under the given login. For each + consumer no longer present — the mesh drops it from the contributions file when its consumer + goes away — the harness calls the adapter to withdraw it. +- The adapter shrinks to the per-software half and nothing else. It is given the login, the + password, and the values, and it makes the resource exist or removes it. It generates no + password, derives no name, seals nothing, and returns no credential: + roughly `create({ as, password, values })` and `remove({ as })`, both returning nothing. +- `$MESH_SEAL_KEY`, the symmetric `seal()`/`writeSealedCredential` path, and the `*.grant.json` + / `*.credential` files are removed from the provisioning path entirely. The credential + reaches the consumer through the mesh's own asymmetric channel, which already crosses node + boundaries and holds no shared secret. + +Identity stays the mesh's to say. The login the provider creates is the name the mesh derived +and gave the consumer to present; the provider never invents a name, because a name the +consumer cannot learn is a name it cannot authenticate with. + +## Consequences + +- A provider module becomes smaller and unable to be wrong in this way: with no password to + generate and no key to seal to, the class of bug where the two ends hold different secrets + cannot be written. A provider added after this inherits the corrected contract and has no + seal to reintroduce. +- The four current providers (redis, postgres, minio, umami) each lose their `generatePassword` + + seal code and gain a `create` that takes the password it is given. Their teardown becomes + "withdraw the login named `As`". +- The symmetric `seal()`/`unseal()` primitive loses its only caller and leaves — checked, not + assumed: nothing else in the sdk or the catalogue called it, so it is removed with the + provisioner it belonged to. +- **How this is verified:** redis is assigned as a provider in the lab, the contributions and + the unsealed password the mesh would deliver are put in its `receives` path, and a client + authenticates as that consumer with the mesh's password and gets PONG — where a provider that + invented its own password answers WRONGPASS — with `$MESH_SEAL_KEY` set nowhere and no + `.credential` file written. Proven: `provider-uses-mesh-credential` is green. + +**What this does not cover — credential provisions, not data provisions.** This decision is about a +provision whose credential is a *secret the mesh mints* — a login and password (redis, postgres, +minio). A provider that instead *generates* the thing the consumer needs, and that thing is not a +secret — umami's `analytics`, where the consumer wants back a `siteId` umami assigned — does not fit, +because a contract that returns nothing has no way to hand that data back. The seal-key fault was +never umami's (it sealed no password; it returned a public id), so removing the seal does not break +it further, and it still reconciles its sites off the mesh's contributions. But delivering +provider-generated data back to a consumer is a *return path* the mesh does not have and this +decision does not build — a separate shape, left to a separate decision. +- Teardown beyond "remove the login" — data an object store leaves behind when a consumer + leaves — is named by each provider's adapter, not by the harness, and is out of scope here + except to say the contract must leave room for it. + +## References + +- [04-ISSUES/032](../04-ISSUES/032-provider-runtime-has-no-seal-key/00-report.md) — the + observation and the cross-repo trace this decision rests on. +- ADR 0047 (the module runtime) — what first ran a provider's provisioner as a delivered + process and exposed this; link to be filled when 0047 lands on the trunk. +- Control-plane mechanisms this relies on already existing: `mesh-control` — + `internal/inventory/secrets.go` (`SecretFor`, `SecretsFrom`), `internal/secrets/seal.go` + (`Make`, the two-blob asymmetric sealing), `cmd/mesh-control/plan.go` (`grantsFor`, the + `Grant.Sealed = ForProvider` delivery), `internal/catalogue/declaration.go` (the `receives` + contribution: `As`, `At`, `Values`, `Secret`). +- The path being removed: `mesh-sdk` — `src/provisioner/index.ts` (`runProvisioner`, `sealKey`, + `writeSealedCredential`) and the symmetric `src/primitives/index.ts` `seal()`/`unseal()`. diff --git a/02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md b/02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md new file mode 100644 index 0000000..4194c23 --- /dev/null +++ b/02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md @@ -0,0 +1,130 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-05 +deciders: jochen +reconstructed: false +--- + +# 49. A consumer's identity is bounded by the tightest backend that must accept it + +## Context + +The mesh says who a consumer is, once, and hands the same name to the provider (to create) and the +consumer (to present), so the two ends agree by construction rather than by two conventions (the +principle behind `ConsumerIdentity`, 04-ISSUES/023). The name is `mesh__`, cleaned to +lower-case letters, digits and underscore. + +Proving the provider contract per backend (ADR 0048) turned up 04-ISSUES/034: redis and postgres +create that name verbatim, but **minio refuses it** — an S3 access key is capped at 20 characters, +and `mesh_anchor_bucketuser` is 22. The provisioner then retries for ever, per consumer, and the +consumer holding that same too-long name could never present it either. + +Two things about the existing derivation decide most of this: + +- **The charset is already right.** `[^a-z0-9_]` is deliberately conservative, and its own comment + says it reaches "a PostgreSQL role, a MinIO access key, an LDAP uid and a Keycloak client without + quoting." That much is true. +- **The length is wrong.** `CheckIdentity` refuses names over `identityLimit = 63`, commented as + "the shortest identifier limit among the systems these names reach: PostgreSQL's". It is not the + shortest — S3's 20 is shorter — so the guard that was meant to catch exactly this lets it through, + and the failure lands at provision time as a silent retry instead of at assignment as a refusal. + +So this is a small wrong constant with a real cost attached: whatever bound we set, `mesh_` (5) plus +a node name plus `_` plus a module name has to fit inside it. + +## The options + +**A — Bound the identity by the true minimum, and refuse early.** Lower `identityLimit` to the real +shortest (20, S3's), so `CheckIdentity` refuses an over-long name *at assignment* with a clear +message, the way it already refuses over-63 names. The derivation does not change; long names are +simply rejected before anything is provisioned. +- *For:* smallest change; keeps "the mesh says the identity once, verbatim" intact; the failure + moves from a per-consumer provision-time retry to an up-front, legible refusal — which is what + `CheckIdentity` exists to do. +- *Against:* a hard budget. `mesh_` + node + `_` + module ≤ 20 means node + module ≤ 14 characters. + `anchor` + `bucketuser` (16) is already over. It pushes the constraint onto how machines and + modules are named, which is a real limitation on legible names. + +**B — Keep the readable name when it fits, compact it when it does not.** Below the bound, the name +is `mesh__` as today; over it, the mesh substitutes a deterministic short form (e.g. +`mesh_` + a truncated hash of node+module) — still one derivation, so both ends still agree. +- *For:* no naming constraint; short backends always satisfied; the common case stays legible. +- *Against:* some identities become opaque, and a provisioner tracing "whose login is this" loses + the answer for exactly the consumers that overflowed. The mesh now owns a fallback format and its + collision properties (a truncated hash is not free of collisions at 15 characters). + +**C — Let each interface declare its identifier bounds, and derive within the tightest a consumer +reaches.** `s3-bucket` states `identifier: { max: 20 }`; `postgres-database` states 63; the mesh +derives a name that fits the **minimum** bound across the providers a given consumer is granted. +- *For:* the most precise — each provision gets exactly the room it has, and a database consumer + keeps long legible names while an S3 consumer gets a short one; the constraint lives where the + fact does (on the interface). +- *Against:* the most work, and a consumer of two interfaces with different bounds must satisfy the + smaller — so its name shortens for both, reintroducing B's opacity in a narrower case. It also + means one consumer can hold **different** identities per provision, which the "said once" model + currently forbids. + +**D — Let the provider generate a backend-valid identity and hand it back (rejected).** minio mints +its own access key and returns it to the consumer. This is the data-provision return path this era +keeps meeting — but it directly contradicts 023 and ADR 0048: the identity would no longer be the +mesh's single derivation the two ends share, it would be a value one side invents and the other must +be told. Listed for completeness; not recommended. + +**E — A module (and a node) may declare a short slug; the identity is built from it.** The identity +becomes `mesh__`: where a slug is declared it is used, +otherwise the cleaned name. A slug is a deliberately short, operator-chosen identifier — `kc` for +keycloak, `wkstn` for a workstation. It is optional: short names (`anchor`, `redis`) need none. +- *For:* this is the escape hatch B wanted to be, without the opacity. The name stays legible — a + provisioner can read `mesh_wkstn_kc` and know who is asking — because a person chose it, not a + hash function. And it makes an early refusal *palatable*: if even the slug-built identity overflows, + the refusal points at the slug, a field made for exactly this, rather than at the machine's name. + Both ends still derive it from one declared thing, so they agree by construction. +- *Against:* a new optional manifest field, and someone must pick the slug — but only for names that + would otherwise overflow, and picking a short legible identifier is a better job than being handed + a hash. + +## What implementing A revealed + +A was tried first. At `identityLimit = 20`, the readable budget is `mesh_` (5) + node + `_` + module +≤ 20, i.e. **node + module ≤ 14 characters** — far tighter than it looked. The catalogue's own +existing tests use `workstation`+`keycloak` (25), which compacts to `mesh_dbbc02f8dde34d3`; common +mesh names (`home-server`, `the-build-node`, `workstation`) blow the budget with any module. So B's +compact fallback would fire for the *common* case, not the rare overflow — which inverts A+B: most +identities would be opaque hashes. A alone (hard refusal at 20) would refuse most realistic names. +This is what moved the recommendation to E: the problem is not the limit, it is that the *readable +name* is the wrong source when it is long, and a slug is a better source than either a hash or a ban. + +## Recommendation + +**E, over a per-consumer bound (start with the global minimum, 20).** Build the identity from an +optional slug, keep it when it fits, and refuse at assignment with "declare or shorten ``'s +slug" when it does not — no hash, no lost legibility, and the fix is a first-class field. Set the +bound to the true minimum (20) now; it needs no per-interface machinery to unblock S3, and a module +that consumes S3 simply declares a short slug. Graduate to **C** (per-interface bounds) later if it +turns out that non-S3 consumers are paying for S3's limit often enough to mind — E and C compose: +slugs are the mechanism, per-interface bounds refine where the ceiling sits. **B is dropped**: a +declared slug is a strictly better escape hatch than an opaque hash. **D stays rejected.** + +## Consequences (of E) + +- A module manifest gains an optional `slug`; a node may carry one too. `ConsumerIdentity` prefers + the slug over the cleaned name for each half. `identityLimit` becomes 20 (the true minimum), and + `CheckIdentity` refuses at `module add` / assignment — now with a message naming the slug to set. +- The common case stays legible; only names that overflow the budget need a slug, and what they get + is a name a person chose, not a hash. +- Existing modules/nodes whose names overflow declare a slug once — a migration cost paid as a clear + refusal with an obvious remedy, not a silent hash or a silent truncation. +- minio (04-ISSUES/034) is unblocked: an S3 consumer declares a short slug and its access key fits. +- **How it is checked:** the minio grant e2e — a consumer whose (slugged) identity fits reaches its + bucket with the credential the mesh delivered — plus unit tests that a slug is preferred, that an + un-sluggable over-long identity is refused (naming the slug), and that two consumers never collide. + +## References + +- [04-ISSUES/034](../04-ISSUES/034-mesh-login-exceeds-s3-access-key-limit/00-report.md) — the + observation. +- ADR 0048 — a provider creates the credential the mesh minted; the identity it creates it under is + the one this decision bounds. +- `mesh-control` `internal/catalogue/identity.go` — `ConsumerIdentity`, `identityUnusable`, + `identityLimit`, `CheckIdentity` — where the constant and the check live. diff --git a/02-DECISIONS/README.md b/02-DECISIONS/README.md index 64a4a6f..586deeb 100644 --- a/02-DECISIONS/README.md +++ b/02-DECISIONS/README.md @@ -107,6 +107,17 @@ python3 00-META/checks/index.py fail if stale - **0026** — [The mesh has a session of its own, and it is the node session's mechanism](0026-the-mesh-has-a-session-of-its-own.md) - **0027** — [A provision names what the consumer is coupled to, not the role it plays](0027-a-provision-names-what-the-consumer-is-coupled-to.md) - **0035** — [One implementation, several surfaces, and what that costs](0035-one-implementation-several-surfaces.md) +- **0038** — [The mesh assigns the port, and a module does not care](0038-the-mesh-assigns-the-port.md) *(proposed)* +- **0040** — [What a module is](0040-what-a-module-is.md) +- **0041** — [Events are a relationship, the lighter sibling of provisioning](0041-events-are-a-relationship.md) +- **0042** — [The shape of an event on the wire](0042-the-shape-of-an-event-on-the-wire.md) +- **0043** — [A module's broker account is scoped by what it emits and consumes](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) +- **0044** — [A public name is provisioned, not registered by hand](0044-a-public-name-is-provisioned-like-any-capability.md) +- **0045** — [A machine's firewall is the sum of what its modules listen on](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) +- **0046** — [A module's configuration is its assignment's, not its manifest's](0046-a-module-configuration-is-its-assignments-not-its-manifest.md) +- **0047** — [A module runs its code as its own process, with its own account](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) +- **0048** — [A provider creates the credential the mesh minted, and seals nothing](0048-a-provider-creates-the-credential-the-mesh-minted.md) +- **0049** — [A consumer's identity is bounded by the tightest backend that must accept it](0049-a-consumers-identity-fits-the-tightest-backend.md) ### How it is built @@ -116,6 +127,8 @@ python3 00-META/checks/index.py fail if stale - **0014** — [No workspace — each module is a standalone package consuming published dependencies](0014-no-npm-workspace.md) - **0015** — [Applications live in their own repository; the monorepo is for the mesh](0015-applications-live-in-their-own-repository.md) - **0016** — [The lab](0016-the-lab.md) +- **0037** — [Where a module lives](0037-where-a-module-lives.md) *(proposed)* +- **0039** — [What the SDK holds, and what it refuses](0039-what-the-sdk-holds-and-refuses.md) ### How it is checked diff --git a/04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md b/04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md index 7b670e1..5ff8754 100644 --- a/04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md +++ b/04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md @@ -1,9 +1,9 @@ --- -status: resolved +status: fixed opened: 2026-08-22 -located-in: [mesh-control] -fixed-by: mesh-control — a machine's filtering is computed from what it was assigned -amended-design: 03-DESIGN/01-to-be/08-connectivity.md +located-in: [mesh-control/internal/catalogue, mesh-catalog/modules/firewall] +fixed-by: the manifest refuses unknown keys, `from` is the field that scopes a port and it is rendered to nftables, and the firewall module applies it +amended-design: 0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md --- # 003 — A firewall rule's `scope:` is read by no code diff --git a/04-ISSUES/032-provider-runtime-has-no-seal-key/00-report.md b/04-ISSUES/032-provider-runtime-has-no-seal-key/00-report.md new file mode 100644 index 0000000..c0c70f0 --- /dev/null +++ b/04-ISSUES/032-provider-runtime-has-no-seal-key/00-report.md @@ -0,0 +1,135 @@ +--- +status: resolved +opened: 2026-09-04 +located-in: [mesh-sdk, mesh-catalog] +fixed-by: mesh-sdk src/provisioner rework + redis/postgres/minio/umami adapters (ADR 0048) +amended-design: 0048-a-provider-creates-the-credential-the-mesh-minted.md +--- + +# A provider's provisioner seals with a key the mesh has no way to deliver — and does not need to + +## What was observed + +Building the vertical slice for the module runtime (the module runs its own code as its own +process under its own account), a **provider** module — one that stands up a per-consumer +resource and hands back a credential — was assigned to a node and run as a broker-bound +runtime. The runtime hosts the module's provisioner (the sdk's `runProvisioner`), and the +harness opens by reading a **seal key** from `$MESH_SEAL_KEY`, failing immediately without +one. Every credential it produces for a consumer is sealed to that key with the sdk's +symmetric `seal()` (AES-256-GCM, `mesh-sdk/src/primitives/index.ts`) before being written. + +Nothing in the mesh sets `$MESH_SEAL_KEY`. It is read in exactly two places in the sdk and +set nowhere — no manifest, no control-plane code, no host code. So a provider runtime, as +delivered, aborts at start-up. The slice proved the mechanism only by setting a lab-local key +in the manifest by hand. + +## What a trace of the credential path turned up + +The seal key is not a missing delivery. **The whole symmetric-seal provisioner is orphaned, +and it duplicates — badly — a job the mesh already does.** + +- `runProvisioner` reads request files named `*.grant.json`. **Nothing writes those.** +- It writes sealed credential files named `..credential`. **Nothing reads + those** — not the host, not the control plane. The host reports applied-resource digests + upward and never ships credentials; the control plane has no reference to that filename. +- No consumer ever calls the symmetric `unseal()`. Consumers receive **plaintext**. + +Meanwhile the mesh already carries a provider→consumer credential across nodes, with **no +shared key anywhere**: + +- The control plane mints the password once (`secrets.Make`) and seals it **twice, + asymmetrically** — `ForConsumer` to the consumer node's X25519 public key, `ForProvider` to + the provider node's (`mesh-control/internal/secrets/seal.go`, `mesh-host/internal/identity/ + sealing.go`, NaCl box). +- Each host opens its own copy with its own private key on the machine; the plaintext exists + only for the length of one function call (`mesh-host/internal/apply/apply.go`, the + `${secret:name}` substitution — ADR 0024's "the host is the only thing that ever holds + both"). +- `serves` carries no credential and says so; `receives`/`bound` tell each side *where* its + sealed secret is, never the value. + +The two models also **contradict** each other. The sdk's `seal()` comment says the key is "a +per-node passphrase the host holds"; the host holds no such passphrase — it holds an X25519 +private key, and the control plane's own code refuses a shared symmetric key on principle: +"a key both ends hold is a key the mesh would have to distribute, which is this problem again +one level down" (`secrets/seal.go`). A symmetric `MESH_SEAL_KEY` shared between a provider +node and a consumer node is exactly the thing the mesh was built not to have. + +And the provisioner's model is wrong in a second way: its adapter **generates its own +password** (`generatePassword()`) and creates the resource with it — a different password from +the one the mesh mints and hands the consumer. Even with a seal key delivered, a consumer +would authenticate with the mesh's password against a resource created with the provisioner's. + +## Why it matters beyond this instance + +This is not a four-module problem. The provider contract lives in **one place** — the sdk's +`runProvisioner(resource, adapter)` harness — and every provider is built on it. Four exist +today (redis, postgres, minio, umami); a mesh of any size ends up with many. Whatever the +provisioner harness does, every present and future provider inherits, so the orphaned +symmetric seal is a fault stamped into the interface, not into four adapters. That also sets +the cost of getting it wrong: a contract N providers depend on is N migrations to change +later, which is the argument for settling it deliberately now rather than patching around it. + +As written, each provider carries a provisioner that cannot start (no key), and that, if it +did, would create resources with a password it invented — a *different* password from the one +the mesh minted and handed the consumer — and seal them for a reader that does not exist. The +rule the design states, "a consumer receives a sealed credential and unseals it," is enforced +by nothing: no consumer unseals, and no shared key exists to unseal with. + +## The mesh already does this — confirmed + +The premise the fix rests on is not a hope; it is in the control plane today. For a served +interface, `Inventory.SecretFor` mints one password per (consumer, provider) pair via +`secrets.Make`, sealing it to **both** node keys — `ForConsumer` and `ForProvider`. +`SecretsFrom(provider)` is documented as "every credential a provider node was issued, so it +can be told what to create," and `grantsFor` (plan.go) hands the provider node one `Grant` per +consumer carrying `Sealed: ForProvider`. The provider receives, at the path its `receives` +names, one `Contribution` per consumer: the login to create (`As`, derived by the mesh so both +ends agree — 04-ISSUES/023), the consumer's address (`At`) and requested `Values`, and +`Secret`, the file holding that consumer's password sealed to this provider and unsealed by +its host. Everything the provisioner needs is delivered. It reads the wrong files +(`*.grant.json`, which nothing writes) and invents a password instead of reading the one in +`Secret`. + +## The fix this points to + +A **one-place contract change in the sdk harness**, plus re-pointing today's adapters at it — +not per-provider surgery, and inherited correctly by every provider after them: + +- `runProvisioner` reconciles the mesh-delivered `receives` contributions (not `*.grant.json`): + for each consumer, create the resource under the login `As` with the password read from the + delivered `Secret` file, for its `Values`; withdraw the login when a consumer leaves the file. +- The adapter stops generating a password and stops returning a credential — it is handed the + name and the password and only makes the resource exist. Roughly `create({as, password, + values})` / `remove({as})`, no return. +- `sealKey`, `seal()`, `writeSealedCredential`, `MESH_SEAL_KEY`, and the `.credential` file + leave entirely; the consumer already receives its copy through the mesh's own channel. + +This is proposed as ADR 0048, which defines the corrected provider contract, for ratification. + +## Resolution + +ADR 0048 was accepted and implemented on the branches this issue is fixed by: + +- `mesh-sdk` `src/provisioner/index.ts` now reconciles the mesh's `receives` contributions and, + per consumer, reads the mesh-minted password from the file the host unsealed, calling the + adapter to create the resource under the mesh's login. `$MESH_SEAL_KEY`, the symmetric seal, + `writeSealedCredential`, and the `*.grant.json` / `*.credential` files are gone. The symmetric + `seal()`/`unseal()` primitive had no other caller and was removed. +- The four adapters (redis, postgres, minio, umami) were re-pointed at the new contract — + `create({ as, password, values })` / `remove({ as })`, returning nothing. minio's client gained + a secret-key argument so it sets the mesh's secret rather than generating one. +- Proven in the mesh-lab: `provider-uses-mesh-credential` is green — redis creates the consumer's + login with the password the mesh minted, a client authenticates as that consumer and gets PONG, + with no seal key set anywhere. + +Two things were carved out deliberately, neither blocking: + +- **Data provisions are a separate shape.** umami's `analytics` returns a `siteId` umami + *generates*, not a secret the mesh mints, and a contract that returns nothing cannot hand that + back. ADR 0048 is scoped to credential provisions and says so; the provider→consumer return + path for generated data is left to a separate decision. umami compiles and reconciles under the + new harness; only that return is unaddressed, and it never had the seal-key fault. +- **Teardown beyond "remove the login"** — an object store's leftover data — is each adapter's to + name (minio leaves a non-empty bucket for an operator rather than deleting a consumer's data), + not the harness's. diff --git a/04-ISSUES/033-runtime-config-change-does-not-restart/00-report.md b/04-ISSUES/033-runtime-config-change-does-not-restart/00-report.md new file mode 100644 index 0000000..b7089ab --- /dev/null +++ b/04-ISSUES/033-runtime-config-change-does-not-restart/00-report.md @@ -0,0 +1,67 @@ +--- +status: open +opened: 2026-09-04 +located-in: [] +fixed-by: +amended-design: +--- + +# Changing a module's settings does not restart its runtime — config is stale until recreated + +## What was observed + +Rolling the module runtime out to the catalogue (the runtime that serves a module's tools and +runs its events under the module's own account), each tools+events module receives its +configuration the way the design intends: a mergeable config file the module declares, into +which the assignment's settings are merged. The runtime container mounts that file and reads +it once at start-up, when it builds its API client. + +The design for settings says a config file a module owns can be changed **without editing +it** — a person states an intention, the file is regenerated, and the change takes effect. +The decision that config is the assignment's, not the manifest's, is explicitly so that +configuration can be updated *on the fly* and managed from a dashboard. + +For a runtime delivered as a **container**, that last part does not hold. When settings +change, the control plane re-renders the config file on the node — but the runtime container +is only ever recreated when its **spec** changes, and the spec is image, name, env, ports, +volumes and args. The *content* of a mounted file is not part of it. So the file on disk +updates and the process that already read it keeps the value it read at start-up. The new +configuration does not take effect until something changes the container's spec, or it is +recreated by hand. + +A **service** resource has `restart-on`, which names the resources whose change forces a +restart — exactly this problem, already solved, for units. A **container** resource has no +equivalent field, and the apply path for containers never consults the set of resources that +changed this pass. So the one kind of resource that hosts a module's runtime is the kind that +cannot say "restart me when my config changes." + +The effect is quiet, which is the worst part: setting a value appears to succeed (the file is +correct on disk), and the running tools keep answering with the old configuration, or keep +failing to load because the value that would fix them is present but unread. + +## Why it matters beyond this instance + +Every tools+events module converted to the runtime model now takes its URL and credentials +this way, so this is not one module's quirk — it is the config path for the whole catalogue. +The gap turns the headline promise of the settings design ("change it without editing it, on +the fly") into "change it, then recreate the container by hand," which is the manual step the +design existed to remove. And because the file is genuinely updated, nothing surfaces the +staleness; a dashboard that set the value would report success while the mesh kept doing the +old thing. + +Config set **before** the runtime first starts (settings, then assign, then push) does work — +the file is right when the process reads it. So the gap is specifically about *updates* to an +already-running runtime, which is precisely the case the "on the fly" promise is about. + +## Open questions + +- Should a `container` gain `restart-on`, mirroring the service field, so a module can point + it at its config resource? +- Or should the apply path recreate a container when a file it mounts changed this pass — + making mounted-file content behave like part of the spec, without a new field to declare? +- Should the config file's content (or a hash of it) fold into the container spec, so an + ordinary spec-diff already catches it? That restarts on every change with no new mechanism, + at the cost of a spec that is no longer only the container's own declaration. +- Is a restart even the right primitive for a runtime that could instead watch its config + file and rebuild its clients in place — and if so, is that each module's job or the + runtime host's? diff --git a/04-ISSUES/034-mesh-login-exceeds-s3-access-key-limit/00-report.md b/04-ISSUES/034-mesh-login-exceeds-s3-access-key-limit/00-report.md new file mode 100644 index 0000000..01eec09 --- /dev/null +++ b/04-ISSUES/034-mesh-login-exceeds-s3-access-key-limit/00-report.md @@ -0,0 +1,77 @@ +--- +status: resolved +opened: 2026-09-05 +located-in: [mesh-control, mesh-catalog] +fixed-by: ADR 0049 (a slug for the login) + a shorter minted secret (mesh-control) +amended-design: 0049-a-consumers-identity-fits-the-tightest-backend.md +--- + +# The mesh's derived login does not fit every backend's identity rules — S3 rejects it + +## What was observed + +Proving the provider/consumer contract per backend (ADR 0048), redis and postgres passed: a +consumer authenticated against the provider with the login the mesh derived and the password +the mesh minted. **minio failed**, and not on the credential — on the *name*: + +``` +mc: Unable to add a new service account. The access key is invalid. + (access key length should be between 3 and 20). +``` + +The mesh derives a consumer's login as `mesh__` — here `mesh_anchor_bucketuser`, +22 characters. That is a valid postgres role and a valid redis ACL user, so those providers +create it verbatim. S3 access keys are capped at **20 characters**, so minio refuses to create +the service account under it, and the provisioner retries forever while the consumer, holding +that same too-long access key, could never present it either. + +## Why it matters beyond this instance + +ADR 0048 says a provider creates *exactly* the login the mesh derived, so that the two ends +agree by construction — the mesh hands the same name to the provider (to create) and the +consumer (to present). That only holds if the derived name is one every provider can accept. +It is not: the mesh's `as` is a single format with no knowledge of a backend's identity rules, +and S3's are stricter than a database's. Any provider whose backend constrains identifiers more +tightly than postgres — a length cap, a charset, a required prefix — inherits this, and the +failure lands at provision time, per consumer, as an infinite retry rather than a refusal at +assignment. + +This also shows the seam is real, not cosmetic: `as` is doing two jobs — a stable per-consumer +identity the two ends must agree on, and a literal identifier a specific backend must accept — +and those are not always the same string. + +## The shape of a fix (open, not decided) + +- **Constrain the derivation** so `as` is broadly acceptable — short (≤ 20), a conservative + charset, deterministic. This keeps "the provider creates exactly what the mesh derived" true + everywhere, at the cost of a less legible name, and it is a mesh-wide identity change (every + provider that already created the longer name would see it change). +- **Let a provider map `as` to a backend-valid identifier** it derives the same way on create + and on the consumer's behalf — but the consumer is generic and cannot run minio's mapping, so + this only works if the mapped identifier is *delivered back* to the consumer. That is the + data-provision return path this era keeps meeting (umami's siteId, cloudflare's record) and + does not yet have. +- **Declare the constraint on the interface** (`s3-bucket` states its identifier bounds) and + have the mesh derive within them — the most honest, the most work. + +## Open questions + +- Is `as` meant to be human-legible, or is a short opaque token acceptable — i.e., can the + derivation simply be shortened without anyone minding? +- Do redis/postgres actually want the long name, or did it only survive because they are + permissive? If nothing needs it long, the cheap fix is to cap it. +- Does this fold into the same decision as the data-provision return path, or is it separate? + +## Resolution + +Accepted **ADR 0049** (option E): a module declares an optional short `slug`, and the mesh derives +`mesh__`, bounded by the tightest backend (an S3 access key's 20) and refused at +assignment — naming the slug as the remedy — when it still would not fit. The minio grant e2e proved +it: `bucketuser` declares `slug: bkt`, so its access key `mesh_anchor_bkt` (15) is accepted where +`mesh_anchor_bucketuser` (22) was refused. + +Proving that surfaced a **second S3 length constraint on the same credential** — the secret. The +mesh minted a 43-character password (32 random bytes, base64url), and an S3 secret key is 8–40. Fixed +in `mesh-control` `internal/secrets/seal.go` by minting 30 bytes → exactly 40 characters (240 bits, +ample), which fits S3 and every other backend. Both halves of an S3 credential — the access key +(login) and the secret key (password) — now fit the tightest backend, by the same rule. diff --git a/04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md b/04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md new file mode 100644 index 0000000..8c8bd49 --- /dev/null +++ b/04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md @@ -0,0 +1,48 @@ +--- +status: open +opened: 2026-09-02 +located-in: [] +fixed-by: +amended-design: +--- + +# 035 — Reconciling a seed file wipes what grew in it + +## The symptom, as observed + +Found by review of the catalogue examples (2026-09-02), not by an outage — the outage is the +part the design permits to be silent. + +The cache module declares its access-control file as an ordinary file resource with fixed, +empty content. The program that consumes the file requires it to exist at startup, which is +why the manifest declares it at all. But the same file is the one the provisioner writes +consumer users into, and the one the running program persists ACL changes back to. + +A declaration is complete for what the host owns, and the host reconciles what is declared +([ADR 0010](../../02-DECISIONS/0010-delivery.md)). +So every apply that revisits this resource restores the declared content — empty — behind the +running program. Every consumer credential granted since the last apply is removed, the apply +reports success, and nothing anywhere says a grant vanished. + +## Why it matters beyond the instance + +The manifest needed *the file to exist before first start*, and the only vocabulary available +was *the file has this content, forever*. Those are different intentions, and the gap between +them is generic: any resource that a module seeds and something else then legitimately mutates +— an ACL file, a bootstrap configuration a program rewrites, an htpasswd a provisioner appends +to — has the same two owners and the same silent loss on reconcile. + +It is also the mirror image of the boundary ADR 0010 draws so carefully on the *removal* side: +the host never removes what it did not create, but it happily overwrites what it *did* create, +even when what grew inside since is somebody else's work the mesh asked for. + +## Open questions + +- Is the missing thing a create-once file semantic ("present with this content if absent, + untouched otherwise"), or is the real fault that two owners share one file — and the + provisioner, not the declaration, should own it entirely, with first-start ordering solved + some other way? +- ADR 0010 treats every added resource type as a security artefact. Does a create-once + semantic widen what a compromised control plane can express, or narrow it? +- Are there other seeded-then-mutated files already in the catalogue that this failure is + waiting inside? diff --git a/04-ISSUES/036-six-modules-own-what-they-must-share/00-report.md b/04-ISSUES/036-six-modules-own-what-they-must-share/00-report.md new file mode 100644 index 0000000..7513890 --- /dev/null +++ b/04-ISSUES/036-six-modules-own-what-they-must-share/00-report.md @@ -0,0 +1,51 @@ +--- +status: open +opened: 2026-09-02 +located-in: [] +fixed-by: +amended-design: +--- + +# 036 — Six modules own what they must share + +## The symptom, as observed + +Found by review of the catalogue examples (2026-09-02). The media stack is several modules — +a library server, the acquisition managers, a download client and their satellites — and each +of them declares the same library and download directories as its own resources. + +The resolver refuses two modules that declare one path on one node, with no exemption for +identical content and no merge. That rule is right in general: two owners of one path is the +class of fault this repository keeps recording. But sharing those directories on one machine +is the entire point of this stack — the download client and the managers must see the same +downloads, the library server must see the same libraries. So the set, as written, refuses +its own only sensible assignment. + +No test co-resolves any two of them, which is why the manifests pass today. The first machine +to be assigned the stack together is where the refusal would have surfaced. + +## Why it matters beyond the instance + +The manifests can express *a directory I own* and nothing else, so a directory that is the +shared workspace of several modules was written six times as six private ones. The intention +— several modules, one filesystem contract between them — has no vocabulary, and this is not +a media-stack peculiarity: any pipeline of modules handing files to each other on one machine +(an ingest directory, a spool, a drop folder) hits the same wall. + +It is also a fork in the design the catalogue has otherwise avoided: the fix could be a new +owning module the others depend on, a shared-resource concept in the manifest, or a statement +that co-located file handoff is not a thing the mesh supports and these modules are one +module. Each answer changes what a module *is*, which is why this is an issue and not a patch. + +## Open questions + +- Is the unit wrong — is a stack that must share a filesystem one module with several + containers, the way the mail module already is? +- If it stays several modules: does one of them own the directories and the rest require + them, and is *requiring a directory from a neighbour* a provision, a claim, or a third + thing? +- The duplicate-path rule protects against genuinely rivalrous owners. Whatever expresses + sharing must not weaken it for the cases where refusal is the right answer — what + distinguishes the two, machine-checkably? +- The mesh's own rule is that a rule states how it is checked: whichever shape is chosen, + what test co-resolves the stack so this class of refusal is caught before a machine is? diff --git a/04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md b/04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md new file mode 100644 index 0000000..2c26c1f --- /dev/null +++ b/04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md @@ -0,0 +1,57 @@ +--- +status: open +opened: 2026-09-05 +located-in: [] +fixed-by: +amended-design: +--- + +# 037 — A module cannot run its own code at a lifecycle phase + +## The symptom, as observed + +Found while converting the catalogue (2026-09-05), across several modules at once. A module can +declare *things that exist* — a directory, a file with fixed content, a network, a container — but +it cannot declare *a step that runs* at a defined point in its own lifecycle. Three converted +modules need exactly that and have nowhere to put it: + +- **mosquitto.** Its Dynamic Security plugin will not start unless `dynamic-security.json` already + contains an admin client *before the broker's first start* — the broker loads the plugin at + boot. Seeding it is a run-once step that must happen after the file resource exists and before + the container starts. The vocabulary has no "before first start." +- **The database providers (postgres/mongodb/mssql).** First-boot seeding works today only because + the *image* happens to do it from an env var. Anything the mesh itself must run once against the + server — a schema migration, an extension enable, a health gate before the module is announced + ready — has no home. +- The seed-then-mutate family already recorded in [035](../035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md) + is the same shape seen from the *content* side; this is it seen from the *timing* side. + +## Why it matters beyond the instance + +This is not a defect in a module — it is a **capability the module system does not yet offer.** A +real class of modules needs to run their own code at points in the build/install/run lifecycle: +seed-before-start, migrate, post-start health-gate, pre-remove drain. The declarative resource +model deliberately describes *state*, not *steps*, and that is right for what it covers; the gap is +that some modules genuinely have a step. + +**Prior art, and its warning.** An earlier mesh had exactly this as a feature: event-driven +**hooks** that ran custom code at phases of the build/publish/deploy pipeline. It was powerful and +it was **complex to set up and flaky** — which is the real content of this record. The need is not +in question; the cost of the obvious answer is. Whatever shape this takes must not reproduce that +fragility, or it will be worse than the gap. + +## Open questions + +- Is the right unit narrow — a **run-once / init resource** ("run this once, here, in the + lifecycle") — or general — a **per-phase lifecycle hook** on a module, and if so which phases + (build / publish / install / pre-start / post-start / pre-remove)? +- Where does a hook's code run — in the module's own runtime container under its scoped account + (ADR 0043/0047), so it inherits the same isolation as its tools and events? Or is some of it the + host's, before a container exists? +- How is a step made **idempotent and reconcilable** so a re-apply does not re-run it + destructively — the same discipline the resource model gets for free and a step does not? +- What is the smallest version that unblocks the three modules above without rebuilding the old + flaky hook engine? Is "seed-before-first-start" alone enough for now, with the general case + deferred? +- A rule states how it is checked: whatever shape is chosen, what lab scenario proves a hook runs + exactly once, at the right phase, and converges on re-apply?