Re-home this session's new ADRs (0039-0049) and issues (032-037) onto the consolidated scheme; flip issue 003; port repos.md sdk line + feature-branches playbook (07); regenerate index

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
This commit is contained in:
2026-09-05 12:24:07 +02:00
parent 546caa31c4
commit e269f9a185
22 changed files with 1643 additions and 5 deletions
@@ -1,9 +1,9 @@
---
status: resolved
status: fixed
opened: 2026-08-22
located-in: [mesh-control]
fixed-by: mesh-control — a machine's filtering is computed from what it was assigned
amended-design: 03-DESIGN/01-to-be/08-connectivity.md
located-in: [mesh-control/internal/catalogue, mesh-catalog/modules/firewall]
fixed-by: the manifest refuses unknown keys, `from` is the field that scopes a port and it is rendered to nftables, and the firewall module applies it
amended-design: 0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md
---
# 003 — A firewall rule's `scope:` is read by no code
@@ -0,0 +1,135 @@
---
status: resolved
opened: 2026-09-04
located-in: [mesh-sdk, mesh-catalog]
fixed-by: mesh-sdk src/provisioner rework + redis/postgres/minio/umami adapters (ADR 0048)
amended-design: 0048-a-provider-creates-the-credential-the-mesh-minted.md
---
# A provider's provisioner seals with a key the mesh has no way to deliver — and does not need to
## What was observed
Building the vertical slice for the module runtime (the module runs its own code as its own
process under its own account), a **provider** module — one that stands up a per-consumer
resource and hands back a credential — was assigned to a node and run as a broker-bound
runtime. The runtime hosts the module's provisioner (the sdk's `runProvisioner`), and the
harness opens by reading a **seal key** from `$MESH_SEAL_KEY`, failing immediately without
one. Every credential it produces for a consumer is sealed to that key with the sdk's
symmetric `seal()` (AES-256-GCM, `mesh-sdk/src/primitives/index.ts`) before being written.
Nothing in the mesh sets `$MESH_SEAL_KEY`. It is read in exactly two places in the sdk and
set nowhere — no manifest, no control-plane code, no host code. So a provider runtime, as
delivered, aborts at start-up. The slice proved the mechanism only by setting a lab-local key
in the manifest by hand.
## What a trace of the credential path turned up
The seal key is not a missing delivery. **The whole symmetric-seal provisioner is orphaned,
and it duplicates — badly — a job the mesh already does.**
- `runProvisioner` reads request files named `*.grant.json`. **Nothing writes those.**
- It writes sealed credential files named `<consumer>.<resource>.credential`. **Nothing reads
those** — not the host, not the control plane. The host reports applied-resource digests
upward and never ships credentials; the control plane has no reference to that filename.
- No consumer ever calls the symmetric `unseal()`. Consumers receive **plaintext**.
Meanwhile the mesh already carries a provider→consumer credential across nodes, with **no
shared key anywhere**:
- The control plane mints the password once (`secrets.Make`) and seals it **twice,
asymmetrically** — `ForConsumer` to the consumer node's X25519 public key, `ForProvider` to
the provider node's (`mesh-control/internal/secrets/seal.go`, `mesh-host/internal/identity/
sealing.go`, NaCl box).
- Each host opens its own copy with its own private key on the machine; the plaintext exists
only for the length of one function call (`mesh-host/internal/apply/apply.go`, the
`${secret:name}` substitution — ADR 0024's "the host is the only thing that ever holds
both").
- `serves` carries no credential and says so; `receives`/`bound` tell each side *where* its
sealed secret is, never the value.
The two models also **contradict** each other. The sdk's `seal()` comment says the key is "a
per-node passphrase the host holds"; the host holds no such passphrase — it holds an X25519
private key, and the control plane's own code refuses a shared symmetric key on principle:
"a key both ends hold is a key the mesh would have to distribute, which is this problem again
one level down" (`secrets/seal.go`). A symmetric `MESH_SEAL_KEY` shared between a provider
node and a consumer node is exactly the thing the mesh was built not to have.
And the provisioner's model is wrong in a second way: its adapter **generates its own
password** (`generatePassword()`) and creates the resource with it — a different password from
the one the mesh mints and hands the consumer. Even with a seal key delivered, a consumer
would authenticate with the mesh's password against a resource created with the provisioner's.
## Why it matters beyond this instance
This is not a four-module problem. The provider contract lives in **one place** — the sdk's
`runProvisioner(resource, adapter)` harness — and every provider is built on it. Four exist
today (redis, postgres, minio, umami); a mesh of any size ends up with many. Whatever the
provisioner harness does, every present and future provider inherits, so the orphaned
symmetric seal is a fault stamped into the interface, not into four adapters. That also sets
the cost of getting it wrong: a contract N providers depend on is N migrations to change
later, which is the argument for settling it deliberately now rather than patching around it.
As written, each provider carries a provisioner that cannot start (no key), and that, if it
did, would create resources with a password it invented — a *different* password from the one
the mesh minted and handed the consumer — and seal them for a reader that does not exist. The
rule the design states, "a consumer receives a sealed credential and unseals it," is enforced
by nothing: no consumer unseals, and no shared key exists to unseal with.
## The mesh already does this — confirmed
The premise the fix rests on is not a hope; it is in the control plane today. For a served
interface, `Inventory.SecretFor` mints one password per (consumer, provider) pair via
`secrets.Make`, sealing it to **both** node keys — `ForConsumer` and `ForProvider`.
`SecretsFrom(provider)` is documented as "every credential a provider node was issued, so it
can be told what to create," and `grantsFor` (plan.go) hands the provider node one `Grant` per
consumer carrying `Sealed: ForProvider`. The provider receives, at the path its `receives`
names, one `Contribution` per consumer: the login to create (`As`, derived by the mesh so both
ends agree — 04-ISSUES/023), the consumer's address (`At`) and requested `Values`, and
`Secret`, the file holding that consumer's password sealed to this provider and unsealed by
its host. Everything the provisioner needs is delivered. It reads the wrong files
(`*.grant.json`, which nothing writes) and invents a password instead of reading the one in
`Secret`.
## The fix this points to
A **one-place contract change in the sdk harness**, plus re-pointing today's adapters at it —
not per-provider surgery, and inherited correctly by every provider after them:
- `runProvisioner` reconciles the mesh-delivered `receives` contributions (not `*.grant.json`):
for each consumer, create the resource under the login `As` with the password read from the
delivered `Secret` file, for its `Values`; withdraw the login when a consumer leaves the file.
- The adapter stops generating a password and stops returning a credential — it is handed the
name and the password and only makes the resource exist. Roughly `create({as, password,
values})` / `remove({as})`, no return.
- `sealKey`, `seal()`, `writeSealedCredential`, `MESH_SEAL_KEY`, and the `.credential` file
leave entirely; the consumer already receives its copy through the mesh's own channel.
This is proposed as ADR 0048, which defines the corrected provider contract, for ratification.
## Resolution
ADR 0048 was accepted and implemented on the branches this issue is fixed by:
- `mesh-sdk` `src/provisioner/index.ts` now reconciles the mesh's `receives` contributions and,
per consumer, reads the mesh-minted password from the file the host unsealed, calling the
adapter to create the resource under the mesh's login. `$MESH_SEAL_KEY`, the symmetric seal,
`writeSealedCredential`, and the `*.grant.json` / `*.credential` files are gone. The symmetric
`seal()`/`unseal()` primitive had no other caller and was removed.
- The four adapters (redis, postgres, minio, umami) were re-pointed at the new contract —
`create({ as, password, values })` / `remove({ as })`, returning nothing. minio's client gained
a secret-key argument so it sets the mesh's secret rather than generating one.
- Proven in the mesh-lab: `provider-uses-mesh-credential` is green — redis creates the consumer's
login with the password the mesh minted, a client authenticates as that consumer and gets PONG,
with no seal key set anywhere.
Two things were carved out deliberately, neither blocking:
- **Data provisions are a separate shape.** umami's `analytics` returns a `siteId` umami
*generates*, not a secret the mesh mints, and a contract that returns nothing cannot hand that
back. ADR 0048 is scoped to credential provisions and says so; the provider→consumer return
path for generated data is left to a separate decision. umami compiles and reconciles under the
new harness; only that return is unaddressed, and it never had the seal-key fault.
- **Teardown beyond "remove the login"** — an object store's leftover data — is each adapter's to
name (minio leaves a non-empty bucket for an operator rather than deleting a consumer's data),
not the harness's.
@@ -0,0 +1,67 @@
---
status: open
opened: 2026-09-04
located-in: []
fixed-by:
amended-design:
---
# Changing a module's settings does not restart its runtime — config is stale until recreated
## What was observed
Rolling the module runtime out to the catalogue (the runtime that serves a module's tools and
runs its events under the module's own account), each tools+events module receives its
configuration the way the design intends: a mergeable config file the module declares, into
which the assignment's settings are merged. The runtime container mounts that file and reads
it once at start-up, when it builds its API client.
The design for settings says a config file a module owns can be changed **without editing
it** — a person states an intention, the file is regenerated, and the change takes effect.
The decision that config is the assignment's, not the manifest's, is explicitly so that
configuration can be updated *on the fly* and managed from a dashboard.
For a runtime delivered as a **container**, that last part does not hold. When settings
change, the control plane re-renders the config file on the node — but the runtime container
is only ever recreated when its **spec** changes, and the spec is image, name, env, ports,
volumes and args. The *content* of a mounted file is not part of it. So the file on disk
updates and the process that already read it keeps the value it read at start-up. The new
configuration does not take effect until something changes the container's spec, or it is
recreated by hand.
A **service** resource has `restart-on`, which names the resources whose change forces a
restart — exactly this problem, already solved, for units. A **container** resource has no
equivalent field, and the apply path for containers never consults the set of resources that
changed this pass. So the one kind of resource that hosts a module's runtime is the kind that
cannot say "restart me when my config changes."
The effect is quiet, which is the worst part: setting a value appears to succeed (the file is
correct on disk), and the running tools keep answering with the old configuration, or keep
failing to load because the value that would fix them is present but unread.
## Why it matters beyond this instance
Every tools+events module converted to the runtime model now takes its URL and credentials
this way, so this is not one module's quirk — it is the config path for the whole catalogue.
The gap turns the headline promise of the settings design ("change it without editing it, on
the fly") into "change it, then recreate the container by hand," which is the manual step the
design existed to remove. And because the file is genuinely updated, nothing surfaces the
staleness; a dashboard that set the value would report success while the mesh kept doing the
old thing.
Config set **before** the runtime first starts (settings, then assign, then push) does work —
the file is right when the process reads it. So the gap is specifically about *updates* to an
already-running runtime, which is precisely the case the "on the fly" promise is about.
## Open questions
- Should a `container` gain `restart-on`, mirroring the service field, so a module can point
it at its config resource?
- Or should the apply path recreate a container when a file it mounts changed this pass —
making mounted-file content behave like part of the spec, without a new field to declare?
- Should the config file's content (or a hash of it) fold into the container spec, so an
ordinary spec-diff already catches it? That restarts on every change with no new mechanism,
at the cost of a spec that is no longer only the container's own declaration.
- Is a restart even the right primitive for a runtime that could instead watch its config
file and rebuild its clients in place — and if so, is that each module's job or the
runtime host's?
@@ -0,0 +1,77 @@
---
status: resolved
opened: 2026-09-05
located-in: [mesh-control, mesh-catalog]
fixed-by: ADR 0049 (a slug for the login) + a shorter minted secret (mesh-control)
amended-design: 0049-a-consumers-identity-fits-the-tightest-backend.md
---
# The mesh's derived login does not fit every backend's identity rules — S3 rejects it
## What was observed
Proving the provider/consumer contract per backend (ADR 0048), redis and postgres passed: a
consumer authenticated against the provider with the login the mesh derived and the password
the mesh minted. **minio failed**, and not on the credential — on the *name*:
```
mc: <ERROR> Unable to add a new service account. The access key is invalid.
(access key length should be between 3 and 20).
```
The mesh derives a consumer's login as `mesh_<node>_<module>` — here `mesh_anchor_bucketuser`,
22 characters. That is a valid postgres role and a valid redis ACL user, so those providers
create it verbatim. S3 access keys are capped at **20 characters**, so minio refuses to create
the service account under it, and the provisioner retries forever while the consumer, holding
that same too-long access key, could never present it either.
## Why it matters beyond this instance
ADR 0048 says a provider creates *exactly* the login the mesh derived, so that the two ends
agree by construction — the mesh hands the same name to the provider (to create) and the
consumer (to present). That only holds if the derived name is one every provider can accept.
It is not: the mesh's `as` is a single format with no knowledge of a backend's identity rules,
and S3's are stricter than a database's. Any provider whose backend constrains identifiers more
tightly than postgres — a length cap, a charset, a required prefix — inherits this, and the
failure lands at provision time, per consumer, as an infinite retry rather than a refusal at
assignment.
This also shows the seam is real, not cosmetic: `as` is doing two jobs — a stable per-consumer
identity the two ends must agree on, and a literal identifier a specific backend must accept —
and those are not always the same string.
## The shape of a fix (open, not decided)
- **Constrain the derivation** so `as` is broadly acceptable — short (≤ 20), a conservative
charset, deterministic. This keeps "the provider creates exactly what the mesh derived" true
everywhere, at the cost of a less legible name, and it is a mesh-wide identity change (every
provider that already created the longer name would see it change).
- **Let a provider map `as` to a backend-valid identifier** it derives the same way on create
and on the consumer's behalf — but the consumer is generic and cannot run minio's mapping, so
this only works if the mapped identifier is *delivered back* to the consumer. That is the
data-provision return path this era keeps meeting (umami's siteId, cloudflare's record) and
does not yet have.
- **Declare the constraint on the interface** (`s3-bucket` states its identifier bounds) and
have the mesh derive within them — the most honest, the most work.
## Open questions
- Is `as` meant to be human-legible, or is a short opaque token acceptable — i.e., can the
derivation simply be shortened without anyone minding?
- Do redis/postgres actually want the long name, or did it only survive because they are
permissive? If nothing needs it long, the cheap fix is to cap it.
- Does this fold into the same decision as the data-provision return path, or is it separate?
## Resolution
Accepted **ADR 0049** (option E): a module declares an optional short `slug`, and the mesh derives
`mesh_<node>_<slug|name>`, bounded by the tightest backend (an S3 access key's 20) and refused at
assignment — naming the slug as the remedy — when it still would not fit. The minio grant e2e proved
it: `bucketuser` declares `slug: bkt`, so its access key `mesh_anchor_bkt` (15) is accepted where
`mesh_anchor_bucketuser` (22) was refused.
Proving that surfaced a **second S3 length constraint on the same credential** — the secret. The
mesh minted a 43-character password (32 random bytes, base64url), and an S3 secret key is 8–40. Fixed
in `mesh-control` `internal/secrets/seal.go` by minting 30 bytes → exactly 40 characters (240 bits,
ample), which fits S3 and every other backend. Both halves of an S3 credential — the access key
(login) and the secret key (password) — now fit the tightest backend, by the same rule.
@@ -0,0 +1,48 @@
---
status: open
opened: 2026-09-02
located-in: []
fixed-by:
amended-design:
---
# 035 — Reconciling a seed file wipes what grew in it
## The symptom, as observed
Found by review of the catalogue examples (2026-09-02), not by an outage — the outage is the
part the design permits to be silent.
The cache module declares its access-control file as an ordinary file resource with fixed,
empty content. The program that consumes the file requires it to exist at startup, which is
why the manifest declares it at all. But the same file is the one the provisioner writes
consumer users into, and the one the running program persists ACL changes back to.
A declaration is complete for what the host owns, and the host reconciles what is declared
([ADR 0010](../../02-DECISIONS/0010-delivery.md)).
So every apply that revisits this resource restores the declared content — empty — behind the
running program. Every consumer credential granted since the last apply is removed, the apply
reports success, and nothing anywhere says a grant vanished.
## Why it matters beyond the instance
The manifest needed *the file to exist before first start*, and the only vocabulary available
was *the file has this content, forever*. Those are different intentions, and the gap between
them is generic: any resource that a module seeds and something else then legitimately mutates
— an ACL file, a bootstrap configuration a program rewrites, an htpasswd a provisioner appends
to — has the same two owners and the same silent loss on reconcile.
It is also the mirror image of the boundary ADR 0010 draws so carefully on the *removal* side:
the host never removes what it did not create, but it happily overwrites what it *did* create,
even when what grew inside since is somebody else's work the mesh asked for.
## Open questions
- Is the missing thing a create-once file semantic ("present with this content if absent,
untouched otherwise"), or is the real fault that two owners share one file — and the
provisioner, not the declaration, should own it entirely, with first-start ordering solved
some other way?
- ADR 0010 treats every added resource type as a security artefact. Does a create-once
semantic widen what a compromised control plane can express, or narrow it?
- Are there other seeded-then-mutated files already in the catalogue that this failure is
waiting inside?
@@ -0,0 +1,51 @@
---
status: open
opened: 2026-09-02
located-in: []
fixed-by:
amended-design:
---
# 036 — Six modules own what they must share
## The symptom, as observed
Found by review of the catalogue examples (2026-09-02). The media stack is several modules —
a library server, the acquisition managers, a download client and their satellites — and each
of them declares the same library and download directories as its own resources.
The resolver refuses two modules that declare one path on one node, with no exemption for
identical content and no merge. That rule is right in general: two owners of one path is the
class of fault this repository keeps recording. But sharing those directories on one machine
is the entire point of this stack — the download client and the managers must see the same
downloads, the library server must see the same libraries. So the set, as written, refuses
its own only sensible assignment.
No test co-resolves any two of them, which is why the manifests pass today. The first machine
to be assigned the stack together is where the refusal would have surfaced.
## Why it matters beyond the instance
The manifests can express *a directory I own* and nothing else, so a directory that is the
shared workspace of several modules was written six times as six private ones. The intention
— several modules, one filesystem contract between them — has no vocabulary, and this is not
a media-stack peculiarity: any pipeline of modules handing files to each other on one machine
(an ingest directory, a spool, a drop folder) hits the same wall.
It is also a fork in the design the catalogue has otherwise avoided: the fix could be a new
owning module the others depend on, a shared-resource concept in the manifest, or a statement
that co-located file handoff is not a thing the mesh supports and these modules are one
module. Each answer changes what a module *is*, which is why this is an issue and not a patch.
## Open questions
- Is the unit wrong — is a stack that must share a filesystem one module with several
containers, the way the mail module already is?
- If it stays several modules: does one of them own the directories and the rest require
them, and is *requiring a directory from a neighbour* a provision, a claim, or a third
thing?
- The duplicate-path rule protects against genuinely rivalrous owners. Whatever expresses
sharing must not weaken it for the cases where refusal is the right answer — what
distinguishes the two, machine-checkably?
- The mesh's own rule is that a rule states how it is checked: whichever shape is chosen,
what test co-resolves the stack so this class of refusal is caught before a machine is?
@@ -0,0 +1,57 @@
---
status: open
opened: 2026-09-05
located-in: []
fixed-by:
amended-design:
---
# 037 — A module cannot run its own code at a lifecycle phase
## The symptom, as observed
Found while converting the catalogue (2026-09-05), across several modules at once. A module can
declare *things that exist* — a directory, a file with fixed content, a network, a container — but
it cannot declare *a step that runs* at a defined point in its own lifecycle. Three converted
modules need exactly that and have nowhere to put it:
- **mosquitto.** Its Dynamic Security plugin will not start unless `dynamic-security.json` already
contains an admin client *before the broker's first start* — the broker loads the plugin at
boot. Seeding it is a run-once step that must happen after the file resource exists and before
the container starts. The vocabulary has no "before first start."
- **The database providers (postgres/mongodb/mssql).** First-boot seeding works today only because
the *image* happens to do it from an env var. Anything the mesh itself must run once against the
server — a schema migration, an extension enable, a health gate before the module is announced
ready — has no home.
- The seed-then-mutate family already recorded in [035](../035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md)
is the same shape seen from the *content* side; this is it seen from the *timing* side.
## Why it matters beyond the instance
This is not a defect in a module — it is a **capability the module system does not yet offer.** A
real class of modules needs to run their own code at points in the build/install/run lifecycle:
seed-before-start, migrate, post-start health-gate, pre-remove drain. The declarative resource
model deliberately describes *state*, not *steps*, and that is right for what it covers; the gap is
that some modules genuinely have a step.
**Prior art, and its warning.** An earlier mesh had exactly this as a feature: event-driven
**hooks** that ran custom code at phases of the build/publish/deploy pipeline. It was powerful and
it was **complex to set up and flaky** — which is the real content of this record. The need is not
in question; the cost of the obvious answer is. Whatever shape this takes must not reproduce that
fragility, or it will be worse than the gap.
## Open questions
- Is the right unit narrow — a **run-once / init resource** ("run this once, here, in the
lifecycle") — or general — a **per-phase lifecycle hook** on a module, and if so which phases
(build / publish / install / pre-start / post-start / pre-remove)?
- Where does a hook's code run — in the module's own runtime container under its scoped account
(ADR 0043/0047), so it inherits the same isolation as its tools and events? Or is some of it the
host's, before a container exists?
- How is a step made **idempotent and reconcilable** so a re-apply does not re-run it
destructively — the same discipline the resource model gets for free and a step does not?
- What is the smallest version that unblocks the three modules above without rebuilding the old
flaky hook engine? Is "seed-before-first-start" alone enough for now, with the general case
deferred?
- A rule states how it is checked: whatever shape is chosen, what lab scenario proves a hook runs
exactly once, at the right phase, and converges on re-apply?