From 1ce9ffe79f91f4079d93ca2012bf1f35d311476b Mon Sep 17 00:00:00 2001 From: jochen Date: Sun, 20 Sep 2026 20:29:18 +0200 Subject: [PATCH] =?UTF-8?q?Issue=20067=20=E2=80=94=20a=20provision=20canno?= =?UTF-8?q?t=20name=20which=20provider=20serves=20it?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The mesh models provisions as mesh-scoped (one provider of a kind, a single mesh-store). But node-specific services delivered to the mesh was the plan from the start: both nodes already run their own postgres, SQL server, redis and object store, and identity — currently single — already serves apps on a second node. The model cannot express which provider serves a consumer, so it collapses a deliberately per-node fleet to one. Provider scoping is a whole-mesh decision across postgres/s3-bucket/oidc, not an SSO patch. Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx --- .../00-report.md | 78 +++++++++++++++++++ 1 file changed, 78 insertions(+) create mode 100644 04-ISSUES/067-a-provision-cannot-name-which-provider-serves-it/00-report.md diff --git a/04-ISSUES/067-a-provision-cannot-name-which-provider-serves-it/00-report.md b/04-ISSUES/067-a-provision-cannot-name-which-provider-serves-it/00-report.md new file mode 100644 index 0000000..20a7bcd --- /dev/null +++ b/04-ISSUES/067-a-provision-cannot-name-which-provider-serves-it/00-report.md @@ -0,0 +1,78 @@ +--- +status: open +opened: 2026-09-20 +located-in: [] +fixed-by: +amended-design: +--- + +# A provision cannot name which provider serves it + +## Symptom, as observed + +The mesh models every provision — `postgres-database`, `s3-bucket`, `amqp`, a +future `oidc` — as **mesh-scoped**: there is one provider of a given kind for the +whole mesh, and a consumer that requires the provision is bound to *that* one. +`scope: "mesh"` is written into the provision definitions, and the adopted store +is a single `mesh-store`. + +The mesh being migrated onto is not shaped that way, and never was meant to be. +**Node-specific services delivered to the mesh was the plan from the start.** Each +node already runs its own provider of the same kinds: + +- Both control-capable nodes run their **own general-purpose postgres server** + (the same image, one per node), serving that node's own applications. +- Each node runs its **own** SQL server, its **own** redis, its **own** object + store — infrastructure is per node, by design, not a single mesh-wide instance. +- Identity is the *only* provision that is currently single (one realm on one + node), and even that already serves applications hosted on a **second** node. + +So the multi-provider reality is not a future edge case that appears "the day a +second provider is added" — it is the founding topology, true today, on every +provision kind. What is missing is any way to **say it**. A consumer requires +`postgres-database`; it cannot require *this node's* postgres rather than *that +node's*. It requires `oidc`; it cannot name which node's identity provider. The +mesh model collapses a deliberately per-node fleet down to one mesh-scoped +provider, and a consumer has no field in which to choose. + +## Why it matters beyond this instance + +- **The model regressed an intended topology, it did not merely miss a corner + case.** "Node-specific services delivered to the mesh" is the design; `scope: + "mesh"` with a single `mesh-store` expresses the opposite. This is a gap between + a stated intent and what the manifests can represent, which is exactly what the + issues process is for. +- **It is not an identity special case.** The missing concept — *a provision has a + provider, providers are per-node, and a consumer names the provider* — is the + same for databases, object stores, brokers and identity. A fix aimed only at SSO + would leave the same wall standing behind postgres and minio, both of which are + *already* multi-provider on the live mesh. +- **The single-provider assumption is silent.** Nothing rejects a second provider + of a mesh-scoped provision; the model just cannot address it, so a consumer binds + to whichever one is "the" provider — by accident of there being one, or by a race + when there are two. A rule enforced by nothing ("there is one provider per + provision") reads as true until the second node's provider makes it false, with + no diagnostic at the seam. +- **It blocks the migration concretely.** The mesh already has applications on one + node depending on another node's provider (identity today; databases the moment + an app is assigned to a node whose local postgres is not "the" mesh store). + Modelled as mesh-scoped, that topology is expressible only by accident. To carry + it deliberately the provision must be able to name its provider. + +## Open questions + +- Where does the provider name live — on the provision definition (`scope: "node"` + with a provider identity), on the requirement in the consumer's manifest, or + supplied only at assignment time so the same module can be bound to different + providers on different assignments? +- Is "mesh-scoped" still a legitimate scope for some provisions (a single mesh CA, + say), or does every provision become node-scoped, with a single instance + expressed as "there happens to be one"? +- What is the default when a consumer names no provider — bind to the node the + consumer is assigned to (co-located provider), require the name always, or fall + back to a mesh-wide default provider where one is declared? +- How does a provider's identity survive being moved between nodes, so a consumer's + recorded choice does not silently rebind when the provider relocates? +- Does this interact with secret rotation (the issue-scope of `rotate`) — must + rotation address a specific provider's credential holders rather than "the + provision's"?