Issue 067 — a provision cannot name which provider serves it
The mesh models provisions as mesh-scoped (one provider of a kind, a single mesh-store). But node-specific services delivered to the mesh was the plan from the start: both nodes already run their own postgres, SQL server, redis and object store, and identity — currently single — already serves apps on a second node. The model cannot express which provider serves a consumer, so it collapses a deliberately per-node fleet to one. Provider scoping is a whole-mesh decision across postgres/s3-bucket/oidc, not an SSO patch. Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
This commit is contained in:
@@ -0,0 +1,78 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-20
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# A provision cannot name which provider serves it
|
||||
|
||||
## Symptom, as observed
|
||||
|
||||
The mesh models every provision — `postgres-database`, `s3-bucket`, `amqp`, a
|
||||
future `oidc` — as **mesh-scoped**: there is one provider of a given kind for the
|
||||
whole mesh, and a consumer that requires the provision is bound to *that* one.
|
||||
`scope: "mesh"` is written into the provision definitions, and the adopted store
|
||||
is a single `mesh-store`.
|
||||
|
||||
The mesh being migrated onto is not shaped that way, and never was meant to be.
|
||||
**Node-specific services delivered to the mesh was the plan from the start.** Each
|
||||
node already runs its own provider of the same kinds:
|
||||
|
||||
- Both control-capable nodes run their **own general-purpose postgres server**
|
||||
(the same image, one per node), serving that node's own applications.
|
||||
- Each node runs its **own** SQL server, its **own** redis, its **own** object
|
||||
store — infrastructure is per node, by design, not a single mesh-wide instance.
|
||||
- Identity is the *only* provision that is currently single (one realm on one
|
||||
node), and even that already serves applications hosted on a **second** node.
|
||||
|
||||
So the multi-provider reality is not a future edge case that appears "the day a
|
||||
second provider is added" — it is the founding topology, true today, on every
|
||||
provision kind. What is missing is any way to **say it**. A consumer requires
|
||||
`postgres-database`; it cannot require *this node's* postgres rather than *that
|
||||
node's*. It requires `oidc`; it cannot name which node's identity provider. The
|
||||
mesh model collapses a deliberately per-node fleet down to one mesh-scoped
|
||||
provider, and a consumer has no field in which to choose.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
- **The model regressed an intended topology, it did not merely miss a corner
|
||||
case.** "Node-specific services delivered to the mesh" is the design; `scope:
|
||||
"mesh"` with a single `mesh-store` expresses the opposite. This is a gap between
|
||||
a stated intent and what the manifests can represent, which is exactly what the
|
||||
issues process is for.
|
||||
- **It is not an identity special case.** The missing concept — *a provision has a
|
||||
provider, providers are per-node, and a consumer names the provider* — is the
|
||||
same for databases, object stores, brokers and identity. A fix aimed only at SSO
|
||||
would leave the same wall standing behind postgres and minio, both of which are
|
||||
*already* multi-provider on the live mesh.
|
||||
- **The single-provider assumption is silent.** Nothing rejects a second provider
|
||||
of a mesh-scoped provision; the model just cannot address it, so a consumer binds
|
||||
to whichever one is "the" provider — by accident of there being one, or by a race
|
||||
when there are two. A rule enforced by nothing ("there is one provider per
|
||||
provision") reads as true until the second node's provider makes it false, with
|
||||
no diagnostic at the seam.
|
||||
- **It blocks the migration concretely.** The mesh already has applications on one
|
||||
node depending on another node's provider (identity today; databases the moment
|
||||
an app is assigned to a node whose local postgres is not "the" mesh store).
|
||||
Modelled as mesh-scoped, that topology is expressible only by accident. To carry
|
||||
it deliberately the provision must be able to name its provider.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Where does the provider name live — on the provision definition (`scope: "node"`
|
||||
with a provider identity), on the requirement in the consumer's manifest, or
|
||||
supplied only at assignment time so the same module can be bound to different
|
||||
providers on different assignments?
|
||||
- Is "mesh-scoped" still a legitimate scope for some provisions (a single mesh CA,
|
||||
say), or does every provision become node-scoped, with a single instance
|
||||
expressed as "there happens to be one"?
|
||||
- What is the default when a consumer names no provider — bind to the node the
|
||||
consumer is assigned to (co-located provider), require the name always, or fall
|
||||
back to a mesh-wide default provider where one is declared?
|
||||
- How does a provider's identity survive being moved between nodes, so a consumer's
|
||||
recorded choice does not silently rebind when the provider relocates?
|
||||
- Does this interact with secret rotation (the issue-scope of `rotate`) — must
|
||||
rotation address a specific provider's credential holders rather than "the
|
||||
provision's"?
|
||||
Reference in New Issue
Block a user